In a lasso coefficient path drawn with lambda decreasing left to right, what do the two ends show?
answer
- one line per predictor, not per row
- the horizontal axis is penalty strength
- one end is a constant-only model
- the other end pays no penalty at all
basics
~20 sAt the far left the penalty is largest and every coefficient is exactly zero, so the model predicts a constant. Moving right, predictors enter one by one, and the far right is the unpenalised least-squares fit.
solid answer
~50 sA coefficient path draws one line per predictor: the fitted coefficient as a function of the penalty strength lambda. With lambda decreasing left to right, the left edge sits at `lambda_max`, the smallest penalty that drives every coefficient to zero, so the model there is just the intercept. As lambda falls, predictors switch on one at a time and their lines move away from zero; at the right edge, where lambda is effectively zero, you recover the ordinary least-squares fit. On a used-car listing-price model you would typically see mileage enter first, then model year, then trim level, with colour arriving only at the far right. That axis is a bias-variance dial: the all-zero left end is maximally biased with almost no variance, the unpenalised right end has the lowest bias and the highest variance, and a useful model normally sits somewhere in between.
go deeper
Be ready to say what sits at each end: an all-zero, intercept-only model at the largest penalty and the plain least-squares fit at zero penalty. Knowing that lasso lines can hit exactly zero while ridge lines do not is enough here.
Explain the mechanics: why a finite lambda_max zeroes everything, why predictors enter one at a time, and why the lasso path is straight lines between knots while a ridge path is smooth. Expect to contrast the two penalties on the same plot.
Show how you use the plot in real work — sanity-checking that the predictors you expect enter early, spotting that a near-unpenalised end is unstable, and refusing to pick lambda by eye when an out-of-sample estimate is what decides it.
Own the framing: the path is the cheapest artefact for arguing capacity tradeoffs with non-specialists, and the temptation to present it as a feature-importance ranking is the thing to head off before it reaches a slide deck.
## What a coefficient path is A penalised linear model does not produce one set of coefficients; it produces a *family* of them, one for every value of the penalty strength lambda. The lasso minimises ``` (1 / 2n) * sum_i (y_i - b0 - sum_j x_ij * b_j)^2 + lambda * sum_j |b_j| ``` and the solution `b(lambda)` moves continuously as you turn lambda. A **coefficient path plot** simply draws that family: the horizontal axis is the penalty (usually `log(lambda)`, or the total L1 norm of the solution), the vertical axis is coefficient value, and each predictor contributes one line. It is the single most informative picture of what a penalty is doing to your model. ## The left end: lambda_max and the null model There is a finite penalty above which nothing survives. Call it `lambda_max`: it is the smallest lambda for which the whole coefficient vector is zero, and it equals the largest absolute inner product between any single predictor column and the response, scaled by the sample size (with the `1 / 2n` least-squares convention above). Intuitively, a predictor only earns a nonzero coefficient when the squared-error it removes outweighs the `lambda * |b_j|` it must pay; when lambda exceeds the best predictor's payoff, no predictor can pay, and every coefficient is zero. So at the far-left edge the fitted model is a constant — the intercept, which is left unpenalised — and it predicts the mean response for every row. That model has no variance worth mentioning (resample the data and it barely moves) and enormous bias if the predictors really do carry signal. ## The right end: the unpenalised fit At the other extreme lambda goes to zero, the penalty term vanishes, and the objective reduces to plain least squares. Every predictor is in the model at its full unshrunk coefficient. This end has the lowest bias the linear family can offer and the highest variance: with many predictors, or with correlated ones, small changes in the training rows swing these coefficients a lot. When predictors outnumber rows the least-squares fit is not even unique, which is why paths are usually stopped short of lambda = 0. ## The middle: entry order and shrinkage Between the ends, two things happen at once. Predictors **enter** — their coefficients leave zero — and predictors already in the model keep **growing** away from zero as the penalty relaxes. For a used-car listing-price model, mileage might enter first, model year next, trim level later, and colour only near the unpenalised end; a reader can see at a glance which handful of predictors carries most of the signal at any given penalty. Two mechanical facts are worth knowing. First, the lasso path is *piecewise linear* in lambda: between the penalty values where the set of nonzero coefficients changes (the knots), every coefficient traces a straight line. Second, the ridge (L2) path looks quite different — coefficients shrink smoothly toward zero but generically never reach it, so all lines are nonzero everywhere and there is no entry order at all. If a path plot shows lines flattening onto zero at distinct points, you are looking at an L1 penalty. Paths are conventionally drawn after standardising the predictors, so that the vertical positions of different lines are comparable; that convention is a topic in its own right. ## Reading the axis as a dial The most useful habit is to read the horizontal axis as a capacity dial. Sliding left buys stability at the price of systematic error; sliding right buys fit at the price of instability. Neither end is the answer. Where exactly to stop is decided by out-of-sample error, not by looking at the coefficient lines — the path tells you *what the model becomes* at each penalty, not *which penalty generalises best*. ## Common misreadings The path is not a training history: nothing is iterating, and there is no time axis. Moving right is not "the model learning"; it is you choosing a weaker penalty. Also, a coefficient sitting at zero over most of the path does not prove its predictor is unrelated to the response — it may simply be redundant given the predictors that entered earlier.
- How does the same plot look if you swap the lasso's L1 penalty for a ridge L2 penalty?Every line is nonzero across the whole plot. Ridge shrinks coefficients smoothly toward zero as the penalty grows but generically never sets one exactly to zero, so there is no entry order and no sparsity: all predictors are present at every penalty, just progressively smaller. The lines are also smooth curves rather than the lasso's straight segments joined at knots.
- What exactly is lambda_max, and why is the solution there known without fitting anything?It is the smallest penalty that makes the whole coefficient vector zero, equal to the largest absolute predictor-response inner product scaled by the sample size. Above it, no predictor removes enough squared error to pay its own penalty, so the optimum is the all-zero vector. That makes the left edge a free, exact starting point rather than something you have to solve for.
- Can you pick the deployed model by eye from the path plot alone?No. The path shows what the model becomes at each penalty, not how well each one generalises. Choosing a penalty needs an out-of-sample estimate of error over the same lambda grid; the path is then read at the winning lambda to see which predictors survived and how large their coefficients are.
It is a dimmer switch photographed at every setting: full dark is the intercept-only model, full brightness is the unpenalised fit, and each predictor's light comes on at its own point on the dial.
saying these in an interview costs you the question
- Treats the unpenalised right end as the best model
- Reads the horizontal axis as training iterations or epochs
- Says ridge paths also drive coefficients to exactly zero
- Claims a coefficient stuck at zero proves no relationship exists
- Thinks every predictor enters the path at the same penalty