In a lasso, what happens to the fitted coefficients as the penalty strength lambda increases?
answer
- one dial, two effects at once
- lambda = 0 is plain least squares
- constant pull all the way to zero
- above lambda_max only the intercept
- picked by cross-validation, never training error
basics
~20 sRaising lambda shrinks every coefficient toward zero and pushes more of them to exactly zero, so the fit uses fewer predictors. Past a large enough lambda every slope is zero and only the intercept survives.
solid answer
~40 sLambda weights the absolute-size penalty in `loss = RSS + lambda * sum(|w_j|)`. At `lambda = 0` you recover the ordinary least-squares fit. As lambda grows, every coefficient is pulled toward zero, and because the L1 penalty pushes with constant force `lambda` right up to zero, any predictor whose least-squares signal is weaker than that threshold snaps to exactly zero instead of merely becoming small. So the fit gets simultaneously more biased and sparser: the count of nonzero coefficients trends down as lambda rises, though individual predictors can drop out and re-enter along the way. Above some finite `lambda_max` nothing is left but the unpenalised intercept, and the model just predicts the mean. Because bias rises while variance falls, lambda is picked by cross-validated out-of-sample error, not by looking at training fit.
go deeper
Be able to say what the dial does in both directions: lambda = 0 is plain least squares, large lambda drives coefficients to exactly zero until only the intercept is left, and the value is chosen by cross-validation.
Explain why the L1 penalty produces exact zeros while a squared penalty does not — the constant-size pull of an absolute value versus a pull that fades near zero — and be able to describe the regularisation path.
Show judgment about picking a point on the path in production: cross-validating over a sensible lambda grid, using the one-standard-error rule when stability matters, and reading an unstable active set as a warning about correlated predictors.
Own the tradeoff between predictive accuracy and the operational value of a short, stable feature list — fewer inputs mean less pipeline to maintain and monitor, which can be worth accepting a slightly worse cross-validated score.
## The objective lambda sits in A lasso is ordinary linear regression with one extra term bolted onto the thing being minimised: ``` loss(w) = sum_i (y_i - w0 - w'x_i)^2 + lambda * sum_j |w_j| ``` The first term is the residual sum of squares (RSS) — the usual squared prediction error that plain least squares alone minimises. The second is the **L1 penalty**: the sum of the absolute values of the slope coefficients, multiplied by a non-negative number `lambda`. The intercept `w0` is normally left out of the penalty sum. `lambda` is not learned from the fit; it is a hyperparameter you set from outside, and it is the single dial that decides how much simplicity the fit is willing to buy with accuracy. ## The two ends of the dial At **lambda = 0** the penalty term vanishes and the objective is exactly RSS, so the lasso returns the ordinary least-squares solution: no shrinkage, no zeros, maximum fit to the training data and maximum variance across resamples. At the other extreme there is a finite value — call it `lambda_max` — above which *every* slope is exactly zero and the fitted model is just the intercept, predicting the same number (the mean of y) for every row. For centred, standardised predictors this threshold is roughly `max_j |x_j' y| / n`: the strongest single correlation between a predictor and the target. Once the penalty's pull exceeds even the best predictor's pull, nothing can hold a nonzero value. ## What happens in between: shrink and select at the same time Between those ends, raising lambda does two things at once. **Shrinkage.** Every surviving coefficient is smaller in absolute value than its least-squares counterpart. The penalty makes size itself expensive, so the fit only pays for magnitude that buys enough RSS reduction. **Selection.** Unlike a squared (L2 / ridge) penalty, whose pull weakens to nothing as a coefficient approaches zero, the absolute-value penalty pushes with the same force `lambda` no matter how small the coefficient is. The tug-of-war therefore has a threshold: if a predictor's marginal contribution to reducing RSS is smaller than `lambda`, the penalty wins outright and the coefficient lands on exactly zero, not merely near it. In the clean orthogonal case this is literally the soft-thresholding rule ``` w_j <- sign(z_j) * max(|z_j| - lambda, 0) ``` where `z_j` is the unpenalised least-squares coefficient: subtract lambda from the magnitude, and clip at zero. That is why the lasso does variable selection as a side effect of fitting rather than as a separate step. ## The path, and why it is not perfectly tidy Sweeping lambda from `lambda_max` down to 0 and recording the coefficients gives the **regularisation path** (or lasso path): a plot of each coefficient against lambda, starting all-zero on the left and fanning out to the least-squares values on the right. The number of nonzero coefficients trends downward as lambda increases, but it is not strictly monotone — with correlated predictors, one variable can leave the active set and another enter as lambda changes, and a coefficient can even re-enter after having been zeroed. Treat "more lambda means fewer variables" as a strong tendency, not a guarantee. ## Bias, variance, and choosing lambda Lambda traces the bias-variance tradeoff. Small lambda gives a low-bias, high-variance fit that can chase noise; large lambda gives a stable, heavily biased fit that can miss real signal (underfitting). Because training RSS is monotonically *worsened* by any lambda above zero, you can never pick lambda by looking at training error — it would always choose zero. The standard procedure is k-fold cross-validation: fit at a grid of lambda values, score each on held-out folds, and take the lambda with the best out-of-sample error, or the largest lambda within one standard error of it when you deliberately want a simpler, sparser model. ## Two things lambda does not do It does not rank importance by itself — a coefficient surviving to a high lambda is evidence of a strong, non-redundant signal, but among near-duplicate predictors the survivor is somewhat arbitrary. And it does not make the fit unit-free: lambda multiplies raw coefficient magnitudes, so what counts as "a large lambda" depends entirely on the scale the predictors are measured on.
- Why can't you choose lambda by whichever value minimises training error?Training RSS is smallest at lambda = 0 by construction — any positive penalty moves the fit away from the least-squares optimum, so that criterion always returns zero regularisation and defeats the point. You need an estimate of error on data the fit has not seen, which is why k-fold cross-validation over a grid of lambda values is the standard choice, sometimes with the one-standard-error rule to prefer a simpler model.
- At the same lambda, how does a ridge penalty's effect on the coefficients differ?A ridge penalty penalises squared coefficients, so its pull shrinks in proportion to the coefficient and fades to nothing as the value approaches zero. Coefficients therefore get asymptotically small but essentially never land on exactly zero. Ridge shrinks without selecting: you still carry every predictor in the model, just with damped weights, so it cannot hand you a short list of variables.
- If you double lambda and the number of surviving predictors goes up, is something broken?Not necessarily. The active set is not a strictly monotone function of lambda when predictors are correlated: a variable can be dropped and a correlated one picked up, and a coefficient can re-enter the path. Small increases are normal on real data. What would be suspicious is a large, erratic swing between adjacent grid points, which usually means unstable near-duplicate predictors or a fit that has not converged.
Lambda is a budget cut applied to a team of predictors: everyone's hours get trimmed, and the ones who cannot justify even their minimum hours are cut to zero rather than kept on part-time.
saying these in an interview costs you the question
- Says lambda only rescales predictions, not coefficients
- Claims larger lambda always yields a lower training error
- Thinks lambda is learned during fitting like a weight
- Assumes ridge also zeroes coefficients at large lambda
- Picks lambda by training fit instead of held-out error
- Believes the nonzero count falls strictly monotonically
- Says lambda = 0 still leaves some shrinkage in place