Why does an L1 penalty on a linear model's coefficients drive some of them to exactly zero?
answer
- compare the penalties near zero
- slope of |w| versus slope of w^2
- one force fades, the other does not
- soft threshold: subtract lambda, clip
- diamond corners sit on the axes
basics
~20 sAn L1 penalty adds lambda times the absolute value of each coefficient, and the slope of that term stays at lambda right up to zero. That constant pull can push a coefficient exactly to zero; an L2 penalty's pull fades away and never does.
solid answer
~50 sLasso minimises `RSS + lambda * sum|w_j|`; ridge minimises `RSS + lambda * sum w_j^2`. The difference lives at zero. The derivative of `lambda*|w|` is plus or minus `lambda` no matter how small `w` gets, so the penalty keeps pushing with full force; the derivative of `lambda*w^2` is `2*lambda*w`, which vanishes as the coefficient shrinks. So under L1 a coefficient is set to zero whenever the data's pull on it — the absolute correlation between that column and the current residual — is no bigger than `lambda`. On an orthonormal design this is exactly soft thresholding: `w_j = sign(b_j) * max(|b_j| - lambda, 0)`, where `b_j` is the least-squares value. The usual picture says the same thing: the L1 budget region is a diamond whose corners lie on the axes, and the loss contours tend to touch it at a corner, where some coordinates are zero.
code
python · 12 linesdef soft_threshold(b, lam):
if b > lam:
return b - lam
if b < -lam:
return b + lam
return 0.0
lam = 0.30
for b in (0.80, -0.45, 0.22, 0.30):
lasso = soft_threshold(b, lam)
ridge = b / (1 + lam)
print(f"ols={b:>6.2f} lasso={lasso:>6.2f} ridge={ridge:>6.3f}")go deeper
Recall the two objectives and the headline difference: absolute values give exact zeros and therefore a shorter model, squares only shrink. Being able to state which penalty is which already puts you ahead on a screening call.
This is your tier. Explain the mechanics: the constant slope of the absolute-value term versus the vanishing slope of the squared term, the soft-threshold formula, and the diamond-versus-disc picture. Work a small numeric example unprompted.
Show you know what the zeros cost you: survivors are biased toward zero by roughly lambda, the sparse set shifts with lambda and with the sample, and a zero is not evidence of irrelevance. Say how you would choose lambda and check the result holds up.
Own the framing decision: whether the team wants a sparse model because it is genuinely the better predictor or because a short coefficient list is easier to socialise, and what you owe stakeholders when those two goals disagree.
## The two objectives Ordinary least squares chooses coefficients `w` to minimise the residual sum of squares, `RSS = sum over rows of (y - w.x)^2`. Two penalised variants add a term that charges you for coefficient size: - **Lasso**: minimise `RSS + lambda * sum_j |w_j|` (an L1 penalty). - **Ridge**: minimise `RSS + lambda * sum_j w_j^2` (an L2 penalty). `lambda >= 0` is a tuning knob: at `lambda = 0` both reduce to least squares; as `lambda` grows, coefficients are pulled toward zero. The only difference between the two is absolute value versus square, and that single change is what decides whether coefficients merely get small or actually reach zero. ## The slope argument — the one to say out loud Think about the force each penalty exerts on a coefficient that is already tiny. - For the squared penalty, the derivative of `lambda*w^2` is `2*lambda*w`. As `w` approaches zero, that force approaches zero too. Any non-zero pull from the data is therefore enough to hold the coefficient at some small non-zero value. Ridge shrinks proportionally and, for finite `lambda`, never quite arrives. - For the absolute-value penalty, the derivative of `lambda*|w|` is `+lambda` for positive `w` and `-lambda` for negative `w`. The force does not shrink as the coefficient does; it stays at full strength all the way in. There is no derivative at zero at all — the function has a kink — and that kink is what lets zero be a genuine optimum. Formally, zero is optimal for coefficient `j` exactly when the magnitude of the correlation between column `j` and the current residual is at most `lambda`. Below that threshold the data cannot pay the fixed toll the penalty charges for leaving zero, so the coefficient stays at zero rather than taking a small value. ## Worked mechanics: soft thresholding When the predictors are orthonormal (uncorrelated columns of unit length) the lasso solution has a closed form, and it is the cleanest way to see the effect: `w_j = sign(b_j) * max(|b_j| - lambda, 0)` where `b_j` is the least-squares coefficient. This is called the **soft-threshold operator**: move every coefficient `lambda` units toward zero and stop it there. With `lambda = 0.30`, a least-squares coefficient of `0.80` becomes `0.50`, one of `-0.45` becomes `-0.15`, and one of `0.22` becomes exactly `0` because the shrinkage overshoots. Ridge on the same design gives `b_j / (1 + lambda)`: `0.80` becomes `0.615` and `0.22` becomes `0.169` — smaller, but never zero. Two consequences fall straight out. First, sparsity is a threshold effect: whether a coefficient survives depends on its size relative to `lambda`, not on any separate selection rule. Second, the survivors are **biased** — each is pulled about `lambda` toward zero — so lasso coefficients understate effect sizes and should not be read as unbiased estimates. ## The geometric picture The standard diagram draws the penalty as a budget on total coefficient size. For two predictors, the L1 budget region `|w1| + |w2| <= t` is a diamond with its corners sitting on the coordinate axes; the L2 region `w1^2 + w2^2 <= t` is a disc. The squared-error loss has elliptical contours centred on the least-squares solution. You expand those contours until they first touch the region. A diamond's corners are sharp, so a wide range of ellipse orientations touches a corner first — and at a corner one coordinate is exactly zero. A disc has no corners; the first touch happens at a smooth boundary point where both coordinates are generally non-zero. In higher dimensions the L1 region is a many-cornered polytope whose faces, edges and vertices each set different subsets of coefficients to zero, which is why lasso produces sparse solutions rather than just one zero. ## Sparsity as selection built into fitting Because the zeros arrive during the fit, lasso does variable selection and estimation in one pass: the fitted model simply ignores the zeroed columns. How many survive is a function of `lambda` — with `lambda = 0` you get the full least-squares fit, and with `lambda` large enough every slope is zero and the model predicts the intercept alone. In an HR attrition model built on a 400-item engagement survey, a well-chosen `lambda` can leave roughly a dozen non-zero drivers out of the 400, which is exactly the appeal: a short, readable model straight out of the optimiser. ## What a zero does not mean A zero coefficient is not a significance test and not proof of irrelevance. It says only: *given the other columns currently in the model, at this value of lambda, this column does not earn its penalty*. Drop a correlated partner and the zeroed column can come straight back. It is also worth remembering that the penalty sums raw coefficient magnitudes, so the whole procedure depends on the scale of the columns — predictors are standardised before fitting so the toll is comparable across them.
- How does reading lasso as a MAP estimate under a Laplace prior explain the exact zeros?Putting a zero-mean Laplace (double-exponential) prior on each coefficient makes the log-prior proportional to minus the sum of absolute values, so the posterior mode is the lasso solution. That density has a sharp spike at zero instead of a smooth top, and a spike is exactly what lets the mode sit at zero rather than merely near it. A smooth bell-shaped prior gives a mode that is shrunk but non-zero.
- Are the coefficients lasso keeps unbiased estimates of their effects?No. Every surviving coefficient is pulled toward zero — by exactly lambda on an orthonormal design — so magnitudes systematically understate effect sizes and standard inference on them is not valid. If the numbers themselves matter, one common move is to refit an unpenalised model on the selected columns to undo the shrinkage, accepting that the selection step was still made on the same data.
- What does the lasso solution look like as lambda grows very large?Every slope is driven to zero and the model collapses to the intercept, predicting the mean response for everyone. Coming the other way, at lambda = 0 you recover the ordinary least-squares fit. So lambda traces a family of models from fully dense to fully empty, and choosing it is choosing where on that range you want to sit.
The squared penalty is a spring: the closer to zero, the weaker the pull. The absolute-value penalty is friction: it pushes with the same force right up to the stop, so things actually come to rest at zero.
saying these in an interview costs you the question
- Says a large enough L2 penalty also produces exact zeros
- Explains it only as 'L1 is more aggressive than L2'
- Thinks the zeros come from numerical rounding of tiny values
- Describes L1 shrinkage as proportional to coefficient size
- Swaps the shapes: circle for L1, diamond for L2