skip to content

What does an L2 (ridge) penalty do to a linear regression's coefficients?

level: middleimportance: must knowfreq 78%

answer

  1. a price tag on coefficient size
  2. the penalty is squared, not absolute
  3. lambda added along the diagonal
  4. divide by 1 + lambda, never reach zero

basics

~20 s

A ridge penalty adds lambda times the sum of squared coefficients to the squared-error objective. Every coefficient is pulled proportionally toward zero, none lands exactly on zero, and the fit trades a little bias for much steadier estimates.

solid answer

~50 s

Ridge minimises `RSS + lambda * sum(w_j^2)` - ordinary least squares plus a squared penalty on the weights, with the intercept left out of the penalty. It has a closed form, `w = (X'X + lambda I)^-1 X'y`: adding `lambda` along the diagonal makes the system well conditioned even when the columns carry nearly the same information. The shrinkage is proportional rather than absolute - with an orthonormal design each weight becomes the least-squares weight divided by `1 + lambda` - so coefficients get smaller but never land exactly on zero, and ridge does no feature selection. At `lambda = 0` you are back at least squares; as `lambda` grows all slopes tend toward zero and predictions flatten to the intercept. You accept some bias in exchange for far less sensitivity to the particular sample you drew.

go deeper

for a junior

Recall the objective in words - squared error plus lambda times the sum of squared coefficients - and the headline consequence that every coefficient shrinks but none is set to zero.

for a middle

Be ready to write the closed form with lambda added along the diagonal, say why that makes the solve stable, and show the orthonormal case where each weight is simply divided by 1 + lambda.

for a senior

Expect to justify a chosen lambda from held-out error rather than training fit, describe what proportional shrinkage does to reported coefficient magnitudes, and carry the same penalty over to a penalised log-loss model.

for a principal

Own the call of when shrinkage is the right instrument at all, versus better features or a smaller predictor set, and set the rule for what a shrunken coefficient may be reported as outside the team.

Ridge regression is ordinary least squares with a price tag attached to the size of the coefficients. ## The objective Least squares picks the weight vector `w` minimising the residual sum of squares, `RSS = sum_i (y_i - x_i . w)^2`. Ridge minimises `RSS + lambda * sum_j w_j^2` - the same data-fit term plus `lambda` times the sum of squared weights. `lambda >= 0` is a hyperparameter you choose, not something the fit learns. Two conventions travel with it: the intercept is excluded from the penalty, and the predictors are put on a common scale first, because the penalty charges coefficient magnitude and magnitude depends on the units a feature happens to be measured in. ## The closed form Unlike most penalised fits, ridge has an exact solution: `w = (X'X + lambda I)^-1 X'y` where `X` is the n-by-p feature matrix, `X'` its transpose, and `I` the p-by-p identity. Compare it with the least-squares solution `(X'X)^-1 X'y`: the only change is `lambda` added along the diagonal of `X'X`. That single change is what makes ridge well behaved numerically. `X'X` is symmetric with non-negative eigenvalues; adding `lambda I` lifts every eigenvalue by `lambda`, so the matrix is strictly positive definite and invertible for any `lambda > 0` - even when two columns carry almost identical information, and even when there are more features than rows. ## Why the shrinkage is proportional Take the singular value decomposition `X = U D V'`, with singular values `d_1, ..., d_p`. Along the direction `v_j`, ridge multiplies the least-squares coordinate by `d_j^2 / (d_j^2 + lambda)` Every such factor lies strictly between 0 and 1: nothing is left untouched and nothing is driven to zero. Directions in which the data varies a lot - large `d_j` - have a factor near 1 and survive almost unchanged. Directions the data barely pins down - small `d_j` - are crushed. That is the whole idea: ridge spends its shrinkage where the data has least to say. In the special case of an orthonormal design, where `X'X = I` and every singular value is 1, the formula collapses to `w_ridge = w_ols / (1 + lambda)`. Every weight is divided by the same constant. This is the crispest statement of what proportional means, and it sits next to the obvious arithmetic fact that dividing a nonzero number by `1 + lambda` never yields zero. ## No exact zeros, therefore no selection Because the shrinkage is multiplicative, a coefficient that was nonzero under least squares stays nonzero under ridge for every finite `lambda`. Ridge keeps all p predictors in the model; it makes them all small rather than making some of them disappear. If you want a fit that discards features, the squared penalty is the wrong tool. If you want a fit that uses everything but trusts nothing too much, it is the right one. ## What you are buying Shrunken coefficients are biased, deliberately so: their expected values are pulled toward zero. In exchange, they move far less from sample to sample. For a wide range of problems there exists some `lambda > 0` whose out-of-sample squared error beats the least-squares fit's, and finding it is the job of held-out validation, not of theory. ## Choosing lambda Tune on held-out error - cross-validated error over a grid of `lambda` values spaced logarithmically, since the useful range spans orders of magnitude. Training error rises monotonically as `lambda` rises, so training fit can never be used to choose it. Plotting each coefficient against `lambda` gives a coefficient path: weights slide smoothly toward zero, approaching it asymptotically rather than snapping to it. At `lambda = 0` you recover least squares; at very large `lambda` the slopes vanish and the model predicts the intercept for everyone, which with centred predictors is the target's mean. ## The same penalty on other losses The squared penalty is not tied to squared error. On a classifier fitted by log loss, the objective becomes `sum_i -[y_i * log(p_i) + (1 - y_i) * log(1 - p_i)] + lambda * sum_j w_j^2` with `p_i` the predicted probability. There is no closed form, so it is minimised numerically, but the objective stays convex and the behaviour is identical in kind: on a 40-feature manufacturing quality model, every log-odds weight shrinks proportionally as `lambda` rises, none reaches exactly zero, and predicted probabilities are pulled toward the base rate carried by the unpenalised intercept. ## Common confusions The penalty is squared, not absolute - that difference is exactly why there are no zeros. Ridge does not remove collinearity or features from the data; it changes how the fit responds to them. `lambda` is not on any natural scale, so a value that is strong for one dataset is negligible for another. And a large `lambda` is not safer by default: it is simply a more heavily biased model, and only held-out error can say whether the trade paid off.

  • What happens to the ridge fit as lambda goes to zero and as it goes to infinity?
    At `lambda = 0` the penalty disappears and you recover ordinary least squares, with all its sample-to-sample instability when columns overlap. As `lambda` grows every slope is pulled toward zero and the model flattens to a constant equal to the intercept, predicting the training mean for everyone. The useful value sits between the two extremes and is chosen by held-out error, since training error only ever worsens as lambda rises.
  • How does the penalty change when ridge is applied to a logistic regression's log loss?
    The objective becomes `sum(-[y*log(p) + (1-y)*log(1-p)]) + lambda * sum(w_j^2)` - the same squared penalty bolted to a different data term. There is no closed form, so it is solved numerically, but the objective stays convex. On a 40-feature manufacturing quality model every log-odds weight shrinks proportionally, none becomes exactly zero, and predicted probabilities move toward the base rate held by the unpenalised intercept.
  • Can a ridge fit ever set a coefficient to exactly zero?
    Only in degenerate cases, such as a feature whose least-squares coordinate is already exactly zero. For any finite `lambda` the shrinkage is multiplicative, so a nonzero weight stays nonzero, however small it becomes. That is why ridge keeps every predictor in the model and cannot serve as a feature selector - exact zeros require an absolute-value penalty instead.

Think of a spring attached from each coefficient back to zero. The further out a coefficient is pulled, the harder the spring resists - but a spring never snaps a value exactly onto zero, it only holds it closer in.

saying these in an interview costs you the question

  • Says ridge performs feature selection by zeroing weights
  • Confuses the squared penalty with the absolute-value penalty
  • Claims ridge estimates are unbiased, just lower variance
  • Penalises the intercept along with the slopes
  • Assumes a larger lambda always improves test error

context