In ridge regression, what happens to the coefficients as lambda goes to 0 and as lambda grows very large?
answer
- one dial between two familiar models
- look at the closed form's diagonal term
- zero penalty means nothing changes
- huge penalty flattens the model
- shrinks proportionally, never exactly zero
basics
~20 sRidge with lambda = 0 reproduces the ordinary least squares fit. As lambda grows, every slope coefficient shrinks smoothly toward zero and the model flattens toward a constant, trading variance for bias. No coefficient reaches exactly zero at finite lambda.
solid answer
~40 sRidge minimises `sum((y_i - x_i.w)^2) + lambda * sum(w_j^2)`, and its solution is `w = (X'X + lambda I)^-1 X'y`. At `lambda = 0` the added `lambda I` term disappears and you get plain least squares: lowest bias, highest variance. As lambda grows, the penalty term dominates the inversion, so every slope is pulled proportionally toward zero and the model degenerates toward predicting a single constant — maximum bias, near-zero variance. In between, the shrinkage is smooth and multiplicative: with orthonormal predictors each coefficient becomes `w_ols / (1 + lambda)`. Crucially it never hits exactly zero at any finite lambda, so ridge shrinks but never selects — that is the L1 penalty's job. You pick lambda by cross-validation, because training error rises monotonically with it while validation error is U-shaped.
go deeper
Recall the two endpoints cleanly: lambda = 0 is ordinary least squares, very large lambda flattens the model to a constant, and everything in between is a smooth shrink toward zero without ever reaching it.
Be ready to derive the endpoints from w = (X'X + lambda I)^-1 X'y, and to state the proportional form w_ols / (1 + lambda) for orthonormal predictors. Explain why the penalty never yields exact zeros.
An interviewer expects you to connect lambda to bias-variance in operational terms: how you search the grid, why training error is useless for the choice, and what an optimum sitting at the edge of your grid is telling you.
Own the tradeoff framing: how much predictive stability you are willing to buy with bias, when a shrinkage-only method is the wrong tool because stakeholders need a short predictor list, and how lambda selection should be nested inside evaluation so it does not leak.
## What lambda controls Ridge regression fits a linear model by minimising a penalised sum of squares: `loss(w) = sum_i (y_i - x_i . w)^2 + lambda * sum_j w_j^2` The first term is the ordinary least squares (OLS) objective: the squared gap between each observed target `y_i` and the model's prediction `x_i . w` (the dot product of that row's predictor values with the coefficient vector `w`). The second term, the squared L2 penalty, charges the fit for the total squared size of its slope coefficients. `lambda >= 0` is the price per unit of squared coefficient — a single dial saying how much the data are allowed to move the weights away from zero. ## The closed form makes both limits obvious In matrix notation, with `X` the n-by-p matrix of predictors, `y` the target vector, `X'` the transpose and `I` the p-by-p identity matrix, the minimiser is `w_ridge = (X'X + lambda I)^-1 X'y` Compare OLS, which is `w_ols = (X'X)^-1 X'y`. The two differ by exactly one thing: ridge adds `lambda` to every entry on the diagonal of `X'X` before inverting. That single term is where all the behaviour comes from. ## lambda -> 0: ridge becomes OLS Set `lambda = 0` and `lambda I` vanishes, leaving the OLS formula verbatim. The fit is unbiased under the usual linear-model assumptions, but it has the highest variance of the whole ridge family: with many predictors, or with predictors that carry nearly the same information, small changes in the training data can swing the coefficients wildly. One nuance worth knowing: if `X'X` is not invertible (more predictors than rows, or an exactly redundant column), OLS has no unique solution at all, while ridge stays well defined for every `lambda > 0` — that is why ridge is often reached for on wide or collinear data. ## lambda -> infinity: everything collapses to a constant As lambda grows, `lambda I` dominates `X'X`, so `(X'X + lambda I)^-1` behaves like `(1/lambda) I` and `w_ridge` behaves like `(1/lambda) X'y`, which tends to the zero vector. Every slope is driven toward zero, and the model's predictions flatten to a single number for every input — the constant term, which is conventionally left out of the penalty. Bias is now at its maximum and variance is essentially nil: the fit ignores the predictors entirely and cannot overfit because it no longer responds to them. ## In between: smooth, proportional shrinkage Between the two extremes ridge does not clip or threshold — it scales. If the predictor columns are orthonormal, each coefficient is simply divided by `1 + lambda`: `w_ridge_j = w_ols_j / (1 + lambda)` In the general case, decomposing `X` by its singular values `d_j` gives shrinkage factor `d_j^2 / (d_j^2 + lambda)` along the j-th direction. Directions in which the data vary a lot (large `d_j`) are barely touched; directions with little variation — the ones responsible for unstable, wildly signed coefficients on near-duplicate predictors — are shrunk hardest. That asymmetry is precisely why ridge stabilises collinear fits. Because every factor `d_j^2 / (d_j^2 + lambda)` is strictly positive for any finite lambda, no coefficient is ever driven exactly to zero. Ridge shrinks; it does not select. Getting exact zeros requires an L1 (absolute-value) penalty, whose non-differentiable corner at zero is what pins coefficients there. ## Complexity as a continuous dial A useful summary is the effective degrees of freedom, `df(lambda) = sum_j d_j^2 / (d_j^2 + lambda)`. At `lambda = 0` this equals `p`, the number of predictors; as lambda grows it decreases continuously toward 0. So lambda does not switch predictors on and off — it interpolates the model's effective complexity between a full least-squares fit and a constant. ## Choosing lambda Training error rises monotonically with lambda, so it can never be used to pick the value. Held-out error is typically U-shaped: too small and variance dominates, too large and bias does. The standard procedure is k-fold cross-validation over a grid of lambdas spaced logarithmically (for example 1e-4 up to 1e4), taking the value with the best average validation score; some practitioners take the largest lambda within one standard error of the best, for a simpler model. If the winning value sits at the edge of the grid, extend the grid rather than accept it — an edge optimum means the search range, not the data, chose the answer. ## What to remember lambda = 0 is OLS. lambda = infinity is a constant model. Everything between is a smooth, proportional pull toward zero that is strongest in the least-informative directions of the data, and never quite arrives at zero.
- Why can't you choose lambda by looking at training error?Training error increases monotonically with lambda, because lambda = 0 already minimises the unpenalised squared error on that exact data. Optimising training error therefore always selects lambda = 0, which is just least squares. You need held-out data — typically k-fold cross-validation over a log-spaced grid — because only out-of-sample error shows the U-shape where reduced variance outweighs the added bias.
- At a large but finite lambda, could a ridge coefficient still be exactly zero?Only by coincidence of the data, not because of the penalty. The shrinkage factor along each direction is `d^2 / (d^2 + lambda)`, strictly positive for any finite lambda, so ridge scales coefficients down without ever pinning them at zero. A coefficient reads as exactly zero only if its least-squares counterpart already was. This is the structural difference from an L1 penalty, which does produce exact zeros.
- How does raising lambda affect the model's bias and variance?Variance falls and bias rises, monotonically in both cases. Larger lambda means the coefficients depend less on the particular training sample, so refits on different samples agree more closely — that is lower variance. But the coefficients are systematically pulled below their true values, which is bias. The best lambda is wherever the drop in variance still outweighs the added squared bias on held-out data.
Lambda is a volume knob on how loudly the predictors are allowed to speak: at zero they shout at full least-squares volume, turned all the way up they are muted to a flat constant, and every setting in between quietens them proportionally rather than silencing any one of them.
saying these in an interview costs you the question
- Says ridge sets unimportant coefficients to exactly zero
- Claims larger lambda reduces both bias and variance
- Thinks lambda = 0 gives a constant-only model
- Picks lambda by whichever value minimises training error
- Believes shrinkage hits every coefficient by the same absolute amount
- Confuses lambda with the number of predictors kept