Why is the intercept left unpenalised in ridge and lasso regression?
answer
- not a slope, a level
- zero is not a neutral value
- shift every target by 100
- huge lambda should predict the mean
basics
~20 sThe intercept only sets the model's overall level, and zero is not a neutral level. Shrinking it would drag every prediction toward zero and make the fit depend on the target's arbitrary origin, so it is fitted freely.
solid answer
~50 sA slope answers "how much does the prediction move per unit of this predictor", and pulling it toward zero is a defensible statement of ignorance about that effect. The intercept answers "what is the overall level of the target", and zero is not a neutral answer to that question: a basket-value model whose target averages 58 currency units would be biased low on every single row if its intercept were shrunk. Penalising it also breaks location invariance: add 100 to every target value and you get different slopes rather than the same fit shifted up. The standard mechanics are to centre the predictors and the target, fit the penalised slopes on the centred problem, then recover `b0 = mean(y) - sum_j b_j * mean(x_j)`. As lambda grows the model then collapses to the mean of the target, which is the sensible null model, instead of collapsing to zero.
go deeper
Recall the rule and one consequence: the intercept is fitted freely, and if it were shrunk toward zero every prediction would be dragged down with it.
Be able to derive it. Show that an unpenalised intercept makes the fit shift cleanly when the target shifts, and that a very large lambda then leaves you predicting the training mean.
Show you can spot the symptom in production: predictions biased low across the board, or a rare-event classifier whose probabilities bunch near 0.5, both point at an intercept being shrunk or a target that was never centred.
Frame the tradeoff for the team. The penalty buys variance reduction on effects, not on the level; a model whose overall level is uncertain needs better data or an explicit prior, not a larger lambda.
## Two different kinds of parameter A penalised linear model has parameters of two distinct kinds, and the penalty is a statement about only one of them. The slopes `b_1 ... b_p` describe *effects*: how much the prediction moves when a predictor moves by one unit. Shrinking a slope toward zero encodes a prior belief that the effect is probably small — that is precisely the bias you accept in exchange for lower variance, and it is the whole point of ridge and lasso. The intercept `b0` describes a *level*: where the whole prediction surface sits on the target's axis. Shrinking it toward zero encodes the belief that the target is probably near zero, which is not a belief about effects at all — it is a belief about the arbitrary origin of the measurement scale. Basket value, house price, blood pressure and revenue are all quantities whose zero point carries no special meaning for the model. ## What penalising the intercept would actually do Consider a basket-value model whose target averages 58 currency units. Add `lambda * b0^2` to the objective and the fitted intercept is pulled below the value that best centres the residuals. Since every prediction is `b0 + sum_j b_j x_j`, the whole prediction cloud slides downward: a systematic, uniform, unnecessary bias on every row, and one that grows as you tune lambda upward for perfectly good reasons about the slopes. Worse, the fit stops being location-invariant. With an unpenalised intercept, adding a constant `k` to every target value leaves all slopes untouched and simply moves the intercept to `b0 + k` — the same model, re-expressed. Penalise the intercept and that clean equivariance disappears: the amount of intercept shrinkage now depends on how large the raw target numbers happen to be, so measuring temperature in Celsius versus Kelvin, or expressing a price in cents versus dollars, would produce genuinely different slopes. A modelling choice would be leaking in from the choice of origin. There is also a dimensional argument. `b0` has the units of the target, while `b_j` has units of target per unit of predictor `j`. Adding `b0^2` into `sum_j b_j^2` sums quantities that are not commensurable — the sum has no consistent interpretation. ## How the mechanics are usually arranged The common formulation removes the intercept from the optimisation entirely rather than special-casing it: 1. Centre every predictor by subtracting its training mean, and centre the target by subtracting `mean(y)`. 2. Fit the penalised slopes on the centred problem. With both sides centred, the optimal intercept of that problem is exactly zero, so there is nothing to penalise or protect. 3. Recover the intercept for the original scale as `b0 = mean(y) - sum_j b_j * mean(x_j)`. One useful consequence falls straight out of this. Push lambda toward infinity and all the slopes go to zero, so `b0 -> mean(y)` and the model predicts the training mean for every row. That is the correct null model — the best constant predictor under squared error — and it is the sensible endpoint of the bias-variance dial. If the intercept were penalised, the same limit would predict zero, a model that is not merely simple but wrong. ## The classification case The argument is, if anything, sharper for penalised logistic regression. There the intercept sets the baseline log-odds, and log-odds of zero means a predicted probability of exactly 0.5. Penalise the intercept on a 1%-prevalence fraud problem and you pull every predicted probability toward a coin flip — the ranking may survive, but calibration is destroyed and any threshold or expected-value calculation built on those probabilities goes with it. Leaving the intercept free lets the model absorb the base rate and spend the penalty budget on the slopes, where it belongs. ## When you might reconsider The rare case where penalising the intercept is defensible is one where zero genuinely is the meaningful level — a target that is already a residual, a deviation, or a deliberately centred quantity, so that a non-zero offset really is suspicious. Even then the cleaner move is to centre the target, which makes the intercept zero by construction, rather than to fight it with a penalty term. ## Symptoms to recognise If a penalised regression's predictions are uniformly biased toward zero, or a penalised classifier's probabilities cluster around 0.5 on an imbalanced problem, suspect an intercept that is being shrunk, or a pipeline that centred the predictors but forgot the target.
- What does penalising the intercept do to a logistic regression on rare events?It pulls the intercept toward zero, and a zero intercept in a logistic model means baseline log-odds of zero, that is a predicted probability of 0.5. On a 1%-prevalence problem that wrecks calibration: predicted probabilities come out far too high on every row, and any expected-value threshold built on them is wrong. Leaving the intercept free lets the model match the base rate and shrink only the slopes.
- Where does the intercept go if you centre both the predictors and the target before fitting?It becomes zero and drops out of the optimisation, which is exactly why the centred formulation is the usual one: there is then no intercept in the penalised problem to protect. You recover it afterwards as `mean(y) - sum_j b_j * mean(x_j)` so that predictions come back on the original scale.
- Is there ever a case for penalising the intercept?Rarely, and only when zero is a genuinely meaningful level for the target — an already-centred quantity, a residual, or a deviation, where a non-zero overall offset really is suspicious. Even then the cleaner move is to centre the target so the intercept is zero by construction, rather than to add a penalty term and accept the loss of location invariance.
The intercept is the sea level of the model, not a current within it. Shrinking the currents makes the model more cautious; shrinking sea level toward zero just drains the ocean.
saying these in an interview costs you the question
- Says the intercept is just another coefficient, so penalise it
- Thinks shrinking the intercept has no effect on predictions
- Claims a very large lambda should make the model predict zero
- Confuses centring the predictors with penalising the intercept
- Ignores calibration damage when a classifier's intercept is shrunk