skip to content

Scaling and the Intercept

A penalty sums coefficients on their raw scales, so grams and kilograms are punished differently and the intercept has to stay out of it. Interviewers use scaling as a quick tell for real practice.

on this pageshow

questions

3

Why must predictors be standardised before fitting a ridge or lasso model?

level: juniorimportance: must knowfreq 72%

answer

  1. the penalty sums coefficients, not effects
  2. units live inside a coefficient
  3. dollars versus thousands of dollars
  4. narrowest-spread predictor is shrunk hardest

basics

~20 s

An L1 or L2 penalty adds up coefficient sizes, and a coefficient's size depends on the units of its predictor. Standardising puts every predictor on a common spread so one lambda penalises them all comparably.

solid answer

~50 s

Penalised regression minimises `RSS + lambda * penalty(coefficients)`, and a coefficient carries units of target per unit of predictor. Measure a predictor in thousands instead of ones and its coefficient becomes a thousand times larger, so a fixed lambda charges it about a million times more under an L2 penalty and a thousand times more under L1. On raw data the penalty therefore falls hardest on the narrowest-spread predictors and barely touches wide-ranging ones such as a daily step count: the shrinkage ordering is decided by measurement units rather than by predictive value. Centring each predictor and dividing by its standard deviation makes lambda mean the same thing for every coefficient and makes the fitted coefficients comparable with each other. Unpenalised least squares needs none of this — rescale a predictor and its coefficient simply rescales back to the identical fit.

code

python · 17 lines
python
import random
random.seed(0)

# raw units: BMI ~27 (spread 4), age ~55 (spread 12), steps ~8000 (spread 2500)
rows = [(random.gauss(27, 4), random.gauss(55, 12), random.gauss(8000, 2500))
        for _ in range(1000)]
lam = 20000.0

for j, name in enumerate(("bmi", "age", "steps")):
    col = [r[j] for r in rows]
    mean = sum(col) / len(col)
    sum_sq = sum((v - mean) ** 2 for v in col)
    print(name, "sum_sq =", round(sum_sq), "fraction kept =", round(sum_sq / (sum_sq + lam), 3))

# bmi   sum_sq = 16079        fraction kept = 0.446
# age   sum_sq = 147689       fraction kept = 0.881
# steps sum_sq = 6195362333   fraction kept = 1.0

go deeper

for a junior

Be ready to state the rule and its reason in one breath: the penalty adds coefficients up, coefficients carry the units of their predictor, so every predictor goes onto a common scale before the fit.

for a middle

Explain the arithmetic out loud — multiplying a predictor by c divides its coefficient by c, which moves its L2 penalty term by c squared — and say which predictor gets shrunk hardest on raw data and why.

for a senior

Expect to talk about operating it: storing training means and spreads with the model, retuning lambda whenever a feature's units or spread shift upstream, and translating standardised coefficients back into raw units for stakeholders.

for a principal

Own the standard. Decide whether feature scales are a contract enforced upstream or a step each model repeats, and be able to say what a silent unit change in a shared feature costs a fleet of penalised models.

## The objective, and where units hide in it Ridge and lasso fit a linear model by minimising a sum of two terms: ``` objective = sum_i (y_i - b0 - sum_j b_j * x_ij)^2 + lambda * P(b) ``` where `P(b) = sum_j b_j^2` is the L2 (ridge) penalty, `P(b) = sum_j |b_j|` is the L1 (lasso) penalty, and `lambda >= 0` is the dial that trades fit against coefficient size. The first term, the residual sum of squares, is measured in squared units of the target. The second term is measured in units of the coefficients — and a coefficient `b_j` has units of *target per unit of predictor j*. That is the whole problem. A predictor's units are a human choice with no statistical content: dollars or thousands of dollars, grams or kilograms, steps or thousands of steps. The residual sum of squares does not care, because whatever you do to `x_j`, the fitted `b_j` moves inversely and the predictions are unchanged. The penalty does care, because it looks at `b_j` directly. ## The algebra of a unit change Multiply predictor `j` by a constant `c`. To reproduce exactly the same predictions, its coefficient must become `b_j / c`. So: - an L2 penalty term moves from `b_j^2` to `b_j^2 / c^2`; - an L1 penalty term moves from `|b_j|` to `|b_j| / c`. Re-expressing household income in thousands of dollars rather than dollars means `c = 1/1000`, so the coefficient grows by 1,000 and its L2 contribution grows by 1,000^2 — a factor of a million. With lambda held fixed, that predictor is now being charged a million times more for the same underlying effect, and it will be shrunk toward zero far harder. Nothing about the data changed; only the label on the axis did. Run it the other direction and you get the equivalence rule: if you rescale *every* predictor by `c`, you must move lambda to `c^2 * lambda` (ridge) or `c * lambda` (lasso) to recover the identical fit. A lambda grid tuned on one encoding of the data is simply meaningless on another. ## Which coefficient gets crushed, and which rides free The direction trips people up, so it is worth pinning down. For centred, mutually uncorrelated columns, ridge keeps a fraction of each least-squares coefficient: ``` b_ridge_j = b_ols_j * S_j / (S_j + lambda), S_j = sum_i (x_ij - mean_j)^2 ``` `S_j` is the column's centred sum of squares, which grows with the predictor's *spread* in its chosen units. A wide-ranging column has a huge `S_j`, so `S_j / (S_j + lambda)` is essentially 1 and the penalty barely touches it. A narrow column has a small `S_j` and is shrunk hard. Take a health-risk model on raw units: BMI with a spread of about 4, age with a spread of about 12, daily step count with a spread of about 2,500. The step count's sum of squares is hundreds of thousands of times larger than BMI's, so at any lambda that meaningfully shrinks BMI, the step-count coefficient is effectively unpenalised. You have not chosen which effects to regularise; the measuring instruments chose for you. Note that it is the **spread**, not the average level, that matters, because predictors are centred before the penalised fit and the mean is absorbed by the intercept. A variable averaging 8,000 that only varies by two units behaves like a small-scale predictor. ## What standardising fixes Centring each predictor and dividing by its training standard deviation gives every column the same spread, so `S_j` is the same for all of them. Now a single lambda applies the same amount of shrinkage pressure everywhere, the coefficients are on a comparable footing (each is an effect per standard deviation), and lasso's variable selection reflects predictive strength rather than which columns happen to be quoted in small units. It also makes the tuned lambda a property of the problem rather than of the spreadsheet. Three operational consequences follow. First, keep the training means and standard deviations with the model — scoring rows must be transformed with exactly those numbers. Second, to report an effect in original units, divide the fitted coefficient by the predictor's standard deviation, and correct the intercept by subtracting the sum of those rescaled coefficients times the predictor means. Third, if a feature's units or spread change upstream, the old lambda is stale and must be retuned. One column is deliberately left out of all this: the intercept, which is fitted without a penalty so the model can still represent the overall level of the target.

  • If you rescale every predictor by the same factor, how must lambda change to reproduce the same fit?
    Multiply every predictor by `c` and each coefficient becomes `b/c`, so an L2 penalty must move to `c^2 * lambda` and an L1 penalty to `c * lambda` to give the identical solution. Dividing every predictor by 1,000 therefore needs lambda divided by a million under ridge. That is the precise sense in which a lambda grid tuned on one encoding of the data is worthless on another.
  • Is it a predictor's average level or its spread that decides how hard the penalty hits it?
    The spread. Predictors are centred before a penalised fit, so the average level is absorbed by the intercept and drops out of the penalty entirely. A variable averaging 8,000 that varies by only a couple of units behaves like a small-scale predictor and gets shrunk hard. Compare standard deviations across columns, never averages.
  • After fitting on standardised predictors, how do you report a coefficient in the original units?
    Divide the fitted coefficient by that predictor's training standard deviation to get an effect per raw unit, then adjust the intercept by subtracting the sum of those rescaled coefficients times the predictor means. Keep the training means and standard deviations alongside the model, because scoring rows must be transformed with exactly those numbers rather than statistics recomputed on fresh data.

Lambda is a single fine levied on every coefficient. If some predictors are quoted in dollars and others in thousands of dollars, the fine lands unevenly for no reason other than the units printed on the price tag.

saying these in an interview costs you the question

  • Says penalised regression is scale-invariant, like least squares
  • Claims scaling only matters for gradient-descent convergence speed
  • Compares average levels across predictors instead of spreads
  • Thinks large-unit predictors are the ones the penalty crushes
  • Tunes lambda first, then changes a feature's units and reuses it

context

open as a page

Why is the intercept left unpenalised in ridge and lasso regression?

level: middleimportance: should knowfreq 48%

basics

~20 s

The intercept only sets the model's overall level, and zero is not a neutral level. Shrinking it would drag every prediction toward zero and make the fit depend on the target's arbitrary origin, so it is fitted freely.

open as a page

Should one-hot dummies and rare binary flags be standardised before a penalised fit?

level: seniorimportance: nice to knowfreq 24%

basics

~10 s

There is no universal answer. Dividing a rare 0/1 flag by its tiny standard deviation stretches its range and lets it escape most of the shrinkage, though it rests on a handful of rows.

open as a page