skip to content

Should one-hot dummies and rare binary flags be standardised before a penalised fit?

level: seniorimportance: nice to knowfreq 24%

answer

  1. binary columns have a fixed, tiny spread
  2. a 2% flag has sd 0.14
  3. dividing by it makes the ones huge
  4. least evidence, least shrinkage

basics

~10 s

There is no universal answer. Dividing a rare 0/1 flag by its tiny standard deviation stretches its range and lets it escape most of the shrinkage, though it rests on a handful of rows.

solid answer

~50 s

The tension is that the penalty compares coefficients, so how a binary column is scaled decides how hard it is shrunk. A 2%-prevalence flag has a standard deviation near 0.14; dividing by that turns its rare 1s into values around 7 and multiplies the column's sum of squares roughly fiftyfold, so the coefficient is shrunk far less than if the flag had stayed 0/1 — the opposite of what you want from an effect estimated on 2% of the rows. Leaving dummies raw next to standardised continuous features has the mirror cost: a 0/1 column has a spread of at most 0.5, so dummies are shrunk harder than everything else. My default is to standardise the continuous predictors, leave dummies at 0/1, and accept that rare levels get shrunk hard — usually the right prior. Then state the choice explicitly and check it on the validation curve.

go deeper

for a junior

Know the two facts this rests on: a 0/1 column has a small fixed spread set by its prevalence, and a penalty shrinks narrow-spread columns hardest.

for a middle

Be able to compute it. A 2% flag has a standard deviation near 0.14, so dividing by it multiplies that column's sum of squares about fiftyfold and largely removes the shrinkage it was getting.

for a senior

Show judgment from experience: a rare flag scaled up, kept by the penalty on a dozen positive rows, unstable across refits and useless in production. Say how you would test the choice on validation rather than assert it.

for a principal

Own the convention. One documented rule for how binaries and continuous features are scaled before penalised fits keeps coefficients comparable across models and makes a lambda grid mean the same thing on every project.

## Why this is a real question and not a detail Once you accept that a penalty compares coefficients and that a coefficient's size is set by its column's spread, the treatment of binary columns stops being cosmetic. Standardising a continuous predictor is uncontroversial: its units were arbitrary. A 0/1 indicator's units are *not* arbitrary — a 1 means the category is present, and the coefficient reads directly as the shift in the target when it is. Dividing that column by anything is a deliberate decision about how much regularisation that indicator should receive, and it deserves a stated rationale. ## The arithmetic for a rare flag A binary column with prevalence `p` has standard deviation `sqrt(p * (1 - p))`. For `p = 0.02` that is about 0.14, and the centred sum of squares over `n` rows is `n * p * (1 - p) = 0.0196 * n`. Standardise it and the sum of squares becomes `n` — about fifty times larger. Concretely, the 2% of rows carrying a 1 now hold the value `(1 - 0.02) / 0.14`, roughly 7, while the other 98% sit at about -0.14. Recall that ridge keeps a fraction `S / (S + lambda)` of the least-squares coefficient for an uncorrelated column: multiplying `S` by fifty moves that column from heavily shrunk to essentially unregularised. Under lasso, the same inflation makes the column much more likely to survive selection. So standardising a rare dummy hands the *least* well-evidenced coefficient in the model — one estimated from perhaps twenty positive rows out of a thousand — the *most* freedom from the penalty. That is a strange prior to adopt by accident, and it is precisely the accident that reflexive "standardise everything" causes. The failure looks familiar in production: a rare flag with a big, confident coefficient, an offline metric that liked it, and no stability at all when the model is refit next month. ## The mirror cost of leaving them raw The opposite default is not free either. A 0/1 column has spread `sqrt(p(1-p))`, at most 0.5 (at `p = 0.5`) and smaller for anything imbalanced, whereas standardised continuous predictors all have spread 1. Dummies are therefore shrunk harder than the continuous features sitting next to them, and a balanced, genuinely predictive category can be shrunk out of a lasso for no reason except that indicators are narrow by construction. If a categorical is your main signal, this is worth knowing before you conclude the categorical does not matter. ## Interpretation, which quietly argues for raw A coefficient on a standardised binary is an effect per standard deviation of that binary, and a one-standard-deviation move in a 2% flag is about a seventh of the way from 0 to 1 — a change that never occurs. The only move that happens in the data is the full 0-to-1 jump, worth roughly seven standard deviations. Coefficients on raw 0/1 columns read directly and honestly; coefficients on standardised ones need mental arithmetic to mean anything to a stakeholder. ## Practical positions people take - **Standardise continuous predictors, leave dummies at 0/1.** The most common default. Accepts stronger shrinkage on rare levels, which is usually the prior you want, and keeps coefficients readable. - **Standardise everything uniformly.** Defensible when interpretability does not matter and the categories are reasonably balanced; dangerous with rare levels. - **Divide continuous predictors by two standard deviations instead of one.** A published suggestion aimed at making a continuous coefficient roughly comparable to the 0-to-1 switch of a balanced binary. It is a convention for comparability, not a fact about regularisation. - **Pool rare levels first.** Often the better fix. If a level appears in 20 rows, collapsing it into an "other" bucket addresses the real problem — thin evidence — rather than arguing about its standard deviation. ## Treat the dummies of one categorical as a block Whatever you choose, apply it uniformly across the indicators derived from a single categorical. Scaling each dummy by its own standard deviation gives rare levels a wider numeric range than common ones, so the penalty ends up preferring whichever level happens to be rare — an artefact with no substantive meaning. If you actually need the whole categorical to enter or leave the model together, a group penalty over the block is the honest tool, rather than per-column scaling tricks. ## How to settle it This is a modelling choice, so decide it the way you decide any other: fix the alternatives, tune lambda separately under each, and compare on the same validation protocol, paying attention to the stability of the rare-flag coefficients across folds as well as to the headline metric. Then write the decision down, because a future reader looking at the coefficient table cannot tell from the numbers alone which convention produced them.

  • What does a per-standard-deviation coefficient actually mean for a 0/1 flag?
    Very little as a description of anything realisable. For a 2% flag, one standard deviation is about a seventh of the way from 0 to 1, a change that never occurs in the data; the only move that happens is the full 0-to-1 jump, worth about seven standard deviations. Coefficients on raw 0/1 columns read directly as the shift in the target when the flag turns on.
  • How should the several dummies derived from one categorical be handled?
    Consistently, as a block — scale them all the same way or not at all. Scaling each by its own standard deviation gives rare levels a wider numeric range than common ones, so the penalty starts preferring levels purely for being rare. If the categorical should enter or leave the model as a unit, a group penalty over the block is the honest tool.
  • Is there a scaling convention that makes continuous predictors comparable with 0/1 binaries?
    One published suggestion is to divide continuous predictors by two standard deviations rather than one, so a one-unit move in the rescaled variable is roughly comparable to the 0-to-1 switch of a balanced binary. It makes coefficient magnitudes easier to compare, but it is a convention for readability — the penalty still treats a rare flag and a dense continuous feature very differently.

saying these in an interview costs you the question

  • Standardises everything reflexively and never inspects rare columns
  • Thinks 0/1 columns are already standardised
  • Believes scaling choices cannot change which features lasso keeps
  • Scales each dummy of one categorical by its own standard deviation
  • Trusts a rare flag's large coefficient without checking fold stability

context