Should one-hot dummies and rare binary flags be standardised before a penalised fit?
answer
- binary columns have a fixed, tiny spread
- a 2% flag has sd 0.14
- dividing by it makes the ones huge
- least evidence, least shrinkage
basics
~10 sThere is no universal answer. Dividing a rare 0/1 flag by its tiny standard deviation stretches its range and lets it escape most of the shrinkage, though it rests on a handful of rows.
solid answer
~50 sThe tension is that the penalty compares coefficients, so how a binary column is scaled decides how hard it is shrunk. A 2%-prevalence flag has a standard deviation near 0.14; dividing by that turns its rare 1s into values around 7 and multiplies the column's sum of squares roughly fiftyfold, so the coefficient is shrunk far less than if the flag had stayed 0/1 — the opposite of what you want from an effect estimated on 2% of the rows. Leaving dummies raw next to standardised continuous features has the mirror cost: a 0/1 column has a spread of at most 0.5, so dummies are shrunk harder than everything else. My default is to standardise the continuous predictors, leave dummies at 0/1, and accept that rare levels get shrunk hard — usually the right prior. Then state the choice explicitly and check it on the validation curve.
go deeper
Know the two facts this rests on: a 0/1 column has a small fixed spread set by its prevalence, and a penalty shrinks narrow-spread columns hardest.
Be able to compute it. A 2% flag has a standard deviation near 0.14, so dividing by it multiplies that column's sum of squares about fiftyfold and largely removes the shrinkage it was getting.
Show judgment from experience: a rare flag scaled up, kept by the penalty on a dozen positive rows, unstable across refits and useless in production. Say how you would test the choice on validation rather than assert it.
Own the convention. One documented rule for how binaries and continuous features are scaled before penalised fits keeps coefficients comparable across models and makes a lambda grid mean the same thing on every project.
## Why this is a real question and not a detail Once you accept that a penalty compares coefficients and that a coefficient's size is set by its column's spread, the treatment of binary columns stops being cosmetic. Standardising a continuous predictor is uncontroversial: its units were arbitrary. A 0/1 indicator's units are *not* arbitrary — a 1 means the category is present, and the coefficient reads directly as the shift in the target when it is. Dividing that column by anything is a deliberate decision about how much regularisation that indicator should receive, and it deserves a stated rationale. ## The arithmetic for a rare flag A binary column with prevalence `p` has standard deviation `sqrt(p * (1 - p))`. For `p = 0.02` that is about 0.14, and the centred sum of squares over `n` rows is `n * p * (1 - p) = 0.0196 * n`. Standardise it and the sum of squares becomes `n` — about fifty times larger. Concretely, the 2% of rows carrying a 1 now hold the value `(1 - 0.02) / 0.14`, roughly 7, while the other 98% sit at about -0.14. Recall that ridge keeps a fraction `S / (S + lambda)` of the least-squares coefficient for an uncorrelated column: multiplying `S` by fifty moves that column from heavily shrunk to essentially unregularised. Under lasso, the same inflation makes the column much more likely to survive selection. So standardising a rare dummy hands the *least* well-evidenced coefficient in the model — one estimated from perhaps twenty positive rows out of a thousand — the *most* freedom from the penalty. That is a strange prior to adopt by accident, and it is precisely the accident that reflexive "standardise everything" causes. The failure looks familiar in production: a rare flag with a big, confident coefficient, an offline metric that liked it, and no stability at all when the model is refit next month. ## The mirror cost of leaving them raw The opposite default is not free either. A 0/1 column has spread `sqrt(p(1-p))`, at most 0.5 (at `p = 0.5`) and smaller for anything imbalanced, whereas standardised continuous predictors all have spread 1. Dummies are therefore shrunk harder than the continuous features sitting next to them, and a balanced, genuinely predictive category can be shrunk out of a lasso for no reason except that indicators are narrow by construction. If a categorical is your main signal, this is worth knowing before you conclude the categorical does not matter. ## Interpretation, which quietly argues for raw A coefficient on a standardised binary is an effect per standard deviation of that binary, and a one-standard-deviation move in a 2% flag is about a seventh of the way from 0 to 1 — a change that never occurs. The only move that happens in the data is the full 0-to-1 jump, worth roughly seven standard deviations. Coefficients on raw 0/1 columns read directly and honestly; coefficients on standardised ones need mental arithmetic to mean anything to a stakeholder. ## Practical positions people take - **Standardise continuous predictors, leave dummies at 0/1.** The most common default. Accepts stronger shrinkage on rare levels, which is usually the prior you want, and keeps coefficients readable. - **Standardise everything uniformly.** Defensible when interpretability does not matter and the categories are reasonably balanced; dangerous with rare levels. - **Divide continuous predictors by two standard deviations instead of one.** A published suggestion aimed at making a continuous coefficient roughly comparable to the 0-to-1 switch of a balanced binary. It is a convention for comparability, not a fact about regularisation. - **Pool rare levels first.** Often the better fix. If a level appears in 20 rows, collapsing it into an "other" bucket addresses the real problem — thin evidence — rather than arguing about its standard deviation. ## Treat the dummies of one categorical as a block Whatever you choose, apply it uniformly across the indicators derived from a single categorical. Scaling each dummy by its own standard deviation gives rare levels a wider numeric range than common ones, so the penalty ends up preferring whichever level happens to be rare — an artefact with no substantive meaning. If you actually need the whole categorical to enter or leave the model together, a group penalty over the block is the honest tool, rather than per-column scaling tricks. ## How to settle it This is a modelling choice, so decide it the way you decide any other: fix the alternatives, tune lambda separately under each, and compare on the same validation protocol, paying attention to the stability of the rare-flag coefficients across folds as well as to the headline metric. Then write the decision down, because a future reader looking at the coefficient table cannot tell from the numbers alone which convention produced them.
- What does a per-standard-deviation coefficient actually mean for a 0/1 flag?Very little as a description of anything realisable. For a 2% flag, one standard deviation is about a seventh of the way from 0 to 1, a change that never occurs in the data; the only move that happens is the full 0-to-1 jump, worth about seven standard deviations. Coefficients on raw 0/1 columns read directly as the shift in the target when the flag turns on.
- How should the several dummies derived from one categorical be handled?Consistently, as a block — scale them all the same way or not at all. Scaling each by its own standard deviation gives rare levels a wider numeric range than common ones, so the penalty starts preferring levels purely for being rare. If the categorical should enter or leave the model as a unit, a group penalty over the block is the honest tool.
- Is there a scaling convention that makes continuous predictors comparable with 0/1 binaries?One published suggestion is to divide continuous predictors by two standard deviations rather than one, so a one-unit move in the rescaled variable is roughly comparable to the 0-to-1 switch of a balanced binary. It makes coefficient magnitudes easier to compare, but it is a convention for readability — the penalty still treats a rare flag and a dense continuous feature very differently.
saying these in an interview costs you the question
- Standardises everything reflexively and never inspects rare columns
- Thinks 0/1 columns are already standardised
- Believes scaling choices cannot change which features lasso keeps
- Scales each dummy of one categorical by its own standard deviation
- Trusts a rare flag's large coefficient without checking fold stability