skip to content

Should a one-hot column drop a reference level before an L2-penalised linear fit?

level: seniorimportance: nice to knowfreq 30%

answer

  1. why the drop exists at all
  2. indicators sum to the intercept
  3. the penalty already picks one solution
  4. toward zero means toward the baseline
  5. keep the intercept out of the penalty

basics

~20 s

Usually no. Dropping a level exists to remove the exact redundancy between the full set of indicators and the intercept, which an L2 penalty already resolves on its own. Keeping every level and leaving the intercept unpenalised treats all levels symmetrically.

solid answer

~50 s

Dropping a level solves a specific problem: with an intercept present, the indicators for all `k` levels sum to 1 in every row, so the fit has no unique solution. An L2 penalty removes that ambiguity by itself, because among all the equivalent weight vectors it prefers the one with the smallest squared norm. So the reason to drop has already been handled. Keeping every level is also better behaved under a penalty: with a level dropped, shrinking a weight toward zero means shrinking it toward the dropped level, so predictions depend on which level you happened to make the baseline. Keep all `k` indicators, exclude the intercept from the penalty, and shrinkage pulls every level toward the common level instead of an arbitrary one. For trees, dropping is simply harmful - the dropped level is only expressible as all other indicators being zero, which no single split can say.

go deeper

for a junior

Know the mechanical fact: with an intercept present, the indicators for all levels add up to 1 in every row, which is why one level is often left out.

for a middle

Explain that the redundancy is what forces the drop and that an L2 penalty resolves it on its own, so keeping all levels is a legitimate choice rather than a mistake.

for a senior

Show you know what shrinkage points at - toward the average with all levels kept, toward the dropped level without - and that trees should never drop a level at all.

for a principal

Own the convention across the codebase: one documented rule per model family, encoder and model versioned together, and no silent divergence between the training and serving column layouts.

## What dropping a level is for With `k` indicator columns for a `k`-level categorical and an intercept column of ones, the indicators add up to the intercept in every single row. The columns are exactly linearly dependent, so infinitely many weight vectors produce identical predictions: add a constant `c` to all `k` level weights and subtract `c` from the intercept, and nothing about the fit changes. An unpenalised least-squares fit therefore has no unique answer. Dropping one level - making it the reference - breaks the dependence and restores a unique solution. (How the surviving coefficients are then read as contrasts against that reference is a regression-interpretation topic in its own right.) That is the entire motivation, and it is worth noticing what it is **not**: it is not a statement that redundant columns hurt predictive accuracy, and it is not a rule about encoding in general. ## Why a penalty changes the calculus An L2 penalty adds a term proportional to the sum of squared weights to the objective. Now the equivalent weight vectors are no longer equivalent - among all the shift-by-`c` variants that give identical predictions, the penalty strictly prefers the one with the smallest sum of squares, and that one is unique. The degeneracy that dropping a level was invented to fix has been fixed by the penalty. Keeping all `k` indicators is fitting fine. More importantly, keeping all `k` is **better** than dropping, for a reason about what shrinkage means: - **All levels kept, intercept unpenalised.** The penalty pulls each level weight toward zero, and because the intercept absorbs the common level of the target, zero means the overall average behaviour. A level with little data is shrunk toward the population, which is the sensible prior. - **One level dropped.** The dropped level has no weight to shrink; it is pinned at zero by construction while every other level's weight is pulled toward zero. Since the retained weights are now offsets relative to the baseline, shrinking them toward zero means shrinking them **toward the dropped level**. A level with little data is pulled toward whichever level you arbitrarily chose as the reference, not toward the average. Change the reference and the fitted predictions change - a dependence on an arbitrary choice that has nothing to do with the data. ## The unpenalised intercept detail This argument only works if the intercept is excluded from the penalty. Penalising the intercept makes the fit depend on where the target happens to be centred - add 100 to every target value and the solution changes in a way that is not about the data's structure. Leaving the intercept out is the standard convention for exactly this reason, and it is what lets the level weights be interpreted as deviations from a common level that shrinkage can safely pull toward zero. ## Where the L1 case differs The clean uniqueness argument is an L2 argument. With an L1 penalty the shift ambiguity is resolved differently and can be resolved non-uniquely, and the sparsity pattern - which levels get zeroed - can depend on the redundancy in ways that are awkward to reason about. Dropping a level is a more common and more defensible choice under L1. ## Trees and other non-linear learners Here dropping is not neutral, it is a loss. There is no intercept, no linear dependence, and therefore no problem to solve. What dropping does is make one level unrepresentable in a single split: a tree can test `is_state_texas == 1` in one node, but the dropped level can only be expressed as the conjunction `all other indicators are 0`, which requires as many nested splits as there are remaining levels. The dropped level becomes the hardest one for the model to isolate, for no benefit. Never drop for a tree ensemble. ## The practical summary - Unpenalised linear fit with an intercept: you must drop a level (or omit the intercept). - L2-penalised linear fit: keep every level, leave the intercept out of the penalty. - L1-penalised fit: dropping is the more common choice. - Trees and ensembles: keep every level, always. - Whatever you choose, the training pipeline and the serving pipeline must agree exactly - a mismatch in which level was dropped shifts every column and, as with any encoder mismatch, fails silently.

  • Should you drop a level before a random forest or gradient boosting?
    No. There is no intercept and no linear dependence, so there is nothing to fix, and dropping makes that level the only one a single split cannot isolate - it is expressible solely as every other indicator being zero. You have made one level structurally harder for the model to use and gained nothing in return.
  • Why should the intercept be left out of the penalty?
    Because a penalised intercept ties the fit to where the target happens to be centred - shift every target value by a constant and the solution changes for no substantive reason. Keeping it unpenalised lets it absorb the overall level, which is what makes shrinking the level weights toward zero mean shrinking toward the average rather than toward an arbitrary point.
  • If you do drop a level under a penalty, which level should it be?
    Since shrinkage now pulls the other levels toward the dropped one, choose a large, stable, well-populated level so the implicit target of that shrinkage is estimated from plenty of data. A rare level as reference means every other coefficient is being pulled toward a noisy anchor, and the same rare level may vanish from a future refit.

saying these in an interview costs you the question

  • Says dropping a level is required for every model and every fit
  • Penalises the intercept alongside the level weights
  • Assumes the choice of dropped level never affects predictions
  • Drops a level before a tree ensemble to avoid a redundancy that does not exist
  • Keeps all levels in an unpenalised fit with an intercept and expects a unique solution

context