skip to content

What does one-hot encoding a 50-level US state column do to a linear model's fit?

level: middleimportance: should knowfreq 56%

answer

  1. count parameters, not just columns
  2. mostly zeros, one 1 per row
  3. each weight sees only that state's rows
  4. effective sample size per level
  5. penalty shrinks the thin levels

basics

~20 s

It adds fifty binary columns that are 98 percent zeros and gives every state its own weight. Each weight is driven only by that state's rows, so small states get unstable estimates and total variance rises. Penalising the weights becomes close to mandatory.

solid answer

~50 s

One-hot turns one column into fifty indicators, so the design matrix gains fifty parameters while the row count stays the same. Each state's coefficient is identified only by the rows from that state - the other 49 indicators are zero there - so a state with 30 rows gets a coefficient estimated from 30 rows, and its variance is large. That is the real cost: not the width of the matrix but the collapse of effective sample size per parameter, which shows up as an overfit gap between training and validation. The matrix is also extremely sparse - one 1 and 49 zeros per row - which sparse storage handles cheaply, so memory is rarely the binding constraint. The usual response is an L2 penalty, which shrinks poorly supported states toward the bulk, plus care with interactions, since crossing 50 states with 12 months gives 600 columns.

go deeper

for a junior

Know the mechanical fact first: k levels become k indicator columns, each row has a single 1, and the resulting block is almost entirely zeros.

for a middle

Be ready to explain why the extra parameters bite - a level's coefficient is identified only by that level's rows, so rare levels get noisy weights and validation error stops tracking training error.

for a senior

Show the diagnosis: check rows per level, compare training and validation curves with the column in and out, and reach for an L2 penalty so thin levels are shrunk rather than trusted.

for a principal

Frame it as a budget question across the whole feature set - how many parameters the row count can support, and which columns earn their width - rather than a per-column encoding preference.

## What actually lands in the matrix Before encoding, `state` is one column of strings. After one-hot encoding with 50 levels, it is 50 binary columns in which every row has exactly one 1 and 49 zeros. Roughly 98 percent of the entries in that block are zero. Two distinct consequences follow, and interviews reward separating them. ## Consequence 1: sparsity (a storage and compute fact) A mostly-zero matrix can be stored by listing only the non-zero positions, so the memory footprint tracks the number of non-zeros - one per row per encoded column - not the number of columns. Many linear-model solvers exploit this and do work proportional to the non-zeros. So the column blow-up is usually **not** a memory crisis, and a candidate whose whole answer is `it uses more memory` has answered the shallow half. What sparsity does cost is a set of second-order annoyances: standardising a sparse column by subtracting its mean destroys the sparsity, so scaling choices interact with the encoding; and any pairwise interaction you build multiplies column counts fast, since 50 states crossed with 12 months is 600 columns. ## Consequence 2: variance (the statistical fact that matters) This is the answer the question is really after. Adding 50 indicators adds 50 free parameters, and the number of rows has not changed. Worse, the rows are **partitioned** across those parameters: the coefficient for Wyoming is identified only by the Wyoming rows, because for every other row that indicator is zero and contributes nothing to its estimate. So the effective sample size per parameter is not `n`, it is `n_state`. If the data has 100,000 rows but Wyoming supplies 40 of them, that state's weight is a 40-row estimate. It will be noisy, it will chase whatever happened to those 40 rows, and it will not replicate on new data. Aggregate this over the long tail of small states and you get the classic symptom: training error keeps dropping as you add the encoded column, validation error does not follow. A second, subtler cost is that one-hot gives the model **no way to borrow strength** across levels. Every state is an independent parameter; nothing in the encoding tells the model that two neighbouring states behave alike. Whatever the levels have in common must be relearned separately for each. ## Why regularisation stops being optional An L2 penalty adds a term proportional to the sum of squared weights to the objective, which pulls every coefficient toward zero. For a well-populated state the data outvotes the penalty and the coefficient survives roughly intact; for a 40-row state the data barely resists and the coefficient is shrunk hard toward the bulk of the data. That is exactly the behaviour you want from a long tail of levels: it trades a little bias on small states for a large variance reduction, and it is why one-hot plus a penalty is a far more reliable combination than one-hot alone. One practical detail follows from this: whether you dropped a reference level changes what shrinking toward zero means. With every level kept and the intercept left out of the penalty, shrinkage pulls each state toward the overall level. With a reference level dropped, shrinkage pulls each remaining state toward the dropped state, which is an arbitrary target. ## How to tell whether it is hurting you Three cheap diagnostics: - **Level frequency table.** Sort the levels by row count. A long tail of levels with tens of rows is the warning sign; 50 roughly balanced levels over a million rows is not a problem at all. - **Learning curve behaviour.** Fit with and without the encoded column. If training error improves and validation error does not, the extra parameters are buying noise. - **Coefficient magnitudes against support.** If the largest absolute weights all belong to the rarest levels, the model is fitting small-sample noise. ## The scale at which this changes character Fifty levels is the regime where one-hot plus a penalty is still the right, boring answer. The picture changes qualitatively when a column has thousands of levels, where the number of parameters approaches or exceeds the number of rows and a different family of encodings takes over - a separate topic. The point to carry from 50 states is the mechanism: **columns are cheap, parameters supported by 40 rows are not.**

  • Does a gradient-boosted tree suffer in the same way from 50 indicator columns?
    It fails differently. Each split on a binary indicator can only separate one state from the other 49, so grouping several states together costs one split each and eats depth. With column sampling per split, any single indicator is also rarely offered as a candidate. The typical symptom is a model that under-uses the column rather than one that overfits it.
  • If sparse storage makes the matrix cheap, is the column blow-up harmless?
    No. Sparse storage solves memory and compute; it does nothing about statistics. The fifty free parameters are still there, still estimated from the rows of their own level, and still able to overfit a long tail of thinly populated states. Cheap to store is not the same as cheap to estimate.
  • Fifty states over a million balanced rows - is this still a problem?
    Barely. With roughly 20,000 rows behind each indicator, every coefficient is well identified and the variance argument mostly evaporates. The cost degrades to width and interaction blow-up. The danger lives in the shape of the frequency table, not in the level count on its own, so look at rows per level before worrying.

saying these in an interview costs you the question

  • Says one-hot is always safe because it assumes nothing
  • Believes extra columns cannot overfit while rows outnumber columns
  • Thinks sparse storage removes the variance problem
  • Assumes every level has enough rows to support its own weight
  • Reports only memory cost and misses the statistical cost

context