skip to content

High-Cardinality Categories

Target and count encoding computed out of fold with smoothing towards the prior, the hashing trick, and folding rare levels into one bucket. Interviewers ask why naive mean encoding leaks.

on this pageshow

questions

5

Why does target-encoding a 40,000-level seller ID leak, and how does out-of-fold encoding fix it?

level: middleimportance: must knowfreq 72%

answer

  1. whose labels built this seller's mean?
  2. optimism that evaporates on holdout
  3. worst when a level has few rows
  4. means from the other folds only
  5. full-train means for validation and test

basics

~20 s

Computing each seller's target mean over all training rows puts a row's own label inside its own feature, so the model reads the answer back. Out-of-fold encoding builds each fold's means from the other folds only.

solid answer

~50 s

Target encoding replaces a category level with the average target among rows of that level. Computed naively over the whole training set, the mean for a seller with three sales contains one third of the row being encoded, and for a seller with one sale it *is* that row's label. The model happily splits on that column and cross-validation looks excellent, but at scoring time a new row's label is not in the mean, so the holdout collapses. The fix is out-of-fold encoding: split the training rows into K folds, and for the rows in fold k compute every level's mean from the other K-1 folds only. Validation and test rows get means computed from the whole training set, which is safe because their labels never entered it. Pair it with smoothing so tiny sellers are shrunk toward the global rate rather than trusted at face value.

go deeper

for a junior

Know what target encoding is: swap the category label for the average target of rows with that label. Remember the headline risk — the row's own outcome ends up inside its own feature.

for a middle

Be ready to write the mean, show that a row's share of it is one over the level count, and describe the fold construction that removes it. Explain why validation rows are encoded differently from training rows.

for a senior

Show how you would diagnose it in a real pipeline: the cross-validation to holdout gap, one encoded feature dominating, encoded values pinned at 0 and 1. Then describe nesting the encoding inside the outer split and versioning the mapping for serving.

for a principal

Own the call on whether a label-derived feature belongs in the stack at all, given that every consumer must reproduce the fold discipline and the refresh cadence. Weigh the lift against a simpler unsupervised encoding that no one can misuse.

## What target encoding is A high-cardinality categorical column — 40,000 seller IDs on a marketplace, 30,000 postal codes, millions of device fingerprints — cannot usefully become 40,000 separate columns. **Target encoding** (also called mean encoding or likelihood encoding) replaces each level with a single number: the average of the target among the rows carrying that level. Seller `S91` with 200 historical orders, 18 of which converted, becomes `0.09`. One dense, informative column instead of tens of thousands of near-empty ones. It is powerful precisely because it is a supervised summary: the encoded value already contains the thing the model is trying to predict. That is also exactly why it is dangerous. ## The leak Write the mean for level `c` as `mean_c = (sum of y over rows with level c) / n_c`. If you compute it over the entire training set and then use it as a feature on those same training rows, the label of row `i` is one of the terms in the numerator of row `i`'s own feature. Its share of the feature is `1/n_c`: - `n_c = 1` (a seller with a single historical row): the encoded value equals that row's label exactly. The feature *is* the answer. - `n_c = 3`: a third of the feature is the row's own answer. - `n_c = 500`: the contamination is negligible. With 40,000 sellers the count distribution is long-tailed — a handful of large sellers and a huge mass of sellers with a handful of rows each. So the contaminated rows are not an edge case; they are most of the table. A tree finds the encoded column instantly, cross-validation reports a striking lift, and the model is worthless on new traffic where the incoming row's label is, of course, absent from the mapping. The tell-tale symptoms: one encoded column dominating the model, an encoded feature with many values sitting exactly at 0.0 or 1.0, and a large unexplained gap between cross-validated score and a clean holdout. ## Out-of-fold encoding Split the *training* rows into K folds (5 or 10 is typical). To produce the encoded feature for the rows in fold k, compute every level's target mean using only the rows in the other K-1 folds. Repeat for every fold, and the training column is assembled fold by fold. No row's own label ever contributes to its own value, and neither do the labels of rows the model will be evaluated on within that fold. Rows outside the training set — the validation fold of an outer split, the test set, live traffic — are encoded with means computed from the *entire* training set. That is legitimate: those rows' labels were never in the training data, so nothing about them is being read back. Two mechanics matter in practice: - **Nesting.** If an outer cross-validation loop is scoring the model, the encoding folds must live strictly inside each outer training split, recomputed per outer fold. Building one encoding over all data and then splitting undoes everything. - **Fold noise.** Training rows are encoded from K-1 folds while inference rows are encoded from K folds, so the training column is slightly noisier and slightly differently distributed. That mismatch is usually a feature, not a bug — it acts like regularisation and stops the model over-trusting the column — but if it bites, averaging several repeats with different fold seeds (bagged encodings) reduces the variance. ## What out-of-fold does not fix Out-of-fold removes the *self-label* leak. It does nothing about the fact that a mean computed from three rows is a terrible estimate of a seller's true rate; it will happily hand you 0.0 or 1.0 from three out-of-fold rows. That is a variance problem, solved by shrinking the level mean toward the global rate (smoothing) or by collapsing tiny levels into a shared bucket. A related trap is **leave-one-out encoding**, which encodes row `i` with the level mean over all rows of that level *except* row `i`. It looks like the honest version, and it leaks worse than it appears: given the level's count and sum, the leave-one-out value is a strictly decreasing function of the row's own label, so a sufficiently flexible model can invert it and recover `y_i` from the feature. Cross-validation inflates again. Adding noise to the encoded value is the usual patch, and it is a patch — fold-based encoding is the cleaner construction. ## Serving Whatever scheme you pick, the artefact that goes to production is a mapping from level to number, built from training data only, versioned alongside the model. A level not present in the mapping falls back to the global prior. Refreshing the mapping without retraining the model, or refreshing it from data that includes labels the model has never seen, quietly reintroduces skew between training and serving.

  • How do you encode the validation and test rows once the training column is built out-of-fold?
    With level means computed from the entire training set. Those rows' labels never entered the training data, so nothing leaks back; using the full training set also gives the most stable estimate per level. Levels absent from the mapping fall back to the global target mean.
  • Leave-one-out encoding excludes the row's own label — why does cross-validation still inflate?
    Because the excluded label is recoverable. Given the level's count and total, the leave-one-out value moves down by a fixed amount when the row's own label is 1, so the encoded feature is a deterministic decreasing function of that label. A tree can invert it and effectively read the target.
  • Your outer cross-validation still shows a gap to the holdout after switching to out-of-fold encoding — what do you check?
    Check that the encoding is recomputed inside each outer training split rather than once over all data, that rows sharing a level are not split across folds when they should be grouped, and that tiny levels are smoothed. An unsmoothed out-of-fold mean from two rows is still mostly noise the model can memorise.

Grading an exam where each student's mark is the class average including their own paper. In a class of 500 nobody notices; in a class of one, the average is the student's own answer.

saying these in an interview costs you the question

  • Says it cannot leak because no test labels were used
  • Computes level means once over all training rows
  • Treats a high cross-validation score as proof it works
  • Claims leave-one-out encoding removes the leakage entirely
  • Builds the encoding before the outer split and reuses it

context

open as a page

What does count encoding do to a 30,000-level ZIP code column, and when does the count itself carry signal?

level: juniorimportance: should knowfreq 41%

basics

~20 s

Count encoding replaces each ZIP code with how many rows carry it, turning 30,000 levels into one numeric column. It helps when frequency proxies something real, like population density. It uses no labels, so it cannot leak.

open as a page

In target encoding, a seller with three historical rows all converted — smooth toward the prior or bucket it as rare?

level: seniorimportance: should knowfreq 46%

basics

~10 s

Three rows cannot support a rate of 1.0. Smoothing shrinks the estimate toward the global rate, weighted by the row count, so some seller signal survives; a rare bucket discards it entirely. Prefer smoothing.

open as a page

How do you choose an encoding for a 40,000-level seller ID that must be refreshed and served daily?

level: principalimportance: should knowfreq 37%

basics

~20 s

Decide on four axes: how much signal the identity carries, what the model family can consume, what state serving can refresh, and who must explain the feature. Start cheap and escalate only when a holdout says it pays.

open as a page

How does the hashing trick encode millions of ad publisher domains, and what do collisions cost?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

A hash function maps each domain to a bucket index modulo a fixed bucket count, say 2^20, and that bucket is the feature slot. Nothing is stored, so new domains need no special case. Colliding domains share one blended weight.

open as a page