In target encoding, a seller with three historical rows all converted — smooth toward the prior or bucket it as rare?
answer
- three rows cannot justify a rate of 1.0
- blend the level mean with the global rate
- a pseudo-count sets how much evidence is needed
- the hard version is one shared bucket
- keep the row count as its own feature
basics
~10 sThree rows cannot support a rate of 1.0. Smoothing shrinks the estimate toward the global rate, weighted by the row count, so some seller signal survives; a rare bucket discards it entirely. Prefer smoothing.
solid answer
~50 sBoth attack the same problem — a level mean from almost no data is mostly noise — and differ in how much they throw away. Smoothing shrinks the level mean toward the global rate: `enc = (n * mean_level + m * prior) / (n + m)`, where `m` is a pseudo-count controlling how much evidence a level needs before it is trusted. With `m = 10`, three converted rows land near the prior rather than at 1.0, while a seller with 2,000 rows is barely moved. A rare bucket instead replaces every level below a count threshold with one shared level — a hard version of the same shrinkage. I default to smoothing because it is continuous and needs no arbitrary threshold, and add a rare bucket only when the mapping would otherwise hold hundreds of thousands of single-row levels.
code
python · 14 linesrows = [("A", 1), ("A", 0), ("A", 1), ("A", 1), ("B", 1), ("B", 0), ("C", 0)]
prior = sum(y for _, y in rows) / len(rows)
m = 10.0 # pseudo-count: evidence a level needs before its own mean dominates
stats = {}
for cat, y in rows:
n, s = stats.get(cat, (0, 0))
stats[cat] = (n + 1, s + y)
for cat, (n, s) in sorted(stats.items()):
raw = s / n
smoothed = (n * raw + m * prior) / (n + m)
print("%s n=%d raw=%.3f smoothed=%.3f" % (cat, n, raw, smoothed))
print("prior=%.3f" % prior)go deeper
Remember that a category seen three times gives an unreliable average, and that the fix is to pull it toward the overall rate. Recognise the blended formula when you see it.
Explain the pseudo-count: what happens as it grows, where the halfway point sits, and what the encoded column looks like before and after. Contrast it with a hard rare bucket in one sentence.
Show judgment about which to apply where — the storage and refresh cost of a huge mapping, the traffic share sitting in the tail, keeping the count as a companion feature, and inspecting the encoded distribution for mass at 0 and 1.
Own the tradeoff between a sharp feature and a stable one across retrains. Decide how much tail resolution the business actually needs and whether the extra hyperparameter and its tuning cost are worth the lift on the metric that matters.
## The problem both mechanisms solve Out-of-fold construction stops a row's own label from entering its own feature, but it does not make a mean computed from three rows trustworthy. Three conversions out of three gives `mean = 1.0`; the honest reading is "this seller might be anywhere from average to excellent, and we have almost no evidence." Handed to a tree, that 1.0 sits next to the 1.0 of a seller with 2,000 rows and a genuinely perfect record, and the model treats them as the same. Across 40,000 sellers, most of which have a handful of rows, this noise floods the column. ## Smoothing toward the prior The standard remedy is a weighted blend between the level's own mean and the global target mean: ``` enc_c = (n_c * mean_c + m * prior) / (n_c + m) ``` - `n_c` is the level's row count in the encoding data, `mean_c` its target mean, `prior` the overall target mean. - `m` is a **pseudo-count**: literally, the number of imaginary rows at the global rate that you add to every level. It is the amount of evidence a level must accumulate before its own mean starts to dominate. At `n_c = m` the encoded value sits exactly halfway between the level mean and the prior. As `n_c` grows the prior term becomes irrelevant; as `n_c` shrinks toward zero the encoding converges on the prior. This is empirical-Bayes shrinkage in its simplest form, and it is self-adjusting — small levels are pulled hard, large levels are left alone, with no threshold anywhere. `m` is a hyperparameter and should be tuned like one, on the same validation scheme as everything else. Small `m` gives a sharper but noisier feature; large `m` collapses the column toward a constant. A variant replaces the hard pseudo-count with a smooth weighting function of `n_c`, but the behaviour is the same and the tuned pseudo-count is easier to explain. ## Collapsing rare levels into one bucket The alternative is a hard cut: every level with fewer than, say, 100 rows becomes the single level `rare`, and that bucket gets its own encoded value — the target mean across all the rows it swallowed. On an app-store category column, dozens of niche categories with a handful of listings each collapse into one bucket with a stable, well-estimated rate. What you gain: a much smaller mapping to store and serve, a stable estimate, and a natural, already-modelled destination for a level the serving path has never encountered. What you lose: every distinction between rare levels, including a genuinely bad actor with 90 rows that would have survived smoothing. And the threshold is arbitrary — 100 is a habit, not a derivation; it should be checked against the count distribution and, if it matters, tuned. ## Choosing, and combining They are not exclusive, and in a real pipeline they usually stack: 1. Collapse levels below a small floor (levels with one or two rows, or the tail that accounts for a negligible share of traffic) into `rare`, mostly to bound the size and refresh cost of the mapping. 2. Smooth everything that survives, so the mid-tail levels degrade gracefully instead of falling off a cliff at the threshold. A useful third move is to keep the level's row count as its own feature alongside the encoded mean. The model can then learn "trust this rate more when the count is high", which recovers some of the information smoothing deliberately blurs — and the count is often independently predictive. ## How to tell it is working Look at the distribution of the encoded column. Spikes of mass exactly at 0.0 and 1.0 mean tiny levels are being trusted at face value. After smoothing, the histogram should be a hump around the prior with tails reaching out only where the counts justify it. Then check the score: if increasing `m` keeps improving validation, the column was noise-dominated; if performance falls off immediately, the level signal is strong and real and you can afford a lighter hand. ## The mistake to avoid Do not use smoothing as a substitute for out-of-fold construction. They fix different failures: shrinkage attacks the *variance* of an estimate from few rows, folds attack the *bias* of reading a row's own label. Heavy smoothing does mask the self-label leak on large levels — with `m = 100` the single-row seller no longer encodes to its own label — but it hides the mechanism rather than removing it, and the smallest levels are precisely where the contamination is largest. Use both.
- How would you pick the pseudo-count rather than defaulting to 10?Tune it on the same validation scheme as the model's other hyperparameters, over a coarse log-spaced grid. The count distribution gives a sensible starting range: pick a value near the count at which you would personally start believing a level's rate, then let validation move it.
- What threshold do you use for the rare bucket, and how do you defend it?Set it from the count distribution and the traffic share, not from habit. A defensible rule is the count below which a level's mean is indistinguishable from the prior at your noise level, or the cut that keeps the mapping within its storage and refresh budget while the bucket stays a small share of rows.
- Does smoothing remove the need for out-of-fold encoding?No. Smoothing reduces the variance of an estimate built from few rows; folds remove a row's own label from its own feature. Heavy smoothing dilutes the leak but does not eliminate it, and it bites hardest exactly on the smallest levels. Use both together.
A restaurant with one five-star review is not better than one with 4.6 stars from 800. Review sites solve it by pulling small-sample ratings toward the site-wide average until enough votes accumulate.
saying these in an interview costs you the question
- Trusts a level mean computed from two or three rows
- Treats smoothing as a replacement for out-of-fold folds
- Picks the rare threshold with no reference to the counts
- Thinks the pseudo-count needs no tuning
- Drops rare levels entirely instead of bucketing them