Why does mean impurity decrease rank a pure-noise, near-unique donor ID column third?
answer
- count the candidate cut-points
- maximum over many random candidates
- no held-out check anywhere
- many tiny deep splits accumulate
- compare against a depth-capped refit
basics
~20 sA near-unique column offers thousands of candidate cut-points, so the greedy splitter can almost always find one that lowers training impurity by chance. Those spurious splits accumulate credit because impurity importance is measured on the training rows only.
solid answer
~50 sImpurity importance rewards a feature for every split it wins, and it is scored on the training data with no held-out check. A donor ID with 5,000 near-unique values gives the splitter thousands of thresholds to maximise over at every node, so even pure noise yields a positive impurity drop somewhere; a two-level flag offers exactly one cut and can only be used once along a path. In a fully grown forest the ID keeps winning small deep splits, and although each is weighted down by node size, there are very many of them, so the total lands it in third place. The diagnosis is cheap: look at the cardinality profile of the top-ranked columns, refit with depth capped at four and watch whether the ID's rank collapses while genuine features hold, and check whether its position is stable across seeds. A noise column that only ranks high in the deep, unconstrained fit is an artefact of the measurement, not a finding.
go deeper
Remember the headline: a column with many distinct values gets many chances to look useful, and impurity importance never checks those splits against held-out data. Identifiers should not be model inputs in the first place.
Explain the mechanism, not just the symptom: the splitter maximises over thousands of candidate thresholds, and each spurious deep split adds its weighted impurity drop to the running total even though it fits noise.
Demonstrate the diagnosis. Profile cardinality next to the ranking, contrast a fully grown fit against a depth-capped one, check stability across seeds, and constrain the splitter with leaf minimums or binning rather than editing the score afterwards.
Decide the standing rule. Rankings that shift with tree depth should never reach a stakeholder deck unaccompanied, and identifier-shaped columns should be blocked at the feature-store boundary so no model or ranking is ever built on them.
## The two ingredients of the bias **1. More candidate cut-points means more chances to win.** A greedy splitter evaluates every admissible threshold for every feature and keeps the best-scoring one. A binary flag offers one candidate split. A continuous or near-unique column with 5,000 distinct values offers on the order of 5,000. Even when the column is independent of the target, the *maximum* impurity drop over thousands of random candidates is comfortably positive - this is a selection-of-maxima effect, not a property of the column. The splitter is comparing one honest candidate against thousands of lottery tickets. **2. The score is a training-fit statistic.** Nothing in mean decrease in impurity (MDI) is validated. A split that isolates eleven donors who happen to have given twice reduces training Gini exactly as legitimately as a split on genuine signal, and it collects credit. On a 5,000-row table a near-unique identifier can eventually isolate almost anything. Put together: the noise ID wins many splits it has no business winning, and every one of them pays. ## Why node weighting does not save you Each split's contribution is scaled by `n_node / N`, the share of training rows reaching the node, so the deep splits where the ID does its damage are individually tiny. The catch is volume. A fully grown tree on 5,000 rows contains hundreds of nodes, the great majority of them deep and small, and the ID is a plausible splitter at nearly all of them, whereas a genuinely decisive feature is consumed near the root and has little left to contribute below. Hundreds of small positive contributions add up to a respectable-looking total. This is also why a genuinely decisive binary flag can be *beaten* by a mediocre continuous feature with 2,000 cut-points. The flag delivers one large drop high in the tree and is then exhausted - below that split, its value is constant within each branch, so it can never be chosen again on that path. The continuous feature is reusable at every node with a fresh threshold, and its many small drops out-total the flag's single big one. MDI measures accumulated split usefulness, not decisiveness. ## The depth-4 contrast as a diagnostic Refit the same data with depth capped at four and re-read the ranking. A depth-4 tree can afford at most fifteen internal splits, and the greedy rule spends them on the highest-gain candidates available at large nodes - which is where real signal lives. A noise ID that ranked third in the fully grown forest usually falls sharply, while features carrying genuine signal keep their positions. That contrast is the tell: **a feature whose importance depends on the trees being allowed to grow deep is a feature earning its score from memorisation.** Two honest caveats. The depth-capped model is a different model, so you are comparing rankings, not proving anything about the original. And capping does not drive the bias to zero - a high-cardinality column can still win a shallow split by luck, particularly on small data. ## The rest of the diagnostic kit - **Profile cardinality alongside the ranking.** Print distinct-value counts next to the top ten scores. When the top of the list correlates with cardinality, treat the whole ranking as suspect. - **Check stability.** Refit with several seeds or across cross-validation folds. Genuine drivers hold their positions; artefacts wander. - **Sanity-check against the validation metric.** If the model's held-out performance is unchanged by whether the ID column is present, the column is contributing nothing real however high it ranks. - **Watch for leakage-shaped columns.** Row identifiers, timestamps stored as integers, and case numbers often correlate with the target through collection order, which produces the same symptom with a different cause and deserves a different fix. ## What to do about it The cheapest structural fix is not to feed near-unique identifiers to the model at all - an ID is not a feature, and its high rank is a warning that it was allowed in. Where a high-cardinality column is genuinely meaningful, constrain the splitter: cap depth, raise the minimum samples per leaf so tiny noise-carving nodes cannot form, and bin the values so the number of candidate cut-points shrinks. Each of these reduces the number of lottery tickets rather than trying to correct the score after the fact. What you must not do is take the ranking at face value and narrate a story about donor identifiers mattering. The number is real arithmetic over a real fitted model; it just does not mean what the reader assumes.
- Does capping depth at four remove the bias or merely shrink it?It shrinks it sharply and is useful as a diagnostic contrast, but it does not remove it. A high-cardinality column can still win a shallow split by chance, especially on small data, and the capped model is a different model - you have changed the fit, not corrected the statistic.
- Two near-duplicate credit-bureau utilisation fields each score mediocre. Is that the same bias?No, that is credit dilution. Once one of them wins a node, the other's remaining impurity drop there is near zero, so the pair splits the credit roughly in half across trees and both look unremarkable. Drop one and the survivor's score jumps. Cardinality bias inflates a weak feature; dilution deflates two strong ones.
- How would you spot this bias before you even look at the ranking?Profile the columns first. Any column whose distinct-value count approaches the row count is an identifier, not a feature, and should not be in the matrix. Timestamps stored as integers and sequential case numbers deserve the same suspicion, because they carry collection order rather than signal.
saying these in an interview costs you the question
- Concludes the ID column is genuinely predictive and keeps it
- Believes node-size weighting already cancels the cardinality bias
- Thinks the score would be the same on held-out rows
- Says only categorical features suffer this, not continuous ones
- Assumes a high-cardinality column always beats a binary flag