Agent ID tops a routing model's training permutation importance but is near zero on held-out data — why?
answer
- two data splits, two different questions
- thousands of levels, few rows each
- memorised, not learned
- held-out version is what ships
- judge near-zero against the spread
basics
~20 sThe model memorised agent ID on the rows it was fitted to. Shuffling it on training data destroys that memorised fit, so the score collapses; on unseen rows the memorisation was never worth anything, so scrambling it costs nothing.
solid answer
~40 sPermutation importance is computed against whatever data you hand it, and the two answers mean different things. On the training rows, a high-cardinality column such as agent ID lets a flexible model carve out tiny groups and fit them almost individually, so scrambling it wrecks a fit that depended on it — large drop. On held-out rows, those memorised per-agent patterns do not transfer, the model's generalising performance never rested on the column, and shuffling it changes almost nothing — near-zero drop. The gap between the two is a direct read-out of overfitting to that feature. For the question 'what does the deployed model actually rely on', use held-out data with the production metric. Keep the training version as a diagnostic, and be explicit about which one any published chart shows.
go deeper
Remember that permutation importance depends on which rows you score, and that training and held-out versions can disagree sharply. Say which data you used whenever you quote a number.
Explain how a high-cardinality column lets a flexible model fit tiny groups almost individually, and why that memorised structure collapses on training shuffles but is worth nothing on unseen rows.
Demonstrate the diagnostic move: use the gap to localise overfitting to a specific column, then act on it through encoding or capacity limits, and defend which version you publish for the production metric.
Own the convention across the team: which split and metric importance is computed on, how it is labelled, and what a train/held-out gap obliges an owner to do before a model passes review.
## Two numbers, two questions Nothing in the permutation procedure specifies which rows you score on. Run it on the training rows and you measure **what the fitted model leans on to reproduce the data it saw**. Run it on held-out rows and you measure **what the model leans on to perform on data it has not seen**. Those are different quantities, and a large gap between them is informative in its own right. ## Why a high-cardinality identifier produces this gap A support-ticket routing model given agent ID as a categorical feature sees thousands of levels, many with only a handful of training tickets each. A flexible learner can isolate those levels and fit their labels closely — effectively storing a per-agent answer rather than learning a rule. That memorised structure is genuinely load-bearing *on the training rows*: shuffle the identifier and the stored answers are attached to the wrong tickets, so training performance falls sharply. Permutation importance on the training set therefore ranks agent ID first, and does so honestly. On held-out tickets the picture inverts. The per-agent answers were fitted to a handful of rows each and carry little that transfers; some held-out agents may not even appear in training. The model's held-out performance is therefore being driven by the features that generalise — ticket text signals, queue, priority, time of day — and agent ID contributes close to nothing. Shuffle it and the held-out score moves by an amount inside the shuffle noise. ## What the gap tells you A feature that is large on train and near zero on held-out is a feature the model **memorised rather than learned from**. That is a concrete, actionable overfitting signal attached to a specific column, which is more useful than a single global train/test score gap. Typical responses: - Reduce the model's ability to isolate rare levels: fewer, deeper-constrained splits, stronger regularisation, a minimum count per leaf. - Re-encode the column so it carries a generalising summary rather than an identity — for example, aggregate historical statistics per agent computed on data strictly before each ticket, with the identity itself dropped. - Remove the column and check whether held-out performance moves at all. If it does not, you have simplified the model for free. The reverse pattern — small on train, large on held-out — is rare and usually points at a distribution difference between the splits rather than at a property of the feature. ## Which one to publish For explanation and stakeholder reporting, publish the **held-out** version with the metric the model is judged on in production, because that is what the deployed system's performance actually rests on. Keep the training version in the diagnostic toolkit. The unforgivable version is publishing one without saying which it is, since the two can rank the same feature first and last. ## Practical cautions on the held-out side - **Noise.** A small validation set gives wide error bars, and 'near zero' has to be judged against those bars, not against zero exactly. Repeat the whole procedure across cross-validation folds and pool the results when the evaluation set is small. - **Split quality.** If the split is not honest — rows from the same ticket thread, the same customer, or overlapping time windows on both sides — held-out permutation importance inherits that dishonesty and can look like training importance. Fix the split first; an importance chart cannot be more trustworthy than the split it is computed on. - **Distribution shift.** If the held-out period differs systematically from training, a low importance may reflect the feature's behaviour in that period rather than a stable property. ## The interview answer Say the mechanism (memorisation of a high-cardinality column), say what the gap diagnoses (overfitting localised to one feature), say which number you would publish and why (held-out, production metric), and add the caveat that 'near zero' must be read against the spread over repeats. That combination — mechanism, diagnosis, decision, caveat — is what separates a senior answer from a textbook definition.
- Which of the two would you put in a stakeholder report, and what must accompany it?The held-out version, computed with the metric the model is judged on in production, because that is what the deployed performance depends on. It has to carry four labels: the metric, the data split it was computed on, the number of repeats, and the spread. Without those, two teams can publish contradictory charts for the same model and both be arithmetically correct.
- What would you change about the model once you see this gap on agent ID?Restrict its ability to isolate rare levels — a minimum number of rows per leaf, stronger regularisation, fewer splits — and re-encode the column as a generalising summary such as historical per-agent statistics computed strictly from earlier tickets, dropping the raw identity. Then re-measure: if held-out performance is unchanged without the column, ship the simpler model.
- The held-out importance of a feature is near zero but with a very wide spread. What do you do?Treat the result as uninformative rather than as evidence of no reliance. Widen the evidence base: repeat the permutation across cross-validation folds and pool, or evaluate on a larger held-out sample. Wide spread usually means the evaluation set is too small for the metric, and that same weakness undermines every other importance number in the same chart.
saying these in an interview costs you the question
- Assumes permutation importance is always on held-out data
- Concludes the feature has no predictive value at all
- Reads the gap as randomness rather than overfitting
- Publishes an importance chart without naming the split
- Calls near-zero importance zero without checking the spread