Why can permutation importance rank features differently from a tree ensemble's built-in importance?
answer
- fit structure versus measured performance
- criterion units versus metric units
- training rows versus held-out rows
- more split points, more chances to score
- check the error bars before reconciling
basics
~20 sThey measure different things. Built-in importance is read off the training-time structure of the fit in split-criterion units; permutation importance measures the drop in an evaluation metric on data you choose. Different quantity, different data, different ranking.
solid answer
~50 sA tree ensemble's built-in importance is accumulated while the trees are grown: it sums how much each split that used a feature improved the split criterion on the training rows. Permutation importance is computed afterwards, by shuffling a column and measuring how far a chosen evaluation metric falls, usually on held-out data. So the two differ on three axes at once — the quantity (split-criterion improvement versus metric loss), the data (training structure versus the rows you score on), and the units. On top of that, features with many possible split points get more chances to look good to the built-in measure, and correlated features share credit differently under the two. When they disagree, check whether the permutation error bars for the disputed features even separate, group correlated features, and then favour the permutation ranking on held-out data if the question is what the shipped model relies on.
go deeper
Know that a tree ensemble gives you an importance ranking for free from its training structure, while permutation importance is measured from performance on data you choose, and that the two need not match.
Explain the three axes of difference — quantity measured, data used, units — and why a feature offering many split points can score highly on the built-in measure without earning a permutation drop.
Walk a disagreement to ground: check the permutation spread, confirm the split, group correlated features, then state which ranking you publish for the decision at hand and why.
Set the standard for which explanation the organisation ships and what must be attached to it, so that model reviews are not settled by whichever chart was easiest to produce.
## Two rankings, computed from different things A gradient-boosted or bagged tree model can hand you an importance ranking for free, because the information was collected while the trees were built. Permutation importance is computed after the fact, from predictions alone. Presenting them side by side and expecting agreement is the mistake; they are not estimates of the same quantity. **Built-in importance** is a property of the fitted structure. It aggregates, over every split in every tree that used a feature, the improvement that split produced in the tree's split criterion on the training rows that reached it (some variants count splits instead). It is measured in the units of that criterion, it is computed on the data the model was fitted to, and it exists whether or not you have any evaluation data at all. The mechanics and biases of that measure are the subject of its own topic; what matters here is the comparison. **Permutation importance** is a property of measured performance. It scrambles one column on a chosen evaluation set and reports how far your metric fell. It is in the units of that metric, it can be computed on held-out data, and it requires no access to the model's internals — the same procedure works on a linear model or anything else that predicts. ## Where the rankings diverge - **Training structure versus generalisation.** Built-in importance is derived from the fit, so a feature the model leaned on to memorise training rows scores highly there even if it contributes nothing on unseen data. Held-out permutation importance will not reward it. - **Number of split opportunities.** A feature that offers many distinct cut points — a continuous measurement, or a categorical with very many levels — gets more chances to be chosen and to register criterion improvements, which inflates its built-in score relative to a binary flag carrying the same information. Permutation importance does not care how many split points a column offers, only what the model loses when it is scrambled. - **Which metric matters.** Built-in importance is tied to the criterion the trees were grown with. Permutation importance is tied to whatever metric you choose, so a feature that matters for ranking quality but not for accuracy at a fixed threshold can move up or down depending on what you measure. - **Correlated features.** Under the built-in measure, correlated columns split the credit because the tree picks one of them at each split more or less arbitrarily. Under permutation, they mask each other and both fall toward zero. Both distort a correlated group, but in different directions, so the two rankings can disagree strongly on exactly those features. - **Cost.** Built-in importance is already computed. Permutation costs one prediction pass per feature per repeat. ## Reconciling a disagreement When the two rankings put different features in the top three, work through this order: 1. **Is the disagreement real?** Look at the permutation spread across repeats. If the intervals for the disputed features overlap, permutation is not making a confident claim and there is nothing to reconcile. 2. **Which data was each computed on?** If your permutation run used training rows, redo it on held-out rows before drawing conclusions — half of apparent disagreements are really this. 3. **Are the features correlated?** Group them and permute the groups. A feature the built-in measure loves and permutation ignores is very often one of several near-copies. 4. **How many split points does the contested feature offer?** A continuous or very high-cardinality column ranked highly by the built-in measure and low by held-out permutation should raise suspicion of the split-opportunity effect rather than a real dependency. 5. **Which question are you answering?** For 'what does the shipped model rely on to perform in production', the held-out permutation ranking with the production metric is the one to publish. For understanding how the trees were grown, the built-in measure is the honest description of the fit. ## What good looks like Agreement between the two is reassuring but not required, and disagreement is diagnostic rather than embarrassing — it usually points straight at correlated groups, a high-cardinality column, or a train/held-out gap. The failure is publishing whichever one was cheaper without stating what it measures, then defending business decisions with a bar chart nobody can interpret.
- A third ranking built from mean absolute attributions matches neither. How do you read that?A global ranking built by averaging the absolute per-row attributions measures how much a feature moves the model's output, not how much your metric falls when it is scrambled. A feature can shift scores substantially while leaving the metric untouched — for instance if the movement never crosses the decision threshold. Treat the three as answers to three questions and pick the one matching the decision you are supporting.
- The two rankings agree completely. Does that prove the ranking is right?No. It shows two measurements of a fit are consistent, which is reassuring but not evidence about the world. Both can be distorted the same way by a correlated group, and neither is causal. Agreement raises confidence that the model really does lean on those columns; it says nothing about whether acting on them would change outcomes.
- When would you not bother computing permutation importance at all?When the built-in numbers are only being used to inspect how the trees were grown, or during rapid iteration where you need a rough signal and prediction passes over a large evaluation set are expensive. The moment a number leaves the notebook — into a report, a feature-removal decision, or a stakeholder conversation — pay for the held-out permutation version with the metric that matters.
saying these in an interview costs you the question
- Treats the two rankings as estimates of the same quantity
- Assumes disagreement means one of them is broken
- Compares a training-data ranking with a held-out ranking unknowingly
- Ignores that correlated features distort the two differently
- Publishes whichever ranking was cheaper without naming what it measures