Your model card lists impurity importance as the key drivers - what claims does that ranking support?
answer
- a claim about the model
- conditional on the other columns
- in-sample, unsigned, unstable
- pin the definition before publishing
- intervention needs designed evidence
basics
~20 sOnly a narrow one: for this fitted model, on its training rows, splits on that column removed the largest share of impurity, given the other columns present. It is not a causal effect, not signed, not stable, and not a property of the data.
solid answer
~50 sThe ranking supports a statement about the model, not about the world. It says which columns this particular fit leaned on to reduce training impurity, conditional on every other column being available, under this depth and tree count. It says nothing about direction, nothing about what happens if you intervene on the feature, and nothing about a differently configured model. So before a ranking goes on a card I want three things: the definition pinned down, since summed gain and split counts rank differently; evidence of stability across seeds and folds, with the spread published, not just the point ordering; and an audit showing the top of the list is not tracking column cardinality or hiding duplicate columns splitting credit. Then the card states the claim in those words. If someone wants to move a lever, that is a question for designed evidence, and an importance table cannot answer it.
go deeper
When you hand a ranking to anyone, say two things with it: it was computed on the training data, and it is relative to the other columns in the model. That sentence prevents most of the misreadings.
Be able to list what the score cannot say - no direction, no causal effect, no statement about a differently configured model - and know that near-duplicate columns split credit, so a low rank does not prove irrelevance.
Show that you productionise the caveats: named definition, recorded seed and snapshot, rank ranges across refits, and a cardinality audit on the top entries before anything is published or monitored.
Own the policy. Decide the house definition, decide what the ranking may and may not be used for, and give the organisation a standard answer for the request to act on a feature so the decision does not get made informally inside a slide.
## The claim the arithmetic actually licenses An impurity importance score is a sum of weighted impurity drops over the splits a fitted model made on one column, computed on the rows it was trained on. Unpacked, a published ranking supports exactly this sentence: > Within this fitted model, trained on this snapshot with this configuration, splits on column X accounted for the largest share of the training impurity the model removed, given that all the other columns were also available. Every qualifier in that sentence is load-bearing, and each one blocks a claim readers routinely make: - **"Within this fitted model"** blocks "X is the most important predictor of the outcome". Refit with a depth cap, a different seed or a different algorithm and the ordering can change. - **"On this snapshot"** blocks any claim of durability. Importance drifts as the population drifts. - **"Share of impurity removed"** blocks percentage-of-outcome readings. It is not variance explained and not a contribution to accuracy. - **"Given the other columns"** blocks "X matters, Y does not". Two near-duplicates split their credit and both look mediocre; drop one and the survivor jumps. Low rank can mean redundant, not irrelevant. - Nothing in the sentence is signed, so it never supports "more X means more outcome". - Nothing in it is causal, so it never supports "raise X and the outcome follows". ## What to require before publishing **Pin the definition.** Summed gain, split counts and coverage-weighted measures rank the same model differently. A card that says "feature importance" without naming which one, how it was normalised, which data snapshot and which seed, has published an unreproducible number. Choose one definition as the house standard so rankings are comparable across the team's models, and record the rest. **Publish stability, not just order.** Refit across seeds and cross-validation folds and report the range of each feature's rank. In most tabular models the top two or three are stable and the middle is noise. A card that shows rank intervals is honest; a card that shows a clean ordered list of twenty features implies a precision that does not exist. **Audit the top of the list before you defend it.** Print distinct-value counts alongside the scores - if rank tracks cardinality, the ranking is measuring the splitter's opportunity set. Look for near-duplicate columns diluting each other. Check whether the ordering survives a depth-capped refit; a feature that only ranks high when trees grow deep is earning its place from memorisation. **State the training-fit basis in the card text itself.** Not in an appendix. The single most common misreading is that these numbers were validated, and one sentence prevents it. ## The organisational call The hard part is rarely the statistics; it is that a ranking is a compelling artefact. It is short, ordered and numeric, so it travels into decks and gets read as a list of causes. Three decisions worth owning as a lead: **Decide what the ranking is allowed to be used for.** Legitimate uses are model debugging, sanity-checking against domain expectation, monitoring for shifts between refits, and documentation. Deciding to intervene on a feature is not among them, and neither is telling a business that a column "drives" an outcome. **Give the escalation a standard answer.** When someone asks to act on the number three feature, the answer is not "the model says so" and not a flat refusal. It is: the model learned an association under the regime that generated this data; whether changing the feature changes the outcome is a different question that needs designed evidence - an experiment, a natural experiment, or an explicit causal argument with its assumptions stated. Offer to help scope that instead. **Accept a boring card over a persuasive one.** Rank intervals, a named definition and a stated in-sample basis make a duller artefact than a clean top-ten bar chart, and they are the reason the card survives contact with a reviewer, an auditor or a model risk function. ## The failure mode to name out loud The worst outcome is not a wrong ranking; it is a right ranking that nobody can reproduce and everybody over-reads. A number computed correctly, published without its qualifiers, and repeated until it becomes an organisational belief is harder to unwind than an obvious error - because there is no moment where it visibly breaks.
- What evidence would you require before a ranking is allowed on a model card?The definition and normalisation named, the data snapshot and seed recorded, rank ranges across several refits or folds rather than one ordering, and an audit showing the top entries are not tracking column cardinality or masking near-duplicate columns that split credit between them.
- A stakeholder wants the number three feature turned into a policy lever. How do you answer?Say what the number is: an association this model exploited under the regime that produced the data. Whether moving the feature moves the outcome is a causal question the ranking cannot answer. Offer to scope evidence that can - an experiment or an explicit causal argument - rather than approving or blocking outright.
- Why is a clean ordered top-ten list a worse artefact than one with rank intervals?Because the ordering below the top few is usually within refit noise. A clean list implies a precision the numbers do not have, and readers anchor on positions that would swap under a different seed. Intervals make the uncertainty visible and make the card defensible under review.
saying these in an interview costs you the question
- Presents the ranking as a list of causal drivers
- Publishes an ordering without saying which definition produced it
- Treats a low-ranked feature as proven irrelevant
- Reports one seed's ordering as if it were stable
- Reads the normalised scores as contributions to accuracy