skip to content

Permutation importance ranks a feature first — what business decisions does that actually license?

level: principalimportance: nice to knowfreq 32%

answer

  1. a claim about the model, not the world
  2. monitoring and data spend, yes
  3. levers and bulk deletion, no
  4. proxies and consequences look identical
  5. refit before you cull features

basics

~10 s

Only model-facing ones: monitoring priority, data-collection spend, sanity checks against domain expectation. It measures one model's reliance on a column under one dataset and metric, not evidence that changing the quantity would change outcomes.

solid answer

~50 s

The number says: scramble this column on this evaluation set and this trained model loses that much of this metric. That licenses model-facing decisions — watch this column hardest for drift and quality breaks, keep paying for the data it comes from, escalate if it disagrees with domain expectation, and document it in model risk. It does not license the two things stakeholders usually want. It is not a causal claim, so 'the top feature is the lever to pull' does not follow: the column may be a proxy for something upstream, or a downstream consequence of the outcome. And it does not license bulk feature deletion by rank, because correlated columns mask each other and a refit redistributes reliance anyway. Causal questions need an experiment or a causal design; feature selection needs iterative refits measured on the metric you care about.

go deeper

for a junior

Remember the boundary: importance describes what the trained model leans on, not what causes the outcome. Never promise a stakeholder that changing the top feature will change results.

for a middle

Be able to explain why a proxy and a genuine driver produce identical importance, and why an importance ranking measured on one fit does not survive a refit on a reduced feature set.

for a senior

Show the two conversions you make in practice: a causal request becomes an experiment design, and a feature-cull request becomes an iterative refit measured with repeated cross-validation.

for a principal

Own the policy — what a published ranking must carry, that correlated features are reported as groups, that causal language is barred from importance reporting, and that removals need refit evidence rather than rank order.

## What the number is a statement about A permutation importance is a measurement on a **model**, conditional on four things: the trained model, the evaluation data, the metric, and the state of every other column. Change any of them and the number can change. Read literally, it says: *if this column's values are scrambled, this model's score on this data falls by this much.* Everything a business wants to do with the chart has to be checked against that sentence. ## What it does license - **Monitoring priority.** The columns the model leans on hardest are the ones whose drift, outage, or schema change will hurt production first. Importance is a sensible way to rank what your data-quality alerts should cover. - **Data-acquisition and vendor spend.** A costly feed whose column earns a negligible held-out drop, and whose correlated group also earns a negligible drop, is a candidate for cancellation. Confirm with a refit before you cancel the contract. - **Sanity checking against domain knowledge.** When a model leans on a column that no domain expert can justify, that is a reason to investigate the data pipeline and the split. Reassuringly boring rankings are weak evidence of a sane model; surprising ones are strong prompts to dig. - **Documentation and model risk.** Regulators and reviewers ask what a model depends on. A labelled importance report — metric, split, repeats, spread — is a legitimate answer to that question. ## What it does not license **It is not a causal effect.** The most common misuse is 'the top-ranked feature is the lever to pull'. Three ways that fails. The column may be a **proxy** for an unmeasured driver, so intervening on the proxy changes nothing about the driver. The column may be a **consequence** of the outcome rather than a cause, recorded before you look but generated after the fact. And even for a genuine cause, the model measures association in the observed regime; a deliberate intervention moves the system to a regime the model never saw, and the relationship it learned need not hold there. Causal questions need experimental design or an explicit causal identification argument. The right move when a stakeholder asks for a lever is to convert the request into an experiment, not to point at a bar chart. **It is not a feature-selection rule.** Deleting the lowest-ranked forty columns by permutation importance fails for two independent reasons. Correlated groups mask each other, so a group carrying real signal can present forty small numbers and be deleted wholesale. And once you refit without those columns, the model redistributes its reliance — the importance you measured belonged to the old fit, not the new one. Honest selection is iterative: remove a block, refit, measure the metric you care about with repeated cross-validation, keep the reduction only if the interval says performance held. **It is not a statement about the data's information content.** A near-zero column may still carry signal the model failed to use, because of encoding, capacity, or a correlated substitute. **It is not stable enough to defend a single ordering.** Two features whose intervals overlap are not ranked relative to each other, no matter what order the bar chart draws them in. ## The organisational call The leadership decision is not which method to run; it is **what the organisation is allowed to conclude from an importance chart**. A workable policy has four parts: every published ranking carries its metric, split, repeat count and error bars; correlated features are reported as groups; causal language is not permitted in importance reporting and requests for levers are routed to experiment design; and feature removals are backed by refit evidence, not by rank. That policy costs a little friction and prevents the two expensive failures — a business initiative aimed at a proxy, and a feature cull that quietly degrades the next model. ## The interview answer Grant the useful uses first so you do not sound obstructive, then draw the line clearly between association measured on a fitted model and effects in the world, and finish with what you would actually do when the stakeholder pushes: run the experiment for the causal question, run the refit for the selection question, and publish the ranking with its four labels either way.

  • A stakeholder wants to act on the top-ranked feature. How do you turn that into a defensible plan?
    Reframe it as an intervention question and design for it: define the action, the population, and the outcome window, then run a randomised test or, if randomisation is impossible, state the causal assumptions explicitly and test them. Use the importance chart only to prioritise which hypotheses are worth the cost of an experiment, and say plainly that the model cannot settle the question on its own.
  • What does a responsible feature-reduction process look like if not ranking by importance?
    Iterative and measured. Group correlated features, remove a block, refit, and evaluate the production metric with repeated cross-validation, keeping the reduction only when the interval shows no meaningful loss. Repeat until removals start to cost. Importance is useful for choosing which block to try first; the refit result, not the ranking, is the evidence.
  • What must every published importance chart carry before it leaves your team?
    The evaluation metric, the data split it was computed on, the number of repeats, and the spread — plus group labels where features are correlated. Without those, two arithmetically correct charts for the same model can contradict each other, and readers have no way to tell which differences are real and which are shuffle noise.

Knowing which ingredient a chef relies on most tells you what to keep in stock. It does not tell you that shipping more of it to every household improves the nation's cooking.

saying these in an interview costs you the question

  • Calls the top-ranked feature the causal driver of the outcome
  • Deletes the lowest-ranked features in bulk without refitting
  • Presents importance as a property of the data, not the model
  • Treats an unlabelled bar chart as sufficient documentation
  • Assumes the ranking will hold after the model is retrained

context