skip to content

Why can two highly correlated features both look unimportant under permutation importance?

level: middleimportance: must knowfreq 62%

answer

  1. the twin column is still intact
  2. each masks the other's credit
  3. low importance is not no signal
  4. shuffling builds impossible rows
  5. permute the correlated group jointly

basics

~20 s

Because shuffling one of them leaves its twin intact, and the model recovers the same information from the twin, so the score barely moves. Each correlated column masks the other's credit, and both end up ranked near zero.

solid answer

~40 s

Permutation importance measures what the model loses when one column is scrambled while everything else stays put. If a manufacturing-yield model carries inlet temperature in both Celsius and Fahrenheit, shuffling the Celsius column leaves an intact Fahrenheit column the model can lean on, so the held-out score hardly drops — and the same happens in reverse. Both features report near-zero importance even though temperature is the model's main driver. Shuffling correlated columns also fabricates rows that cannot exist, such as 5 degrees Celsius next to 200 degrees Fahrenheit, so part of any drop you do see reflects the model extrapolating rather than losing information. The fix is to stop scoring the columns individually: cluster features by correlation and permute each group jointly, or confirm a suspicious zero by refitting without the whole group.

go deeper

for a junior

Remember the one-line reason: the other copy of the information is still there, so shuffling one column costs the model almost nothing. Do not conclude a feature is useless from a low number alone.

for a middle

Explain the masking mechanism precisely, note that it scales with the strength of the correlation rather than needing exact duplicates, and propose permuting correlated groups together as the fix.

for a senior

Show how you would build the groups on real data, how you would present group-level importance to stakeholders, and when you would spend a refit to confirm a surprising zero rather than trusting the shuffle.

for a principal

Frame it as a reporting-contract question: decide what your team publishes when features are correlated, since a per-column bar chart on a correlated table reliably produces wrong decisions downstream about what data to stop collecting.

## The mechanism Permutation importance holds every other column fixed while it scrambles one. That design is exactly what makes it a statement about *this* model rather than about the data — and it is also what breaks when two columns carry the same information. Take a manufacturing-yield model whose training table happens to include inlet temperature twice, once in Celsius and once in Fahrenheit. The two columns are a deterministic transform of each other, so any pattern the model could learn from one it can equally learn from the other. During training, the fit distributes its reliance between them in some arbitrary way. Now permute the Celsius column: the Fahrenheit column is untouched, still carries the full signal, and the model's predictions barely move. Score drop: near zero. Permute Fahrenheit instead: same story. Report both numbers side by side and temperature — the dominant physical driver of yield — appears to be irrelevant. This is not a bug in the arithmetic. The measurement is literally correct: **removing either column alone costs this model almost nothing**, because the other one is a perfect substitute. The mistake is reading 'unimportant column' as 'unimportant quantity'. ## It is a matter of degree Exact duplicates are the clean teaching case, but the effect is continuous. Two features correlated at 0.95 will each absorb most of the other's shuffle, and both drops shrink toward zero. At 0.6 the masking is partial and both features still show some importance, but less than either would show alone. Any correlated cluster — three sensor channels on the same shaft, several ratios built from the same two raw columns — dilutes the same way, spreading a real dependency across several small numbers that individually look like noise. ## The second problem: impossible rows Shuffling one column of a correlated pair does not just leave a substitute in place; it also *creates rows that could never occur*. After the shuffle, a row may show 5 degrees Celsius beside 200 degrees Fahrenheit, or a low-power reading beside a high-vibration reading that never co-occur in the plant. The model is being evaluated on inputs far from anything it was trained on, and its behaviour there is unconstrained. Whatever drop you measure is then a mixture of 'the model lost information' and 'the model extrapolated badly'. Both effects are real, but they are different findings, and the single number does not separate them. ## What to do instead - **Group and permute jointly.** Cluster the features by correlation — a hierarchical clustering on the correlation matrix with a cut at a chosen threshold is the usual route — then permute every column in a group together and record one drop for the group. In the temperature example, shuffling both columns at once produces the large drop that reflects the real dependency. - **Report the ranking at the group level.** Present 'temperature block: 0.09 AUC' rather than two near-zero rows that invite the wrong conclusion. - **Confirm with a refit.** If a group's zero importance is surprising, drop the whole group and retrain. That answers the different but often more useful question of whether the information is available anywhere else in the table. - **De-duplicate before you explain.** If two columns are the same quantity in different units, the honest fix is upstream: keep one. - **Condition instead of marginalise.** Conditional variants permute a feature only within groups of rows that are similar on the correlated columns, which keeps the rows plausible. They cost more and change the question slightly — they measure additional information beyond the correlated peers rather than total reliance — so state which one you ran. ## What to say in an interview Name the mechanism (the intact substitute), name the second-order effect (off-distribution rows), and land on the practical rule: **never conclude 'this feature does not matter' from a low permutation importance without first checking what it is correlated with.** A candidate who volunteers group permutation as the fix, and who notes that group importance and individual importance answer different questions, is showing the judgment the question is testing.

  • How would you produce a defensible ranking on a table full of correlated features?
    Cluster the features on the correlation matrix, cut the tree at a threshold you can justify, and permute each cluster as a unit — one drop per group. Report the ranking at group level, name the members of each group, and keep the per-feature numbers only as a secondary view. Where a group's result is decision-critical, confirm it by refitting without the group.
  • Is the low importance wrong, or is the question wrong?
    The number is right for what it measures: this model does not need that particular column, given the others. It is wrong only when it is read as a claim about the quantity rather than the column. The distinction matters in practice — you can safely stop storing one of two duplicate columns, but you cannot conclude that temperature does not drive yield.
  • Why does shuffling correlated columns push the evaluation off-distribution, and why does that matter?
    Permuting one column independently of its correlated peers produces feature combinations that never occur in the real process. The model has no training support there, so its predictions are arbitrary, and the measured drop blends lost information with extrapolation error. It matters because the two have different remedies: the first is about feature reliance, the second is about where your model is being asked to work.

Fire one of two engineers who both know the same system and nothing breaks; fire either alone and nothing breaks. Conclude that neither matters and the system goes down the day both leave.

saying these in an interview costs you the question

  • Reads a near-zero importance as no predictive signal
  • Recommends deleting both correlated columns as useless
  • Thinks correlation is only a problem for linear models
  • Never mentions permuting correlated groups jointly
  • Ignores that shuffling can create impossible feature combinations

context