How is permutation importance computed for a single feature, and what does the number mean?
answer
- break one column, keep the rest
- score before minus score after
- no refitting anywhere in the loop
- several shuffles, report the spread
- answer lands in metric units
basics
~20 sPermutation importance shuffles one column's values across held-out rows, breaking that column's link to the label, then re-scores the already-trained model. The drop from the baseline score, averaged over several shuffles, is the feature's importance.
solid answer
~50 sYou start from a trained model and an evaluation set it never learned from, and score it with the metric you actually care about — that is the baseline. Then you randomly permute the values of one column across the rows, leaving all other columns and the labels alone, and re-score the same model without retraining it. `importance = baseline_score - permuted_score` for a higher-is-better metric. Shuffling preserves the column's marginal distribution but destroys its relationship to the target, so the drop tells you how much this model's measured performance leans on that column. Because a single shuffle is random, you repeat it ten to thirty times and report the mean and the spread. The number is in the units of your metric — an AUC drop of 0.07 is a real dependency, a drop of 0.002 with a spread of 0.01 is noise.
code
python · 24 linesimport random
random.seed(0)
# 400 held-out rows: (inlet_temp, unrelated_noise); label is 1 when temp > 50
rows = [(random.uniform(0, 100), random.uniform(0, 100)) for _ in range(400)]
y = [1 if t > 50 else 0 for t, _ in rows]
def predict(row): # the trained model leans on column 0 only
return 1 if row[0] > 50 else 0
def accuracy(table):
return sum(predict(r) == label for r, label in zip(table, y)) / len(y)
baseline = accuracy(rows)
for col in (0, 1):
drops = []
for _ in range(20): # 20 shuffles -> mean and spread
table = [list(r) for r in rows]
values = [r[col] for r in table]
random.shuffle(values)
for r, v in zip(table, values):
r[col] = v
drops.append(baseline - accuracy(table))
print('column', col, 'mean drop', round(sum(drops) / len(drops), 3))go deeper
Be ready to state the loop out loud: score the trained model, shuffle one column across rows, score again, take the difference, repeat, average. Remember that the labels and the other columns stay untouched and nothing is retrained.
Explain why shuffling preserves the column's distribution while destroying its link to the target, why the result is in metric units, and how the number changes if you swap AUC for log loss. Know the F x R cost.
Show that you report importance with its metric, split, repeat count and error bars, and that you refuse to rank two features whose intervals overlap. Be ready to say when you would spend the refits on drop-column instead.
Own the reporting standard: which importance method the team publishes, on which data, with what noise floor, and how those numbers are allowed to be used downstream. An unlabelled bar chart is a governance problem, not just a plotting one.
## The idea Permutation importance answers one narrow question: **how much does this already-trained model's measured performance depend on this one column?** It is model-agnostic — it needs only the ability to make predictions, so it works the same for a linear model, a gradient-boosted ensemble, or anything else you can score. Nothing about the model's internals is inspected. ## The procedure, step by step 1. Fix a trained model and an evaluation table of rows the model did **not** learn from, together with their true labels. 2. Score the model on that table with the metric that matches the decision the model supports — AUC, log loss, F1, RMSE, mean absolute error. Call this the **baseline**. 3. Choose one feature. Randomly permute its values **across the rows**, leaving every other column and the label column untouched. The column keeps its own distribution — the same values, the same histogram — but its correspondence with the label, and with the other columns in the same row, is destroyed. 4. Re-score the **same** model on the corrupted table. No refitting happens at any point. 5. Importance = `baseline - permuted_score` when higher is better, or `permuted_error - baseline_error` when the metric is an error. Either way, a bigger positive number means the model relied on the column more. 6. Repeat steps 3-5 with different shuffles — ten to thirty repeats is typical — and report the **mean and the standard deviation**. 7. Restore the column and move to the next feature. ## Reading the number The importance is expressed in the units of the metric you chose, which makes it directly interpretable: 'shuffling this vibration channel costs the wind-turbine failure model 0.07 of held-out AUC, averaged over 20 shuffles, with a standard deviation of 0.01.' A second channel on the same model might cost 0.001 with a standard deviation of 0.004 — indistinguishable from zero, so the model does not measurably lean on it. Because the metric changes what is measured, the ranking can change with the metric. Importance under log loss rewards features that sharpen probabilities; importance under accuracy only rewards features that move rows across the decision threshold. Pick the metric you would defend in production and say which one you used. **Negative values are normal.** A useless column will sometimes score slightly better after shuffling, purely by chance. Read a small negative number as 'zero, within noise', not as 'this feature hurts the model'. If a large negative value appears repeatedly, suspect a broken evaluation set rather than a helpful shuffle. ## Cost With F features and R repeats you pay F x R full prediction passes over the evaluation set, plus one baseline pass. There is no training in the loop, which is what makes the method affordable on wide tables. Contrast this with **drop-column importance**, which removes the column and refits: on a 200-feature crop-yield model that is 200 full retrains versus 200 cheap shuffles. The two also answer different questions — drop-column asks 'how good could a model trained without this column be?', permutation asks 'how much does *this* model lose when the column is scrambled?' ## Choices that change the answer - **Which data.** Permuting the training rows measures what the fit leans on; permuting held-out rows measures what generalises. Say which you did. - **How many repeats.** One shuffle is a single random draw. The spread across repeats is the noise floor that any claimed difference between two features must clear. - **Evaluation-set size.** A small validation set gives wide error bars; repeating the whole procedure across cross-validation folds tightens them. ## What it does not tell you It is a statement about the model, not about the world. A high importance does not mean the feature causes the outcome, and a low importance does not mean the column carries no signal — it may mean a correlated twin column carried the information instead, so shuffling one of them changed nothing. Shuffling also builds rows that never occur in reality when features are correlated, so part of a measured drop can reflect the model being asked about impossible combinations. A good report therefore carries four things: the metric, the data split, the number of repeats, and the spread — not just a sorted bar chart.
- Why not just drop the column and refit the model instead of shuffling it?That is drop-column importance, and it answers a different question: how good a model trained without the column could be, not how much the current model leans on it. It also costs a full retrain per feature — 200 refits on a 200-feature crop-yield model versus 200 cheap prediction passes. Use it when you are deciding whether to stop collecting a feature; use permutation when you are explaining the model you are shipping.
- A feature's permutation importance comes out slightly negative. What do you conclude?Nothing, usually. Shuffling a column the model does not rely on can improve the score by chance, so small negative values scatter around zero. Judge it against the spread over repeats: if the mean is within a standard deviation of zero, report it as 'no measurable reliance'. A persistently large negative value points at a broken evaluation set or a mismatched metric, not at a harmful feature.
- How many repeats do you use, and what do you do with the spread?Ten to thirty is typical, more when the evaluation set is small. The spread across repeats is your noise floor: report mean plus or minus a standard deviation, and refuse to call feature A more important than feature B when their intervals overlap. Repeating the whole procedure across cross-validation folds also folds in the variance from the split itself, which is usually larger than the shuffle variance.
It is like testing which instrument a band depends on by having one player improvise nonsense while everyone else plays the score, then asking the audience how much worse it sounded.
saying these in an interview costs you the question
- Says the model is retrained after each shuffle
- Claims the drop measures the feature's causal effect
- Shuffles the labels instead of the feature column
- Reports a single shuffle with no repeats or spread
- Treats a small negative importance as a harmful feature
- Never states which metric or which data split was used