A sampling-based explainer returns different top-3 features on two runs for one prediction. What do you do?
answer
- two runs is an anecdote
- estimate has a standard error
- some churn no budget can fix
- compare spread against the gaps
- reproducible is not reliable
basics
~10 sSeparate estimator noise from real ambiguity: rerun across many seeds, measure rank agreement, and raise the sample budget until each spread is small next to the gaps you rank.
solid answer
~50 sFirst diagnose. Rank churn on a fixed model and a fixed row has three usual causes: Monte Carlo sampling error in the explainer, genuine ambiguity between redundant features whose credit split has no true value, and a row that sits where several features contribute almost identically. Measure before reacting - rerun the explanation 20-50 times with different seeds on the same ad-auction bid row and report the spread of each attribution plus how often each feature lands in the top 3. If the spread is large relative to the gaps, increase the sample budget until it is not; that is the only cause budget fixes. If it stays large, the ranking is not identifiable at this row and the honest presentation is a grouped or banded explanation - `these four features dominate, in no reliable order` - rather than a numbered list. Fixing the seed buys reproducibility, never reliability.
go deeper
Know that sampling-based explanations are estimates with error, so two runs on the same row can legitimately disagree without anything being broken.
Explain the causes you would separate - sampling error, redundant features, a genuinely knife-edge row - and why only the first responds to a larger sample budget.
Demonstrate the measurement: reruns across seeds, spread per attribution, top-1 agreement, sign flips, and a decision about what you are willing to show a stakeholder.
Own the standard: what churn is acceptable across seeds and across retrains, who monitors it, and when the product stops showing a ranked list at all.
## What the symptom actually tells you Same model, same row, two runs, different top-3. Nothing about the model changed, so what you are seeing is a property of the *estimator*, of the *feature set*, or of the *row* - and the three call for different responses. **Estimator variance.** Attribution methods that would otherwise be intractable estimate their quantity by sampling, so every reported attribution is an estimate with a standard error. Small budget, big error, unstable ranking. This is ordinary Monte Carlo behaviour and it is the one cause that a larger sample budget genuinely fixes. **Non-identifiable credit.** If two or more columns encode overlapping information, their individual shares are not pinned down by the model at all. Extra samples make each estimate more precise around an arbitrary target; the ranking between the redundant columns stays meaningless. Budget does not help here. **A knife-edge row.** Some rows genuinely sit where four or five features contribute nearly equally. The ranking is then unstable because the underlying numbers are within noise of each other, and any ranking would be over-reading. This is a presentation problem, not a bug. ## Measure the instability before you argue about it The mistake is reacting to two runs. Two runs is an anecdote. Turn it into a measurement: - Rerun the explanation for the row with, say, 30 different seeds. - For each feature, report the mean attribution and its spread across reruns (standard deviation, or a 5th-95th percentile interval). - Report a rank-agreement summary: how often the same feature is top-1; the average overlap between the top-3 sets of two runs; whether any feature's sign flips. Now you have numbers to act on. A sign flip is a red alert - the explanation cannot even agree on the direction. Top-1 agreement near chance means the ranking is unusable at any budget you can afford. Tight intervals with stable ordering mean the explanation is fine and you were unlucky with two draws. Do this on a *sample of rows*, not just the one that raised the alarm. Instability is often concentrated - near a decision boundary, or in a segment with unusual feature values - and knowing which rows are unstable is more actionable than a single global verdict. ## Choosing a sample budget The right budget depends on the distinction the explanation is used to make. If the product only ever shows the single strongest driver, you need the top-1 gap to exceed the estimation error - often a modest budget. If someone will compare the third and fourth drivers, you need precision on a much smaller gap, and the cost can be an order of magnitude higher. Set the budget from the decision, and say what distinction the explanation does *not* support. ## What fixing the seed does and does not buy Pinning the seed makes the output reproducible: the same row produces the same list every time, dashboards stop flickering, and a screenshot can be reproduced during an audit. That is worth having. It does not make the estimate any better. The variance is still there; you have chosen one draw from it and hidden the rest. A candidate who answers 'set the seed' and stops has confused reproducibility with reliability - and the moment the model is retrained, the pinned draw moves anyway. The right sequence is: measure across seeds, fix the instability you can, present honestly what remains, and *then* pin a seed for reproducibility. ## Instability that comes from retraining Worth separating explicitly, because stakeholders experience the two identically. Estimator variance is the same model explained twice. Model variance is a different fitted model - a weekly retrain, a different data slice, a different random initialisation of the training procedure - explained once each. Both change the story a user sees, but the remedies differ: budget and grouping for the first, model-stability work and explanation-versioning for the second. If your explanation is shown to users, you should be tracking churn across retrains as its own metric, not just across seeds. ## Presenting an unstable explanation If the ranking does not survive measurement, the answer is not to present it more confidently. Options, roughly in order of preference: group the redundant features and report groups, which frequently removes the churn outright; present bands rather than ranks (`dominant`, `moderate`, `negligible`); report the attribution with its uncertainty interval; or state that this row does not have a small set of separable drivers. A stakeholder who sees a numbered list assumes the numbers are real - the list is a claim, and unstable numbers do not support it.
- Is pinning the random seed a fix?It gives reproducibility, not reliability. The estimator's variance is unchanged; you have frozen one draw from it and stopped seeing the rest, so an audit that reruns the pinned pipeline learns nothing about whether the ranking is real. Measure stability across many seeds first, act on what you find, then pin a seed so outputs are reproducible. And expect the frozen draw to move at the next retrain.
- How do you decide how many samples the explainer needs?From the distinction the explanation is used to make. Increase the budget until each attribution's spread is small compared with the gaps you intend to rank. Showing only the single strongest driver needs the top-1 gap to clear the error, which is usually cheap; letting someone compare the third and fourth drivers demands precision on a far smaller gap and can cost an order of magnitude more compute.
- The explanation also changes after a weekly retrain. Is that the same problem?No. That is model variance, not estimator variance - a different fitted model explained once, rather than the same model explained twice. Both look identical to a stakeholder, but the fixes differ: sample budget and feature grouping address estimator noise, while retrain churn needs model-stability work, warm starts or explanation versioning. Track them as separate metrics if explanations are user-visible.
- How would you present an explanation whose ranking stays unstable after all of that?Stop presenting a rank. Group correlated features and report group contributions, which often removes the churn; otherwise show bands - dominant, moderate, negligible - or attach uncertainty intervals to each contribution. If the row genuinely has no small set of separable drivers, say so. A numbered top-3 is a claim of separability, and you should not make a claim the measurement does not support.
saying these in an interview costs you the question
- Fixes the seed and calls the explanation stable
- Reacts to two runs without measuring across many
- Assumes more samples cure every kind of churn
- Shows a stakeholder a top-3 with no rerun agreement checked
- Blames the model for what is estimator sampling error
- Believes a high-accuracy model implies stable explanations