skip to content

What does it mean for a local explanation of one prediction to be faithful rather than plausible?

level: middleimportance: must knowfreq 60%

answer

  1. two different reference points
  2. one compares to the model
  3. the other compares to the reviewer
  4. convincing story, wrong mechanism
  5. snow in the background

basics

~10 s

Faithful means the explanation reflects what the model actually computed for that row. Plausible means it matches what a human expects. An explanation can read convincingly while the model keys on something else.

solid answer

~50 s

Faithfulness is a property of the explanation relative to the **model**: does the attribution describe the computation that produced this prediction? Plausibility is a property relative to the **reviewer**: does the story sound reasonable given domain knowledge? They are independent, and the failure mode is a plausible-but-unfaithful explanation, because nobody challenges it. The classic illustration is the husky-versus-wolf demonstration from the LIME paper: a classifier that appears to separate the two breeds is in fact keying on snow in the background, and an explanation pointing at fur or ear shape would have been accepted without a second look. Plausibility is what an explanation gets for free from a reviewer who already believes the model; faithfulness has to be tested, for example by degrading the top-ranked features and checking the prediction actually moves more than under a random ranking.

go deeper

for a junior

Be ready to state the difference in one line: faithful is about the model, plausible is about the reader. Know that an explanation looking reasonable is not evidence the model is reasonable.

for a middle

Explain why the plausible-but-unfaithful case is the one that hurts, and name at least one concrete faithfulness test such as degrading top-ranked features and comparing against a random ranking.

for a senior

Show you have caught this in practice: describe how you audited an explanation before letting a stakeholder act on it, and how you handled an explanation that was surprising but held up under testing.

for a principal

Own the framing that plausibility is what an organisation buys by default. Decide what evidence of faithfulness is required before explanations become part of a review process or a customer-facing surface.

## Two different questions When you produce an explanation for a single prediction - a list of features with signed contributions, a highlighted region of an image, a small rule - there are two separate questions you can ask about it, and interviews on this topic exist mainly to check that you keep them apart. **Faithfulness** asks: does this explanation describe what the model actually did? The reference is the model. A faithful explanation of a single row identifies the inputs the model's output genuinely depended on at that point, in roughly the proportions it depended on them. **Plausibility** asks: does this explanation look right to a human? The reference is the reader's domain knowledge. A plausible explanation names features a domain expert would have named. These are independent properties. Every combination exists: - Faithful and plausible - the happy case, and the one people assume they are in. - Faithful but implausible - the model really is leaning on a feature nobody expected. This is uncomfortable but it is *information*, and it is often how leakage and shortcut features get found. - Plausible but unfaithful - the dangerous case. The explanation names exactly the features the reviewer already believed in, so the review stops. - Neither - usually spotted quickly, because implausible explanations get scrutinised. ## Why the dangerous case is dangerous A plausible explanation is *self-ratifying*. The reviewer's own prior is the only check being applied, so an explanation that matches the prior passes review by construction. That is the mechanism behind the husky-versus-wolf demonstration from the LIME paper: a classifier that seems to distinguish the two animals is actually responding to snow in the photo background, because in the training images wolves were photographed on snow. An explanation highlighting fur texture or muzzle shape would be received as obviously correct. The model would still be broken. The lesson is not about images. Any post-hoc explanation of any model can be plausible without being faithful, and the more domain-shaped the feature names are, the more easily a wrong explanation slips through. A churn model that reports 'days since last login' as the driver will be believed whether or not that is what the model used. ## Testing faithfulness instead of assuming it Because plausibility is free and faithfulness is not, faithfulness has to be measured. Two families of check are standard: **Perturbation or deletion checks.** Take the features the explanation ranks highest, degrade them - replace with a baseline value, mask the region, corrupt them - and re-score the row. If the prediction moves substantially more than when you degrade the same number of randomly chosen features, the ranking is carrying real signal. If the two curves overlap, the explanation is not tracking the model. Caveat, and a good one to raise unprompted: deleting or replacing values itself creates input rows the model never saw during training, so the test has some of the same off-distribution weakness as the explanation methods it is auditing. Averaging over many rows and comparing against the random baseline is what makes it usable. **Model-randomisation sanity checks.** Recompute the explanation after scrambling the trained model's parameters. If the explanation looks essentially the same for the trained model and a randomised one, it is a function of the input, not of the model, and it cannot be faithful to anything. Published work on saliency maps found several popular methods failing exactly this check. A third, cheaper option in low-stakes settings: fit a model that is transparent by construction on the same data, and see whether the post-hoc explanation of the complex model at least agrees with it in the region you care about. Disagreement does not prove unfaithfulness, but it is a prompt to look harder. ## The boundary worth stating out loud Even a perfectly faithful explanation tells you about the model, not about the world. 'The model's output depended heavily on this feature for this row' is not 'changing this feature would change the outcome'. The model may be leaning on a proxy that is only associated with the outcome in the training data. Turning attributions into claims about real-world effect requires causal analysis, which is separate machinery entirely; the correct interview move is to name that boundary and stop, not to blur it. ## What to say in the room Define both terms, say they are independent, give the plausible-but-unfaithful case as the one that costs money, and name at least one concrete test you would run before an explanation is trusted. Candidates who say 'the explanation looked sensible so the model is fine' have answered the wrong question.

  • How would you actually test whether an explanation is faithful, rather than trusting that it reads well?
    Two checks. Degrade the top-ranked features and confirm the prediction moves more than when you degrade randomly chosen features of the same count - if the curves overlap, the ranking is noise. Second, recompute the explanation after randomising the trained model's parameters; if it barely changes, the method is describing the input rather than the model. Run both over a sample of rows, not one.
  • If a feature gets a large attribution for a row, does that mean changing it would change the outcome in the real world?
    No. At best it says the model's output depended on that input at that point. The model may be using a proxy that only correlates with the outcome in its training data, so an intervention in the world can do nothing. Moving from a faithful attribution to a real-world effect claim requires causal analysis; treat the attribution as a statement about the model and stop there.
  • Is an implausible explanation a reason to reject the explanation or the model?
    Neither, immediately - it is a reason to investigate. Check faithfulness first: if the explanation survives a deletion check, the model really is leaning on that feature, and you have probably found leakage, a shortcut feature or a genuine effect the experts did not expect. Discarding explanations because they are surprising throws away the only ones that could have told you something.

A witness statement can be entirely believable and still not describe what happened. Believability is the audience's judgment; accuracy has to be checked against the event.

saying these in an interview costs you the question

  • Treats a sensible-sounding explanation as proof the model is correct
  • Says a plausible explanation is by definition a faithful one
  • Reads attributions as real-world causal effects
  • Assumes a popular explanation method is validated for their model
  • Confuses explanation quality with model accuracy

context