skip to content

What would you require before per-decision explanations are shown to customers or regulators?

level: principalimportance: nice to knowfreq 24%

answer

  1. an explanation becomes a public claim
  2. justification versus telling someone what to change
  3. entry criteria before the launch date
  4. measure churn across retrains too
  5. a convincing fiction is the worst outcome

basics

~10 s

Treat the explanation as a product surface with its own bar: a faithfulness test the team runs, a measured stability budget across seeds and retrains, a stated purpose, and a fallback when checks fail.

solid answer

~50 s

I set entry criteria before anything ships. One: **purpose** - is this justification (`why was I declined`) or recourse (`what should I change`)? Recourse is a claim about the world, and a post-hoc attribution does not support it, so either the product promises less or we use a method built for that. Two: **faithfulness evidence** - a documented deletion-style check plus a model-randomisation sanity check, run on a sample of real rows. Three: **a stability budget** - published thresholds for rerun agreement across seeds and churn across retrains, monitored like any other production metric. Four: **presentation that matches the evidence** - grouped or banded contributions rather than a brittle numbered list. Five: **a fallback** - if a decision surface cannot meet the bar, use an inherently interpretable model there or say less. An unfaithful explanation shown to a customer is worse than none: it manufactures confidence and becomes legally load-bearing.

go deeper

for a junior

Know that an explanation shown to a customer is a claim the company is making, and that generating one is not the same as being able to defend it.

for a middle

Be able to list the checks that would go into a release gate: a faithfulness test on real rows, rerun agreement across seeds, and a presentation that does not over-claim a ranking.

for a senior

Show you would measure explanation churn across retrains in production and would push back on a launch whose copy promises recourse that a post-hoc attribution cannot support.

for a principal

Own the tiering: which decision surfaces earn per-decision guarantees, which get sampled auditing, and what the fallback is when a surface cannot meet the bar. Be explicit that a plausible but unfaithful explanation at scale is the costliest outcome.

## The framing to lead with Once an explanation is shown outside the modelling team it stops being a debugging aid and becomes a claim your organisation makes about a decision. People act on it, appeal against it, and quote it back. So the question is not 'can we generate explanations' - a tool will always produce output - but 'what would have to be true for us to stand behind this output'. Leads are hired for that reframe. ## Entry criteria worth defending **1. Name the purpose, because different purposes need different guarantees.** - *Justification*: telling someone which factors drove their outcome. Needs faithfulness to the model and stability of presentation. - *Recourse*: telling someone what to change to get a different outcome. This is a claim about what happens when the world changes, and a post-hoc attribution is a statement about the model's dependence on inputs, not about effects of intervention. Either the copy is written to promise only what the method supports, or the surface uses a method designed to produce actionable changes. - *Compliance artefact*: evidence for a reviewer. Needs reproducibility and an audit trail alongside faithfulness. - *Internal triage*: an analyst deciding what to look at next. The bar is deliberately lower; do not pay customer-grade validation costs here. Getting the purpose stated in writing kills most of the argument, because the usual stakeholder request - 'it makes the model feel trustworthy' - is a request for plausibility, and plausibility is what you get for free from any explanation that names familiar features. **2. Faithfulness evidence, produced by the team, on real data.** At minimum: a degradation check showing that suppressing the top-ranked features moves predictions materially more than suppressing random features of the same count, run over a representative sample of rows; and a model-randomisation check showing the explanation changes when the trained parameters are scrambled. Both should live in the pipeline, with results recorded per model version, not run once during the design review. **3. A stability budget, published and monitored.** Decide before launch what churn is acceptable: agreement of the top-ranked driver across seeds; overlap of the presented set across weekly retrains for a fixed row; frequency of sign flips. These are cheap to measure with a held-out probe set of rows scored at every model release. Explanations that change every week destroy trust faster than no explanation at all, and nobody notices until a customer compares two letters. **4. Presentation matched to the strength of the evidence.** If the measurements support grouped contributions but not a ranking, ship groups. Bands (`major factor`, `contributing factor`) survive noise that a numbered list does not. Where two columns encode the same construct, present the construct. The presentation layer is where over-claiming actually happens. **5. A fallback path.** Some decision surfaces will not meet the bar. The options are to use a model that is interpretable by construction for that surface and accept the accuracy cost, to reduce what the explanation asserts, or to route the case to a human with the underlying evidence rather than a generated rationale. Deciding this in advance is what prevents a launch-date argument in which the only remaining option is to ship something you do not trust. **6. Written limits.** A short internal document: what the explanation does and does not claim, which checks were run, what the measured stability is, who owns it. This is what protects the team when the explanation is quoted back in a dispute. ## Cost discipline All of this scales with stakes, and a lead should say so unprompted. Contested, appealable or regulated decisions justify per-decision guarantees and continuous monitoring. Low-stakes ranking or internal triage justifies sampled auditing on a probe set and nothing more. Applying the maximum bar everywhere is how interpretability work gets cancelled for being expensive. ## The organisational risk to name The worst outcome is not a missing explanation - it is a plausible, unfaithful one at scale. It passes review precisely because it matches what reviewers expect, it gives customers false confidence about what to change, and if it is ever used to defend a decision formally, the organisation has attested to a mechanism the model does not use. That asymmetry is the reason for the entry criteria: the cost of shipping nothing is understood and bounded, and the cost of shipping a convincing fiction is neither.

  • A stakeholder wants explanations because they make the model feel trustworthy. How do you respond?
    I name what is being asked for: that is plausibility, and any explanation naming familiar features delivers it whether or not it is faithful. Then I redirect to the decision - what will a user do differently after reading this, and what would a wrong explanation cost us? The answer sets the validation bar. If nothing changes downstream, we are paying to reassure ourselves, and that is worth knowing before we build it.
  • How do you keep the validation cost proportionate across many models?
    Tier by stakes. Contested, appealable or regulated decisions get per-decision guarantees, monitored stability and an audit trail. Internal triage and low-stakes ranking get sampled auditing on a fixed probe set at each release, and no per-request checking. Publish the tiers so teams self-select rather than negotiating each launch, and re-tier when a surface becomes customer-visible.
  • What do you do when a decision surface cannot meet the faithfulness bar at all?
    Three options, decided in advance rather than at launch. Use a model that is interpretable by construction for that surface and accept the accuracy cost; reduce what the explanation asserts to something the evidence supports, such as broad factor groups; or route the case to a human reviewer who works from the underlying evidence rather than a generated rationale. What I do not do is ship the explanation with hedging copy attached.

saying these in an interview costs you the question

  • Ships explanations because the tooling produces them
  • Treats a plausible-sounding explanation as compliance evidence
  • Promises customers recourse from a post-hoc attribution
  • Never measures the churn users see across retrains
  • Applies the same validation bar to every model regardless of stakes

context