As a counterfactual explanation for a demoted marketplace seller, why prefer '3 fewer late shipments' to '400 more orders'?
answer
- smallest change that flips the decision
- flipping it is only the entry requirement
- prefer few features, small moves
- can the person actually do it
- scale each feature by its spread
basics
~20 sBoth changes may flip the trust model's decision, but a counterfactual is only useful if the seller can act on it. Three fewer late shipments is a small change on one controllable feature; 400 extra lifetime orders may take years.
solid answer
~50 sA counterfactual explanation is the smallest change to a row's feature values that would have produced a different model decision — here, the demotion reversed. Validity is only the entry requirement: both candidates flip the model, so flipping cannot be the tiebreak. What separates them is **proximity** (how far the changed row sits from the real one), **sparsity** (how few features had to move), and **actionability** (whether the person can plausibly make the change, and in the direction proposed). Three fewer late shipments touches one feature, is a small step in the seller's own operating behaviour, and is reachable this quarter. Four hundred additional lifetime orders is a large move on an accumulated quantity that no single decision controls. A search ranking candidates by raw feature-space distance misses this, which is why distances are scaled per feature and paired with an explicit cost of change.
go deeper
Be able to define a counterfactual as the smallest change to the inputs that would have flipped the model's decision, and to say why the recommendation must be something the person can actually do. Naming both candidates as valid but only one as useful earns the credit.
Explain the ranking criteria properly: validity, proximity with per-feature scaling, sparsity in the count of changed features, and actionability. Be ready to say why unscaled distance misranks candidates and why sparsity is encouraged by penalising absolute deviations.
Show that you constrain the search rather than clean up after it — mutability and allowed direction declared per feature before the optimiser runs, plus a diverse set of options returned. Mention what it means when no near counterfactual exists inside the actionable subspace.
Own what the organisation is committing to when it publishes these to users. A change list reads as a promise, so the wording, the feature-mutability policy, who signs off on it, and the boundary between describing the model and predicting the world are decisions you set, not the modelling team.
## What a counterfactual explanation is Most explanation methods answer 'why this decision?'. A **counterfactual** answers a more directly useful question: 'what would have had to be different for the decision to go the other way?' Formally, given an input `x` that the model scored into an unwanted class, you search for a nearby point `x'` such that the model's prediction at `x'` is the desired class, while `d(x, x')` — a distance between the two rows — is as small as possible. The reported explanation is the *difference*: the handful of feature values that moved, and by how much. For a marketplace seller demoted by a trust score, two candidates might both clear the bar: - `late_shipments: 11 -> 8` - `lifetime_orders: 60 -> 460` Both are valid counterfactuals. Only one is worth sending to the seller. ## The four properties that separate good from valid **1. Validity.** The changed row must actually cross the model's decision boundary into the desired class. This is a hard requirement, checked by scoring `x'` with the model — never assumed. A 'counterfactual' that does not flip the prediction is simply wrong. **2. Proximity.** Among valid candidates, prefer the one closest to the original row. Closeness is what makes the explanation intelligible; a point on the far side of the feature space tells the seller nothing about their own situation. Distance is usually computed per feature and then summed, because raw units are not comparable — a change of 3 in a count that ranges 0 to 15 is enormous, and a change of 3 in a count that ranges 0 to 5000 is nothing. The standard fix is to divide each feature's change by a measure of that feature's spread in the data (its standard deviation, or a robust equivalent), so distances are expressed in typical units. Without that scaling, the widest-ranged column dominates and the search returns nonsense. **3. Sparsity.** Prefer changing few features over many. A recommendation touching one behaviour is something a person can hold in their head and act on; a recommendation touching nine features simultaneously is a shrug. Sparsity is encouraged by penalising the *number* of changed features, not only their total size — a penalty on absolute deviations tends to leave most features exactly untouched while moving a few decisively. **4. Actionability.** This is the property the seller example turns on. Even a small, sparse, valid change is useless if it is outside the person's control or unreachable in a relevant timeframe. Late-shipment counts respond to how the seller runs the next few weeks. Lifetime order volume is an accumulated history; going from 60 to 460 is not an action, it is a wait. Ranking by geometry alone cannot see this difference, so actionability enters as an explicit input: each feature is annotated with whether it may change at all, in which direction, and at what cost. ## Immutable and directional constraints The hardest version of the actionability problem is features that must never be proposed. If age is in the feature set, an unconstrained search can happily return 'be twelve years younger', because the optimiser only sees a cheap distance in a numeric column. Nothing in the mathematics knows that age is not a lever. The same applies to features that can move only one way — age increases, a tenure counter increases, a completed qualification is not un-earned. The discipline is to **declare mutability before searching, not filter afterwards**. Every feature is classified up front: freely actionable, actionable in one direction only, or immutable. The search is then constrained to the actionable subspace, so an unusable suggestion is never generated in the first place. Filtering afterwards is worse in two ways: it wastes the search, and it quietly changes what is being optimised, because the nearest valid point in the full space may have no near neighbour in the actionable one — which is itself information worth surfacing. ## Offering more than one A single counterfactual forces one route on the user. Good practice is to return a small **diverse** set of valid, sparse counterfactuals that differ in which features they move, and let the person pick the one that fits their circumstances. Diversity is added as an explicit objective; otherwise the search returns several near-duplicates of the same change. ## The limit to state out loud A counterfactual is a statement about the model's decision boundary given a set of inputs: had these values been these instead, this scoring function would have output the other class. It is not, by itself, a promise about what happens in the world if the seller changes their behaviour — establishing that is a separate question with its own machinery. Saying this plainly is what separates a candidate who understands the method from one repeating a marketing line about actionable explanations.
- A counterfactual search returns 'be twelve years younger'. What went wrong?Mutability was never declared. The optimiser sees age as an ordinary numeric column with a cheap distance and moves it like any other. The fix is to classify every feature before the search as freely actionable, actionable in one direction only, or immutable, and constrain the search to the actionable subspace, rather than filtering unusable suggestions out afterwards.
- Why scale each feature by its spread when measuring how far the counterfactual moved?Because raw units are not comparable. Moving a count with range 0 to 15 by three is a large change; moving one with range 0 to 5000 by three is nothing. Dividing each change by that feature's spread expresses every move in typical units, so the widest-ranged column no longer dominates the distance.
- Why return several counterfactuals rather than the single closest one?Because the algorithm cannot see the person's constraints in full. A diverse set that moves different features lets the seller pick a route they can actually take. Diversity has to be an explicit objective, otherwise the search returns several near-duplicates of the same change and the choice is illusory.
It is like directions out of a dead-end street. The nearest exit on the map is not the useful one if it goes through a wall; you want the shortest route the driver can actually take.
saying these in an interview costs you the question
- Accepts any change that flips the prediction as a good explanation
- Ranks candidates by raw unscaled feature distance
- Proposes changes to immutable features like age
- Changes many features at once and calls it minimal
- Promises the real-world outcome will change, not just the model's decision