skip to content

How does a local surrogate model explain one prediction of a black-box classifier?

level: middleimportance: must knowfreq 60%

answer

  1. explain near one row, not everywhere
  2. make fake neighbours, ask the model
  3. closer neighbours count for more
  4. fit something small and readable
  5. the coefficients are the explanation

basics

~10 s

A local surrogate perturbs the row being explained, labels each perturbation with the black box's prediction, weights it by closeness to that row, then fits a small interpretable model whose coefficients are the explanation.

solid answer

~50 s

The idea is to stop trying to explain the whole model and approximate it only near one point. You build a synthetic neighbourhood around the row being explained — for a text classifier, copies of the post with random words deleted; for tabular data, rows whose values are resampled or nudged. Every synthetic point is labelled by querying the black box, not by ground truth, because the target is the model's behaviour. Points are then weighted by a proximity kernel so close neighbours dominate the fit, and a deliberately simple model — usually a linear one restricted to a handful of terms — is fitted to minimise that weighted loss plus a complexity penalty. Its coefficients are the explanation: `refund` pushed this post toward removal, `thanks` pushed against. The result is valid only inside that neighbourhood; a different row can get the opposite explanation from the same model.

go deeper

for a junior

Be able to state the recipe in order: make many altered copies of the one row, ask the model to score each, weight them by how close they are, fit a small readable model. Knowing the sequence is most of the credit here.

for a middle

An interviewer expects the mechanics: why the labels come from the model rather than the dataset, what the proximity weights do to the fit, and why the surrogate is deliberately restricted to a few terms. Be ready to write the weighted loss plus complexity penalty in words.

for a senior

Show that you know what to check before showing an explanation to anyone — how well the sparse surrogate actually tracks the model over its neighbourhood, and whether the perturbation scheme makes sense for the data type. Say plainly that the explanation is one row deep.

for a principal

Own the question of what the organisation promises when it ships these explanations. Local, model-only, neighbourhood-dependent statements are easy for reviewers and regulators to over-read, so the framing, the retention of how each explanation was generated, and the wording shown to users are your call.

## The problem it solves Suppose a content-moderation classifier removed one forum post, and the author asks why. The model may be a large ensemble or an opaque scoring service; there is no coefficient to read off. A **local surrogate** sidesteps that by refusing to explain the model globally. Instead it answers a narrower, tractable question: *in the immediate vicinity of this one input, what simple rule behaves like the model?* The method is usually taught under the name LIME — local interpretable model-agnostic explanations. Model-agnostic means it needs only the ability to send inputs in and read predictions out; it never inspects weights, trees or gradients. ## The four steps **1. Choose an interpretable representation.** The features a human reads are rarely the features the model consumes. For text, the readable representation is a set of on/off indicators — 'the word refund is present', 'the word refund is absent' — even if the model actually consumes a dense numeric encoding. For images it is the presence or absence of contiguous image segments. For tabular data the raw columns usually serve, though continuous columns are often bucketed so the explanation reads as 'income in the top band' rather than a slope. **2. Perturb around the instance.** Generate many variants of the row in that readable space. For the removed post, that means hundreds of copies with random subsets of words deleted. For tabular data, values are resampled or jittered around the observed ones. This synthetic sample is the neighbourhood. **3. Label the neighbourhood with the black box.** Each variant is passed through the model and its predicted score is recorded. This is the step candidates most often get wrong: the labels are the *model's outputs*, never the dataset's true labels. The surrogate is being trained to imitate the classifier's decision surface, not to predict reality. If you trained it on true labels you would simply have built a second, competing model that explains nothing about the first. **4. Fit a weighted, sparse, simple model.** Each variant gets a weight from a proximity kernel — high for variants close to the original row, decaying with distance. The surrogate minimises `sum_i w_i * (f(z_i) - g(z_i))^2 + complexity(g)`, where `f` is the black box, `g` the simple model, `w_i` the proximity weight, and the complexity term keeps `g` readable. In practice `g` is linear with a small fixed number of non-zero terms, obtained either by selecting the top terms or by an L1 penalty that drives most coefficients exactly to zero. ## Reading the output The explanation is the surviving coefficients, each signed. For the removed post you might get: presence of a slur `+0.41`, presence of a link `+0.12`, presence of the word `thanks` `-0.08`. Read this as: within the deletion neighbourhood of this post, removing the slur moves the model's removal score down by about 0.41 on the surrogate's scale. It is a statement about the model, in one small region, in the readable representation you chose. ## Sparsity is a deliberate cost Restricting the surrogate to, say, five terms is not an accident of implementation — it is the whole point. A local approximation with two hundred terms is as unreadable as the black box. But every term you drop lowers how well the surrogate reproduces the model's outputs across the neighbourhood. The usual practice is to fix a small term budget for human reasons, then check the weighted fit quality to see whether that budget was enough to describe the region at all. A surrogate that fits the neighbourhood badly should not be shown to anyone as an explanation. ## What it does not give you Three limits are worth stating out loud in an interview. First, the explanation is **local**: it does not license any claim about how the model treats other posts, and the same model can produce opposite explanations for two similar rows. Second, it explains the **model**, not the world — a positive coefficient says the model's score rises with that term, not that the term causes anything. Third, the neighbourhood is something *you* defined; the size of it and the perturbation scheme are choices that change the answer, which is why a serious explanation is reported together with how it was generated.

  • Why are the perturbed samples labelled with the model's prediction instead of the true label?
    Because the surrogate's job is to imitate the classifier's decision surface, not reality. Training it on true labels would produce a second independent predictor whose coefficients describe the data, not the black box you were asked to explain. The whole exercise is model mimicry, so the model must supply the targets.
  • What is the interpretable representation for a text or image explanation?
    A set of on/off indicators a human can read: for text, whether each word is present or deleted; for images, whether each contiguous segment is kept or blanked. The surrogate's features are those binary indicators, even though the black box consumes a dense encoding underneath. Coefficients are therefore reported per word or per region.
  • How do you decide how many terms the local explanation should contain?
    Pick the budget for the reader, not the maths — typically five to ten terms, because that is what a moderator or an appeals reviewer can hold in mind. Then check how well the sparse surrogate reproduces the model's outputs on the weighted neighbourhood. If a five-term fit tracks the model poorly there, either widen the budget or admit the region is not linearly describable.

It is like pressing a flat sheet of glass against one spot on a curved sculpture. The sheet describes that spot well and tells you nothing about the far side.

saying these in an interview costs you the question

  • Says the surrogate approximates the model everywhere
  • Trains the surrogate on ground-truth labels
  • Reads the coefficients as real-world causal effects
  • Applies one row's explanation to every other row
  • Suggests deploying the surrogate in place of the black box
  • Forgets the perturbations must be scored by the model

context