How does the neighbourhood kernel width change a local surrogate's explanation of the same prediction?
answer
- how far out do neighbours reach
- no ground truth for this setting
- wide drifts toward global behaviour
- narrow leaves few effective points
- fidelity rises as neighbourhood shrinks
basics
~20 sThe kernel width sets how local the explanation is. Too wide, and the surrogate drifts toward the model's average behaviour; too narrow, and nearly all weight falls on a few near-identical points, so a handful of samples decide the answer.
solid answer
~50 sThe width is a free parameter with no ground truth, and it silently decides what 'local' means. A wide kernel gives distant perturbations real weight, so the fit starts describing the model's general tendencies rather than what tipped this row — for a removed post you recover the model's generally toxic vocabulary instead of the two terms that pushed this one over. A narrow kernel shrinks the effective sample: most weight sits on a few almost-identical variants that all get the same predicted label, leaving little variation to regress on, so the surviving coefficients rest on very few points. Worse, narrowing the kernel makes the weighted fit quality look better almost automatically, so fidelity numbers cannot be used to pick the width. In practice: fix a width by an explicit rule, report it with the explanation, and only publish conclusions that survive a sweep across widths.
go deeper
Remember that a local explanation is built from artificial neighbours of one row, and that how far out those neighbours reach is a number a person chose. Knowing the setting exists, and that it changes the answer, is enough at this level.
Be ready to explain both extremes mechanically: wide weights distant points so the fit describes general behaviour, narrow collapses the effective sample so a few near-identical points carry the coefficients. Mention that the weights come from a distance-decaying kernel.
Show the operating discipline: fix the width by a stated rule, record it with every explanation, sweep it and publish only conclusions that survive. Say out loud that local fit quality cannot select the width because it rises as the region shrinks.
Own the standard your organisation applies before an explanation is shown to a user, a reviewer or a regulator — what metadata is retained, what stability across widths is required, and what happens to a case whose explanation is unstable. That policy matters more than any single method choice.
## What the width actually is A local surrogate weights each perturbation by its distance from the row being explained, typically with an exponential kernel: `w(z) = exp(-d(x, z)^2 / sigma^2)`, where `d` is a distance in the readable representation (for text, often how many words were deleted) and `sigma` is the **kernel width**. Large `sigma` makes the weights decay slowly, so far-away variants still count; small `sigma` makes them decay fast, so only near-identical variants count. Nothing in the data tells you what `sigma` should be. It is a modelling choice, and it is the choice the explanation is most sensitive to. ## The wide end Push `sigma` up far enough and every perturbation carries roughly equal weight. The fit is then no longer local at all; it is a sparse linear approximation of the model over a broad swathe of input space. Applied to the removed forum post, the top terms drift toward the vocabulary the model treats as risky *in general* — slurs, spam markers, common abuse patterns — even if none of them is what pushed this particular post across the line. The explanation reads plausibly, which is exactly why it is dangerous: an appeals reviewer cannot tell from the output that the answer describes the model's average behaviour rather than this decision. ## The narrow end Push `sigma` down and the **effective sample size** collapses. Formally the effective size is roughly `(sum w_i)^2 / sum w_i^2`; when almost all weight sits on twenty of your five thousand variants, you are fitting a linear model to twenty points. Two failure modes follow. First, variance: change the perturbation draw and the surviving terms change, because so few points carry the fit. Second, degeneracy: variants that differ from the original by one deleted word usually receive the *same* predicted class, so the surrogate has almost no output variation to explain and the coefficients shrink toward noise around a constant. ## Why fit quality cannot choose the width for you The tempting fix is to select `sigma` by maximising how well the surrogate reproduces the model on the weighted neighbourhood — a weighted R-squared, say. This does not work, because that quantity is not comparable across widths. As `sigma` shrinks, the weighted evaluation set concentrates on points where the model's output barely varies, and almost any linear fit reproduces a nearly constant target well. Local fidelity therefore tends to rise as the neighbourhood shrinks, all the way down to the degenerate limit. Fidelity is a useful check *at a fixed width* — it tells you whether a five-term linear model can describe that region at all — but it is not a selection criterion for the width itself. ## The sparsity knob interacts with it Width and sparsity are not independent. A wide neighbourhood contains more of the model's structure, so a five-term surrogate has to compress more and fits worse; the same five terms in a narrow neighbourhood fit almost perfectly and mean much less. When an explanation looks crisp, ask which of the two knobs bought the crispness. ## What to do in practice 1. **Fix the width by a stated rule**, not by eyeballing the output until it looks sensible — that is fitting the explanation to your prior. 2. **Record it as metadata** alongside every explanation you store or show. An explanation without its neighbourhood definition is not reproducible and not auditable. 3. **Sweep and report stability.** Recompute the explanation over a range of widths spanning roughly an order of magnitude. Terms that stay in the top set across the range are the ones worth showing; terms that appear only at one width are artefacts of the setting. 4. **Escalate when the sweep disagrees.** If the top term at a narrow width is absent at a moderate one, the honest output is 'this decision is not stably explainable by a local linear model', not the prettier of the two answers. ## The interview point What a senior answer conveys is that a local explanation is not a measurement read off the model — it is the output of a procedure with a tuning parameter that nobody validated against a ground truth. The competent response is to treat the width as a declared assumption, hold it fixed across cases so explanations are comparable to each other, and test conclusions for survival across it.
- Why can't you pick the kernel width by maximising the surrogate's local fit quality?Because that quantity is not comparable across widths. A narrower neighbourhood concentrates on points where the model's output hardly varies, so almost any linear fit reproduces it well, and fit quality climbs as the region shrinks. Optimising it drives you to the degenerate limit. Use fit quality as a sanity check at a fixed width instead.
- How would you report a local explanation so a reviewer can judge it?Ship the neighbourhood definition with the answer: the perturbation scheme, the kernel width, the number of samples, the term budget, and how well the surrogate tracked the model in that region. Add which terms survived a width sweep. Without that metadata the explanation is neither reproducible nor auditable.
- What does it mean when the top terms change completely between a narrow and a wide neighbourhood?It means the model's behaviour near that row is not well described by one sparse linear rule at any scale you tried. The honest report is that the decision is not stably explainable this way, and you should reach for a different form of explanation or escalate the case, rather than publishing whichever answer reads better.
It is a zoom level on a map. Zoom out and you describe the whole city rather than this street corner; zoom in far enough and you are describing three paving stones.
saying these in an interview costs you the question
- Treats the kernel width as a fixed property of the method
- Tunes the width until the explanation looks reasonable
- Picks the width by maximising local fit quality
- Assumes a narrow neighbourhood is always more faithful
- Reports coefficients without the neighbourhood definition