When may a shallow tree fitted to a black-box model's predictions be quoted as its explanation?
answer
- what is the surrogate's target?
- imitate the model, not the world
- agreement measured on held-out inputs
- one global number hides bad segments
- threshold declared before you look
basics
~20 sA global surrogate is trained on the black box's own outputs, so it may be quoted only when it reproduces that model closely on held-out inputs from the deployment population, and closely within each segment you discuss.
solid answer
~50 sA global surrogate is an interpretable model fitted to the black box's **predictions**, not to the ground-truth labels, so it imitates the model rather than the world. Its only credential is **fidelity**: how well it reproduces the black box on inputs it was not fitted on. For a continuous output like an insurance premium that is R-squared against the black-box prediction; for a classifier it is agreement rate or rank agreement against its scores. Fix the threshold before you look, and report the number every time the tree is shown, because a tree quoted without its fidelity reads as if it were the model. Then check fidelity per segment, since a strong global figure routinely hides a region where the two disagree badly. And if the surrogate is both faithful and nearly as accurate, that is an argument for shipping the surrogate instead.
go deeper
Recall the basic recipe: train a small readable model on the complex model's own predictions, then check how closely the two agree. Knowing that the target is the predictions and not the labels is the key fact here.
Explain how fidelity is measured for a continuous and for a categorical output, and why a surrogate fitted on the true labels is a separate model rather than an explanation of the first one.
Show the operating discipline: a fidelity bar set in advance, fidelity reported per segment and alongside every display of the tree, stability checks across refits, and a refusal to make per-row claims from a global average.
Own the governance angle: what fidelity level your organisation requires before a surrogate may be quoted externally, who signs that off, and the standing rule that a highly faithful, nearly-as-accurate surrogate is a reason to replace the black box, not to keep explaining it.
## What a global surrogate is You have a model too complex to read — say a large ensemble that prices insurance premiums — and a committee that wants to know what drives the price. A global surrogate is the standard move: fit a small, readable model (a depth-3 or depth-4 tree, a sparse linear model, a short rule list) whose **target is the black box's own output**. The recipe: 1. Take a set of inputs that reflects the population the model actually scores — ideally recent production traffic, not the original training set. 2. Score them with the black box. Those scores become the labels. 3. Fit the interpretable model to `(inputs, black-box scores)`. 4. Measure how well it reproduces the black box **on inputs held out from step 3**. Note what step 2 changes. The surrogate is not trying to be accurate about reality; it is trying to be accurate about the model. Fitting the readable model to the true labels instead gives you a *different, independent model*, and any statement it makes is about that model, not about the black box. This is the single most common mistake in the whole technique. ## Fidelity is the credential **Fidelity** is agreement between surrogate and black box: - Continuous output: R-squared of the surrogate's output against the black box's output, or mean absolute deviation in the output's own units ("the tree is within 40 currency units of the model's premium on 90 percent of quotes"). - Classifier: proportion of inputs on which the two produce the same predicted class, and, if scores matter, correlation or rank agreement between the two score vectors. Rules for using it honestly: - **Pick the bar before you look.** Deciding after the fact that 0.72 is "good enough" is how a misleading tree gets into a slide deck. Set the bar from consequence: a tree used to *explain a decision to the person it affected* needs a much higher bar than one used internally to form a hypothesis. - **Publish it with the picture.** Any time the tree is displayed, the fidelity number and the population it was measured on travel with it. Without that, an audience reads the tree as the model. - **Segment it.** A global R-squared of 0.93 can decompose into 0.96 on the 90 percent of ordinary cases and 0.4 on the high-premium tail — and the tail is precisely what people ask about. Report fidelity per segment: by geography, by product, by score decile. A surrogate that is unfaithful exactly where the decisions are consequential is worse than no surrogate. - **Never promise anything about an individual row.** High global fidelity is an average statement. It gives no guarantee that the tree's path for one particular quote matches why the black box priced that quote the way it did. If someone needs a per-decision account, a global surrogate is the wrong tool. ## The instability problem Surrogate trees are not unique. Refit on a different sample of inputs, or with a different random seed, and the tree can restructure completely while achieving essentially the same fidelity — a different feature at the root, different thresholds. Before quoting one, refit it several times on resampled inputs and look at whether the top splits are stable. If they are not, you may report the features that recur but must not present specific thresholds as facts about the model. Related: the surrogate inherits the black box's behaviour only where you have data. In regions the input sample does not cover, the tree extrapolates on its own terms and says nothing trustworthy about the model. ## The awkward question the technique raises If a depth-4 tree reproduces the black box with R-squared 0.95, then most of what the ensemble does is capturable by four splits. Measure that tree's own accuracy against real outcomes. If it is close to the black box's accuracy, the honest recommendation is to **ship the tree** — you then get faithfulness for free instead of paying for a surrogate and hoping. Interviewers value candidates who reach that conclusion unprompted, because it shows they are treating the surrogate as a diagnostic rather than as compliance theatre. Conversely, if fidelity stays low no matter how you fit it, that is real information: the black box's behaviour is genuinely not summarisable in a few readable rules, and the right response is to say so and use per-decision methods, tighter model constraints, or a different model — not to show the low-fidelity tree anyway. ## Interview framing Say "fitted to the model's predictions, not the labels" in your first sentence — it is the discriminating detail. Then give fidelity a name, a metric, a pre-declared threshold and a per-segment breakdown, and close with the two honest outcomes: high fidelity is an argument for replacing the black box, and low fidelity means you say nothing rather than showing the tree.
- What is the difference between the surrogate's fidelity and its accuracy?Fidelity is agreement with the black box's outputs; accuracy is agreement with the real outcomes. A surrogate can be highly faithful to a model that is itself wrong, and it can be reasonably accurate while describing the black box badly. Fidelity is what licenses you to speak on the model's behalf; accuracy is what would license shipping the surrogate instead.
- Your surrogate tree reaches R-squared 0.93 globally but 0.45 on the highest-premium decile — what do you do?Do not present it as an explanation for expensive quotes, which is where the questions come from. Report the split explicitly, then either fit a separate surrogate for that segment, restrict the claim to the segment where fidelity holds, or fall back to per-decision methods there. Quoting the global number alone would be misleading.
- Why fit the surrogate on recent production traffic rather than the original training set?Because the explanation should describe the model where it is actually used. If the input mix has shifted, a surrogate fitted on old training data will be faithful in regions that no longer matter and unfaithful where traffic now concentrates. Fidelity is only meaningful relative to a stated population.
- How do you check that a surrogate tree's structure is trustworthy?Refit it several times on resampled inputs or different seeds and compare. Surrogate trees are notoriously unstable: fidelity can hold while the root feature and the thresholds change completely. If only the feature set is stable, report the features and drop the specific cut points from your claims.
A surrogate is a translator. Before you quote the translation in a contract, you check how accurately it renders the original — and you check it on the difficult passages, not just the easy ones.
saying these in an interview costs you the question
- Fits the surrogate to the true labels instead of the model's predictions
- Shows the surrogate tree without any fidelity number
- Reports one global fidelity figure and never checks segments
- Claims the surrogate explains individual predictions
- Chooses the acceptable fidelity threshold after seeing the result
- Treats an unstable surrogate's thresholds as facts about the model