Why can integrated gradients rank a prediction's pixels differently under a black baseline than a blurred one?
answer
- the explanation is always relative to something
- attribution scales with the input-minus-baseline difference
- equal to the baseline means exactly zero credit
- attributions sum to the prediction difference
- blurring keeps colour and layout, removes detail
basics
~20 sIntegrated gradients explains a prediction relative to a chosen baseline input, so the baseline defines the counterfactual. An all-black baseline gives zero credit to already-black pixels; a blurred baseline keeps low-frequency content, so only the added detail earns attribution.
solid answer
~40 sIntegrated gradients attributes to input `i` the quantity `(x_i - b_i)` times the average gradient along a straight path from baseline `b` to input `x`. Two consequences follow. The `(x_i - b_i)` factor is exact, so any value equal to its baseline value gets zero attribution: under an all-black baseline every genuinely black pixel is invisible however much the model relies on it. And completeness says the attributions sum to `F(x) - F(b)`, which makes the map an answer to "why this prediction rather than the one on the baseline" — change the baseline and you change the question. A blurred baseline keeps colour and layout, so credit concentrates on the detail blurring destroyed. Neither ranking is wrong; report the baseline, prefer one whose own prediction is uninformative, and check ranking stability across several.
go deeper
Know the shape of the method: it accumulates gradients along a path from a reference input to the real one, instead of taking a single slope at the real input, and the reference has to be chosen.
Explain the formula and the completeness property, and why a component whose value equals the baseline's value gets exactly zero attribution. Be able to say what the step count controls.
Show that you treat the baseline as the counterfactual being asked about: justify the choice, verify the baseline's own prediction is uninformative, check the completeness gap, and test whether rankings survive a change of baseline.
Own the convention. Decide what baseline a team uses by default and what must be recorded with any published attribution, so that two explanations of the same model are comparable and nobody quietly picks the baseline that flatters the story.
## The method in one line For a model output `F` and an input `x`, integrated gradients picks a baseline input `b` and defines the attribution to component `i` as `IG_i = (x_i - b_i) * average over a in [0,1] of dF(b + a*(x - b))/dx_i` In practice the average is a Riemann sum over `m` interpolation steps between the baseline and the input. It was designed to fix two defects of a single input gradient: saturation (a decisive feature can have zero local slope) and the fact that a point gradient is an arbitrary sample of a jagged surface. Accumulating the gradient along a path handles both. ## Where the baseline enters The baseline is not a technicality; it is half of the definition, in two places. **The multiplicative factor.** Attribution is proportional to `x_i - b_i`. If an input component equals its baseline value, its attribution is exactly zero by construction, regardless of the gradients. With an all-black baseline, a black pixel gets zero. For a model that keys on dark structures — text on a light page, a shadow, a dark lesion — the attribution map systematically blanks out exactly the evidence you care about. The same trap appears for tabular or spectrogram inputs whose "neutral" value coincides with a meaningful value. **The completeness property.** Summing the attributions gives `F(x) - F(b)`. Every explanation therefore accounts for a *difference in prediction between two inputs*. "Why did the model say pneumonia?" is not the question being answered; "why did it say pneumonia here and not on this black image?" is. That is a legitimate and useful question, but it is baseline-dependent by design, and the honest way to present an integrated-gradients map is with its baseline named. ## Black versus blurred, concretely - **All-black.** Cheap, deterministic, and conventional. Two problems: black is not information-free — a network may produce a confidently non-uniform output on a black image, so `F(b)` is not neutral — and dark evidence is zeroed out by the multiplicative factor. - **Blurred input.** The baseline retains the image's colours and coarse layout and destroys the fine structure. Now `x_i - b_i` is small wherever the image is locally flat and large at edges and texture, so attribution concentrates on high-frequency detail. This is the right counterfactual when the question is "what did the sharp detail add?" and the wrong one when the model is keying on large flat regions or on overall colour, which the baseline already contains and which therefore receive no credit. - **Noise, or an average over several baselines.** Sampling baselines from a distribution — random noise, or real inputs drawn from the data — and averaging the resulting attributions reduces the arbitrariness of any single choice, at proportionally more compute. ## How to use this in practice 1. **Choose a baseline that means something.** State the counterfactual you want in words first, then pick the baseline that encodes it. If you cannot describe the counterfactual, you are not ready to interpret the map. 2. **Check the baseline's own prediction.** If the model is confident about the baseline, `F(x) - F(b)` is a difference between two opinionated predictions and the map is harder to read. Prefer a baseline whose output is close to uninformative. 3. **Verify the approximation converged.** Completeness gives a free diagnostic: compare the sum of the attributions with `F(x) - F(b)`. A large gap means the Riemann sum has too few steps; increase them until the gap is small. This catches the single most common implementation error. 4. **Test ranking stability.** Recompute with two or three different baselines. Components that stay near the top under all of them are worth acting on; components that reorder are artefacts of the counterfactual, not findings. 5. **Do not over-read completeness.** It is an accounting identity — the parts sum to the whole — not a guarantee of causal truth or of agreement with human intuition. A method can satisfy it and still produce a map that barely depends on the model's learned parameters, which is why an attribution map should be sanity-checked before it is trusted. ## A common misreading Candidates often describe integrated gradients as "the gradient, but averaged, so it is more reliable". The averaging is along a specific path from a specific reference, and both matter. Two teams explaining the same prediction of the same model with different baselines will produce different top-ranked features and both will be correct about their own counterfactual. Disagreement is a signal to state the question more precisely, not evidence that one of them ran the method wrong.
- How do you check that you used enough interpolation steps?Use the completeness property as a convergence diagnostic: the attributions should sum to the model output at the input minus the output at the baseline. Compute both sides and look at the gap. A large gap means the Riemann sum is too coarse, so raise the step count until the gap is small relative to the prediction difference.
- What baseline would you choose for a non-image input such as an audio spectrogram?Pick something that encodes the counterfactual you actually mean: a silence or noise-floor spectrogram if the question is "what did the sound add over silence", or a time-averaged spectrum if the question is "what did this moment add over the ambient background". Then check the model's prediction on that baseline is uninformative, and confirm no meaningful bin coincides exactly with the baseline value.
- Does completeness mean the attributions are the true causal contributions?No. Completeness only says the numbers add up to the prediction difference between input and baseline — an accounting identity that many different attribution assignments could satisfy. It says nothing about whether removing a highly attributed region would change the prediction, and nothing about whether the map reflects what the model learned. Both need separate checks.
saying these in an interview costs you the question
- Describes the baseline as an arbitrary implementation detail
- Claims a black image is an information-free input
- Cannot explain why black pixels receive zero attribution
- Treats completeness as proof of causal correctness
- Compares two maps computed with different baselines as if equivalent
- Never checks the attribution sum against the prediction difference