Why does a raw input-gradient saliency map for an image classifier look like speckle noise?
answer
- it is a slope, not a contribution
- piecewise-linear surface, tiny linear regions
- saturation kills the derivative
- average over noise-perturbed copies
- cleaner picture is not a truer one
basics
~20 sThe gradient of a class score with respect to the pixels is a purely local slope of a jagged, piecewise-linear function, so it flickers pixel to pixel. Averaging over many noise-perturbed copies of the input makes it readable.
solid answer
~40 sA saliency map is `d(class score)/d(input pixel)`, the first-order term of a linear approximation valid only in a tiny neighbourhood of that one image. A deep rectified network is piecewise linear with an enormous number of pieces, so its surface is continuous but extremely jagged: nudge a pixel imperceptibly and you cross into a different piece with a different slope. Sign and magnitude therefore flicker from pixel to pixel. Saturation makes it worse, since a decisive feature can have a near-zero derivative once the score is confidently high. The standard remedy is SmoothGrad: add Gaussian noise to the input, recompute, and average over a few dozen samples, estimating a locally smoothed gradient rather than one arbitrary point sample. Absolute values, percentile clipping, or a feature-map-level or perturbation method are the alternatives.
go deeper
Know what the map is — the derivative of a class score with respect to each input pixel — and that raw versions look grainy, so what you see displayed has usually been smoothed or percentile-clipped first.
Explain the cause: a rectified network is piecewise linear with tiny regions, so the local slope flickers, and saturation can zero the derivative of a decisive pixel. Then describe averaging over noise-perturbed inputs as the fix.
Demonstrate that you verify rather than admire: cross-check regions across independent attribution methods, mask the highlighted area and measure the probability drop, and confirm the map is aligned with the model's preprocessed input space.
Own where these maps are allowed to appear at all. Set the standard for what evidence must accompany a published or customer-facing heatmap, and push back on smoothing knobs being tuned until the picture matches the story someone wants to tell.
## What the map actually is Input-gradient saliency, the oldest gradient attribution map, is just the partial derivative of a chosen class's score with respect to each input value, displayed as an image (usually as absolute value, or as the maximum absolute value across colour channels). The claim it encodes is narrow and first-order: *if this pixel were nudged infinitesimally, the class score would change at this rate*. ## Why the map is noisy **It is a point estimate of a rapidly varying quantity.** A network built from rectified linear units is a piecewise-linear function of its input. Each unit's on/off pattern defines a linear region, and the number of such regions is astronomically large, so regions are tiny. Inside a region the gradient is exactly constant; step across a boundary and it changes discontinuously. The gradient at one specific image is therefore a sample of a function that varies on a scale far smaller than any visually meaningful perturbation. Two images a human cannot tell apart can produce visibly different saliency maps. **Local slope is not the same as contribution.** A pixel that is essential to the prediction can have a near-zero derivative because the score has saturated: if the evidence is already overwhelming, marginal changes do not move the output. Conversely, a pixel that contributes nothing can sit on a steep local slope. The gradient measures sensitivity in an infinitesimal neighbourhood, not how much of the decision that pixel is responsible for. **Display choices amplify it.** The distribution of gradient magnitudes is extremely heavy-tailed. A handful of pixels dominate the scale, so a linear colour map shows a few bright specks on black. Practitioners clip at a high percentile — say the 99th — before normalising, which is a presentation fix, not a faithfulness fix, and it can make an unstable map look decisive. ## Making it readable **Average over noisy copies (SmoothGrad).** Draw `n` samples `x + e` with `e` Gaussian, standard deviation set as a fraction of the input's value range (a common working range is roughly ten to twenty percent of the range), compute the saliency map for each, and average. This approximates the gradient of a locally smoothed version of the score function, replacing an arbitrary point sample with a neighbourhood estimate. Visual quality improves quickly for the first few dozen samples and then plateaus. The two knobs interact: too little noise and nothing changes; too much and you are explaining a different, blurrier image. **Aggregate along a path instead of at a point.** Integrating the gradient along a path from a reference input to the actual input replaces the single local slope with an accumulation, which sidesteps the saturation problem — that is what integrated gradients does, at the cost of having to choose a reference. **Move up a level of abstraction.** Attributing at a convolutional layer's feature maps rather than at pixels produces a coarse but stable map, because the spatial averaging in the construction throws away exactly the high-frequency component that makes pixel maps unreadable. **Stop using gradients.** Occlusion attribution slides a patch of neutral value across the input and records how far the predicted class probability falls at each position. It is a finite, human-scale perturbation rather than an infinitesimal one, so it answers a question you can actually verify, and it is model-agnostic. The price is one forward pass per patch position and a patch-size tradeoff: small patches give fine maps but a robust model shrugs them off; large patches give smooth maps that cannot localise anything and may cover two objects at once. ## The judgment part The most common failure here is treating visual cleanliness as evidence of correctness. Smoothing is a variance-reduction trick applied to an estimator whose bias you have not measured; a cleaner picture can be a *more* confident lie. Two disciplines protect you. First, agreement: if a smoothed gradient map, an integrated-gradient map and an occlusion map all point at the same region, the region is probably real; if they disagree, believe none of them yet. Second, verification by intervention: mask the highlighted region and measure how much the class probability actually drops. That converts an attribution picture into a measurement. One last practical note: because the map is a derivative with respect to the *input as the model sees it*, it lives in the model's preprocessed input space. If the display pipeline resizes, crops or normalises differently from the training pipeline, the map you are staring at is not aligned with the image you think it explains — a surprisingly common source of maps that look nonsensical for a completely mundane reason.
- How many noisy samples and how much noise would you use when averaging saliency maps?Noise standard deviation of roughly ten to twenty percent of the input's value range is the usual working band, and visual quality typically stops improving after a few dozen samples. Both are tunable by inspection: too little noise leaves the speckle, too much explains a blurrier image than the one you have. Report the settings, because the map changes with them.
- When would you reach for occlusion instead of a gradient-based map?When you want a claim you can verify rather than an infinitesimal sensitivity: sliding a neutral patch and recording the drop in predicted class probability measures the effect of an intervention you could actually perform. It is also model-agnostic and immune to saturation. The costs are one forward pass per position and a patch size that trades resolution against the model's robustness to small occlusions.
- A colleague says the smoothed map is obviously more trustworthy because it looks like the object — how do you respond?Looking like the object is not evidence of faithfulness; edge-like structure appears in maps even when they barely depend on what the model learned. Trust needs a measurement: check that independent methods agree on the region, and mask the region and see how far the class probability actually falls. Visual plausibility is the thing most likely to fool you.
saying these in an interview costs you the question
- Says the noise means the training run is broken
- Treats a smoother map as a more faithful map
- Confuses local sensitivity with a pixel's contribution
- Blames the speckle on the image resolution alone
- Believes a near-zero gradient proves a pixel is irrelevant