skip to content

An attribution map barely changes after a trained network's top layers are randomly reinitialised — what does that mean?

level: seniorimportance: nice to knowfreq 26%

answer

  1. plausible pictures are not evidence
  2. a falsification test, not a quality score
  3. randomise the learned weights layer by layer
  4. faithful maps change when the model does
  5. compare with rank correlation, not by eye

basics

~20 s

The map does not depend on what the model learned, so it cannot be explaining that model's decision — it is acting as an edge or contrast detector on the input. A map that fails this sanity check is not evidence.

solid answer

~50 s

This is the model-parameter randomization sanity check. Take the trained network, reinitialise the weights of the top layers at random — then progressively deeper — and recompute the attribution map for the same input and class. The prediction is now meaningless, so any faithful explanation of it must change dramatically. If the map stays essentially the same, whatever it is showing was computed from the input rather than from the learned parameters; several popular map variants that combine gradients with rectified backward passes are known to be largely invariant this way, which is why they look convincingly object-shaped on almost any network. The companion check is data randomization: retrain the same architecture on randomly permuted labels and confirm the map changes. Judge both with a quantitative similarity measure such as rank correlation between the maps, not by eye — human judgment of "these two heatmaps look different" is exactly what the check exists to replace.

go deeper

for a junior

Take away one habit: an attribution map that looks like the object is not automatically telling you about the model, so a heatmap alone is never proof of anything.

for a middle

Be able to describe the procedure — randomise the trained weights from the top down, recompute the map for the same input and original class, and compare — and say what result would discredit the method.

for a senior

Show you run this before trusting any map, quantify the comparison rather than eyeballing it, and pair it with cross-method agreement and a masking intervention that measures the actual probability drop.

for a principal

Own the bar for evidence. Decide which attribution methods your organisation is permitted to use in reviews, audits and model documentation, and require the sanity-check result to be recorded alongside any published explanation.

## The problem the check solves Attribution maps are evaluated by looking at them, and the thing a human finds convincing is a map that traces the object. But tracing the object is also what a plain edge detector does, and an edge detector explains nothing about a model. So visual plausibility cannot distinguish a faithful explanation from an input-processing artefact. A sanity check is a *falsification test*: a condition an explanation method must pass to be worth interpreting at all, not a measure of how good it is. ## The model-parameter randomization test The logic is a controlled ablation on the model rather than on the input. 1. Compute the attribution map for a fixed input and class on the trained network. 2. Replace the learned weights of the topmost layer with fresh random values, leaving everything else trained. Recompute the map. 3. Continue downward, randomising successively more layers (the *cascading* variant), or randomise one layer at a time with the rest trained (the *independent* variant), recomputing the map at each stage. 4. Measure the similarity between each map and the original with a quantitative metric — rank correlation over pixel scores, structural similarity, or the overlap of the top-k attributed pixels. After randomising the top layers, the network's output for that class is arbitrary. An explanation of an arbitrary output must itself be arbitrary. A method whose maps stay highly correlated with the original has failed: it is a function primarily of the input, with the model's parameters entering only weakly. ## The data randomization test The complementary check attacks the labels instead. Train the same architecture on a copy of the dataset with the labels randomly permuted. Such a model can only memorise; it has learned no real input-label relationship. A faithful attribution method should produce maps on this model that look nothing like the maps from the properly trained model. If they match, the method is again reading the input, not the model. ## What tends to fail, and why Methods that modify the backward pass to suppress negative signals produce beautifully sharp, object-shaped pictures — and that sharpness is largely a property of how the network's early layers process edges, not of what the classifier learned. Empirically, plain input gradients and class-weighted feature-map methods such as Grad-CAM change substantially under parameter randomization, while guided-backpropagation-style maps and their products with a coarse class map are far less sensitive. The lesson generalises past the specific list: **the more a method's visual appeal comes from post-processing the backward signal, the more suspicious you should be**, and any new method should be put through the same test before it is adopted. ## Reading the result correctly - **Passing is necessary, not sufficient.** A map that changes when the weights change is at least a function of the model. That does not make it a correct account of the decision, and it does not license any causal claim. - **Failing does not make the picture useless — it makes it not an explanation.** A method that behaves like an edge detector may still be a fine visualisation of image structure. It simply must not be captioned "this is what the model looked at", shown to a reviewer, or used to justify a deployment decision. - **Judge numerically.** Two heatmaps can look different to the eye while being highly rank-correlated, and vice versa; the normalisation and colour map you use to display them can manufacture apparent differences. Report a similarity number per randomisation stage. - **Check the class you care about.** After randomisation the argmax class changes, so keep explaining the *originally predicted* class index throughout; otherwise you are comparing explanations of two different questions and the test is meaningless. ## Where this sits in a workflow Run the check once per attribution method per architecture family, not once per image — it is a property of the method-and-model pairing, and it is cheap relative to the cost of building a review process on top of an explanation that turns out to be an edge detector. Combine it with two cheap habits: cross-method agreement (do two independent methods point at the same region?) and intervention (mask the highlighted region and measure how far the class probability actually drops). A map that survives randomisation, agrees with a second method, and predicts a measurable drop under masking is worth acting on. One that only looks convincing is not.

  • What is the data randomization version of this check?
    Retrain the same architecture on a copy of the dataset with the labels randomly permuted, so the model can only memorise and has learned no genuine input-label relationship. Then compare its attribution maps with those from the properly trained model. A faithful method should give visibly and measurably different maps; if they agree, the method is reading the input rather than the model.
  • Why compare the maps with a similarity metric instead of just looking at them?
    Because the check exists precisely because human judgment of heatmaps is unreliable. Colour maps, normalisation and percentile clipping can manufacture or hide apparent differences, and two maps that look different can still be highly rank-correlated. Report a number — rank correlation of pixel scores, or top-k overlap — at each randomisation stage.
  • A method fails the check. Does that make its output worthless?
    It makes it not an explanation of that model. The picture may still be a reasonable visualisation of image structure, and nothing stops you looking at it. What you must not do is caption it as what the model attended to, put it in a model card or a review, or let it justify shipping. Swap to a method that passes and re-derive any conclusion drawn from it.

saying these in an interview costs you the question

  • Judges attribution quality by how object-shaped the map looks
  • Thinks passing the check proves the explanation is correct
  • Believes randomising weights should leave a good map stable
  • Compares maps by eye instead of a similarity metric
  • Switches to the new argmax class after randomising the weights
  • Assumes every published attribution method is faithful by default

context