skip to content

How does Grad-CAM turn a convolutional network's feature maps into a class-specific heatmap?

level: middleimportance: must knowfreq 58%

answer

  1. start at the last convolutional layer
  2. one number per channel, not per pixel
  3. spatially average the class-score gradients
  4. weighted sum of the maps, then rectify
  5. coarse grid upsampled to the image

basics

~20 s

Grad-CAM averages the gradients of one class score over each feature map of a chosen convolutional layer to get per-channel weights, sums the feature maps with those weights, and applies ReLU. The coarse result is upsampled onto the image.

solid answer

~50 s

Pick a convolutional layer, normally the last one before the classifier head, and call its output channels the feature maps `A^k`. For the class you want explained, take the gradient of that class's score with respect to every feature map, then spatially average each gradient map to a single number: `alpha_k = mean over positions of d(score)/d(A^k)`. That is one importance weight per channel, not per pixel. Form the weighted sum `sum_k alpha_k * A^k` and pass it through ReLU, which keeps only the regions pushing the score up and discards evidence against the class. The map lives on the spatial grid of that layer, so it is coarse and gets upsampled to the input size for display. Because the weights come from a specific class's gradient, two classes on the same image give two different maps.

go deeper

for a junior

Be ready to say what the picture means: a class-specific heatmap over the image built from a late convolutional layer, showing roughly where the evidence for that class came from, not which exact pixels mattered.

for a middle

You are expected to derive it out loud: gradients of the class score with respect to each feature map, spatially averaged into one weight per channel, weighted sum, ReLU, upsample. Know why the weights are per channel.

for a senior

Show judgment about which layer to explain and which score to differentiate, and treat a map as a hypothesis to confirm by masking or occlusion rather than as a finding you report straight to a stakeholder.

for a principal

Own the policy question: what an attribution map is allowed to justify in a review or an audit, when a coarse map is not enough resolution to support the claim being made, and what evidence has to accompany it.

## What the method is for A convolutional classifier gives you a number per class and nothing else. Grad-CAM (gradient-weighted class activation mapping) answers a narrower question than "why this prediction": it answers **where in the image the evidence for this class was pooled from**. It produces a coarse heatmap over the input, one map per class you ask about. ## The construction, step by step 1. **Choose a layer.** Take a convolutional layer's output activation tensor. Its channels are the feature maps, written `A^k` for channel `k`, each a 2-D grid of spatial positions. The last convolutional layer before global pooling is the usual choice: it is the deepest place that still has spatial structure, so its channels are semantic ("fur texture", "striped pattern") and still localised. 2. **Choose a class and differentiate.** Let `y_c` be the score for the class you want explained. Compute `d y_c / d A^k_ij` for every channel `k` and position `(i, j)`. This is one backward pass, stopped at that layer rather than carried all the way to the input. 3. **Pool the gradients into weights.** Average each gradient map over its spatial positions: `alpha_k = (1/Z) * sum_ij d y_c / d A^k_ij`, where `Z` is the number of positions. `alpha_k` is a scalar per channel that says how much raising channel `k` uniformly would raise the class score. This global average pooling of gradients is the step people most often misremember — the weights are per channel, and all of the spatial information in the final map comes from the activations, not from the gradients. 4. **Combine and rectify.** `L_c = ReLU( sum_k alpha_k * A^k )`. Without the ReLU the map mixes positive and negative evidence; negative regions are typically evidence for *other* classes, and showing them makes the map hard to read. Grad-CAM is defined as the positive part. 5. **Upsample.** `L_c` has the spatial resolution of the chosen layer — a small grid, often around a sixteenth or a thirty-second of the input side. It is bilinearly resized to the input size and overlaid as a heatmap. ## Why it is class-discriminative Everything class-specific enters through `alpha_k`. Run the same procedure for a second class on the same image and the activations `A^k` are identical while the weights change, so the highlighted region moves. That is the property that makes Grad-CAM usable as an audit tool: if the map for the *predicted* class sits somewhere implausible, that is a signal about the model, not about the image. ## Which score to differentiate Use the pre-softmax score for the class. Differentiating the post-softmax probability couples the class to all the others — softmax is a competition, so the gradient also reflects pushing rival classes down, and on a confident, saturated prediction the probability's gradient is close to zero everywhere, washing the map out. ## What it cannot tell you - **Resolution.** The map cannot be sharper than the layer it was computed at. A blob covering a face does not mean "the eyes"; it means "somewhere in this receptive-field-sized region". - **Direction of causation.** The map says the score is sensitive to activity in that region under a local linear approximation. It does not prove that masking the region would flip the prediction — if you need that claim, occlude the region and measure the drop. - **Correctness.** A sensible-looking map is not evidence the model is right, and an implausible map is not automatically a bug in the model: it may be the model exploiting a real but unwanted cue in the data. - **Faithfulness by default.** Before drawing conclusions from any attribution map, check that the map actually depends on the model's learned parameters; some popular variants are much less sensitive to them than they appear. ## Choosing the layer Earlier layers give finer maps that are less class-specific — early channels respond to edges and colours that many classes share, so the weighted sum drifts toward a generic contrast map. Later layers give semantically meaningful but blocky maps. For an architecture with several downsampling stages, the end of the last stage is the default; if you need finer localisation, the honest options are a higher-resolution layer combined with a coarser one, or a perturbation method such as occlusion, not sharpening the picture in post-processing.

  • Why is a ReLU applied to the weighted combination instead of showing the signed map?
    The signed sum mixes regions that raise the target class's score with regions that lower it, and the negative ones are usually evidence for competing classes. Grad-CAM is defined as the positive part so the overlay reads as "support for this class". If you actually want the counter-evidence, compute the map for the rival class instead of un-rectifying this one.
  • What changes if you compute the map at an early convolutional layer instead of the last one?
    You get a much finer grid but a far less class-specific map. Early channels encode edges, colours and textures shared across many classes, so the class-weighted sum degenerates toward a generic contrast or edge map. Later layers trade resolution for semantics, which is why the last convolutional block is the default choice.
  • Two classes on the same image produce nearly identical Grad-CAM maps — what does that suggest?
    The two classes are being separated by channels whose spatial support overlaps almost completely, so the map cannot discriminate them — common for fine-grained classes that share an object and differ in texture. It can also mean the weights are dominated by a few channels active everywhere. Check the per-channel weights and try a perturbation method before concluding anything.

saying these in an interview costs you the question

  • Says the gradients give a weight per pixel rather than per channel
  • Thinks the map is the raw activations with no class involved
  • Uses the post-softmax probability and cannot explain the washed-out map
  • Reads a plausible-looking heatmap as proof the model is correct
  • Forgets the map is upsampled from a coarse spatial grid
  • Claims the map shows which pixels must change to flip the class

context