In the backward pass, how does a max-pooling layer route gradients compared with average pooling?
answer
- one window, one winner
- argmax indices saved in the forward
- three of four get exactly zero
- average splits by one over k squared
- overlapping windows sum their shares
basics
~20 sMax pooling sends a window's entire upstream gradient to the one input that was the maximum and exactly zero to the others. Average pooling spreads it evenly, so each input in a k-by-k window receives one k-squared-th of it.
solid answer
~40 sNeither layer has parameters, so their whole role in the backward pass is routing. Max pooling is a data-dependent router: the forward pass records which input won each window, and the backward hands that winner the full upstream gradient for the window while every loser receives exactly zero. In a 2x2 window that means three of four positions get nothing on this step. Average pooling is a fixed router: every input in the window gets an equal share, `upstream / k^2`, and no forward state is needed beyond the shapes. Two details matter. With overlapping windows, a position that wins several windows accumulates the sum of their gradients, because backprop sums over paths. And ties are broken by picking one winner, not by splitting the gradient across the tied inputs.
go deeper
Recall the two routing rules and be able to state them on a concrete 2x2 window: the maximum takes the whole upstream gradient and the rest take zero, while an average pool gives each of the four a quarter.
Explain why the routing differs — the derivative of a max is one at the argmax and zero elsewhere, the derivative of a mean is a uniform one-over-count — and why max pooling forces the forward pass to store argmax indices.
Show the operational consequences: sparse selective gradient versus diffuse low-variance gradient, the accumulation that overlapping windows create, and the fact that pooling changes the shape of the gradient field without contributing any parameter update.
Frame it as an architectural choice about where credit assignment goes — whether a block should sharpen one detector or spread signal across a region — and be able to argue that case against strided-convolution alternatives on a compute and stability budget.
### Pooling layers have no weights, only a routing rule A pooling layer produces one output per window and has nothing to learn. Its entire contribution to training is therefore *where it sends gradient*, and the two common variants send it to opposite extremes. ### Max pooling: winner takes everything For a window `W`, the forward computes `y = max over W of x`. The derivative of a max with respect to its inputs is 1 at the argmax and 0 everywhere else, so ``` dL/dx[m] = dL/dy if m is the argmax of the window dL/dx[m] = 0 otherwise ``` Concretely, in a video-frame classifier with a 2x2 max-pool, if the four inputs are `1, 7, 3, 2` and the upstream gradient is `4`, the position holding `7` receives `4` and the other three receive exactly `0` — not a small number, not a rounded-down number, zero. Three of every four positions get nothing from this path on this step. Two practical points follow. **The routing is data-dependent, so the forward must record it.** Which input wins is a function of the data, not the geometry, so the backward pass needs the argmax index for every window — often called the pooling switches. That is real stored state proportional to the number of output positions. Average pooling needs none: its routing is fixed by the window layout, so it can be reconstructed from shapes alone. **Ties go to one winner.** When two inputs are exactly equal, the subgradient is not unique; implementations pick one index (conventionally the first) and give it the whole gradient rather than splitting. It is a legitimate subgradient choice, and it means exact ties do not produce a symmetric update. ### Average pooling: everyone gets a share For a `k x k` average pool, `y = (1/k^2) * sum over W of x`, so every input in the window receives `dL/dy / k^2`. The gradient is dense, uniform, and independent of the input values. Global average pooling is the same rule with the window covering the whole feature map: each of the `H*W` positions of a channel receives `1/(H*W)` of that channel's gradient, which is why it spreads a classifier head's signal thinly and evenly over the entire spatial extent. ### Overlapping windows accumulate With a pooling stride smaller than the window (for example, 3x3 windows at stride 2), one input position sits inside several windows. In the computation graph that input has several outgoing edges, and backprop sums over all paths from a node to the loss. So a position that is the argmax of two windows receives the *sum* of both windows' upstream gradients, and under average pooling a position covered by four windows receives four shares. Forgetting this sum is a common hand-derivation bug and shows up as a gradient check that fails only when the pooling stride is less than the pooling size. ### What the sparsity does and does not mean Max pooling's zeros are local to this path, not global. A position that lost its window can still receive gradient through any other branch of the graph that reads it — a parallel path, a residual route, or an earlier layer's other consumers. It is also not the case that the losing units are permanently starved: the argmax is recomputed every forward pass, so as weights move, different positions win on different batches and over many steps most positions receive signal at some point. Still, the contrast is real and it is the reason people reach for one or the other in the backward-pass sense. Max pooling delivers a sharp, sparse, selective signal: only the feature that actually fired gets credit or blame, which sharpens whatever detector produced the maximum. Average pooling delivers a diffuse signal: every position is nudged a little, which keeps weak activations alive and produces smoother, lower-variance gradients but never singles out the responsible unit. ### Neither layer has a weight gradient Because there is nothing learnable, there is no `dL/dw` to compute for a pooling layer at all. If someone claims a pooling layer's weights are being updated, the mental model is wrong: a pooling layer changes the *shape* of the gradient field flowing backward, and that is its only effect on learning.
- With overlapping pooling windows, what happens to a position that is the maximum of two windows?It receives the sum of both windows' upstream gradients. In the graph that node has two outgoing edges, and backprop sums the contributions of every path from a node to the loss. Under average pooling the same rule applies: a position covered by four windows collects four shares.
- Why must the forward pass store the argmax positions for max pooling?Its routing is data-dependent — which input wins depends on the values, not the geometry — so the backward cannot reconstruct it from shapes. The alternative is re-running the forward comparisons. Average pooling needs no such state because its routing is fixed by the window layout.
- Does a max-pooling layer have a weight gradient?No. It has no learnable parameters, so there is nothing to update. Its only effect on training is where it sends the upstream gradient, which is why the argmax routing rule is the whole story for this layer.
- What does a global average pool send back to each spatial position of a channel?One over H times W of that channel's upstream gradient, identically for every position. It is a mean reduction over the spatial axes, so the backward is a uniform broadcast of the scaled gradient back across the whole feature map.
saying these in an interview costs you the question
- Says max pooling splits the gradient evenly across the window
- Thinks a pooling layer has weights that get updated
- Claims non-maximal inputs receive a small nonzero gradient
- Forgets overlapping windows accumulate gradient at a shared winner
- Says an exact tie splits the gradient across the tied inputs
- Assumes the backward needs no state from the forward pass