In a convolution layer, what are the gradients with respect to its input and its kernel weights?
answer
- same family of op as the forward
- rotate the kernel 180 degrees
- input correlated with upstream gradient
- stride becomes dilation of the gradient
- shapes must match x and w exactly
basics
~20 sA convolution's input gradient is itself a convolution: the upstream gradient against the kernel rotated 180 degrees. The kernel gradient correlates the layer input with the upstream gradient, summed over every position the kernel visited.
solid answer
~40 sBoth gradients are convolution-shaped operations, not matrix products. For the input, each position collects a weighted sum of the upstream gradients of every output whose window covered it, weighted by the kernel tap that touched it — written over whole tensors that is the upstream gradient convolved with the kernel rotated 180 degrees, with output channels summed away and padding inverted (`valid` forward becomes `full` backward). For the kernel, `dL/dw[o,c,u,v]` is the correlation of the input channel with the upstream gradient over batch and space: because weight sharing reuses one weight at every position, it accumulates one contribution per position. Bias gradient is the upstream gradient summed over batch and both spatial axes. Stride `s` shows up as dilation: insert `s-1` zeros between the upstream gradient's entries, then run the stride-1 flipped convolution.
go deeper
Be ready to say that a convolution's backward pass produces two gradients — one for the input and one for the kernel — and that both are computed by sliding operations, not by a single matrix multiply.
You are expected to state the mechanics: the input gradient is a full convolution with the 180-degree-rotated kernel summed over output channels, the kernel gradient is a correlation of input with upstream gradient summed over batch and space, and the bias gradient is a plain sum.
Demonstrate the operational consequences: kernel gradients scale with feature-map area, so a resolution change silently rescales the effective step size; stride turns into dilation, and a kernel smaller than the stride leaves input positions with exactly zero gradient.
Own the cost model this implies — backward is roughly two forward-equivalents, which sets the arithmetic budget for training and drives choices about stride, resolution and where in a network to spend compute rather than parameters.
### What the layer computes A convolution layer holds a kernel `w` of shape `(out_channels, in_channels, k, k)` plus one bias per output channel. In deep learning the forward pass is a cross-correlation — no flip: ``` y[o, i, j] = b[o] + sum_c sum_u sum_v w[o, c, u, v] * x[c, i*s + u - p, j*s + v - p] ``` with stride `s` and zero padding `p`. Every output position reuses the same weights. That weight sharing is exactly what makes the backward pass look different from a dense layer's. ### Gradient with respect to the input Let `g = dL/dy` be the upstream gradient. Differentiating the sum above with respect to one input element and collecting terms gives: input position `(c, m, n)` receives a weighted sum of the upstream gradients of every output whose window covered it, each weighted by the kernel tap that touched it, summed over output channels. Written over whole tensors, that is a convolution again: ``` dL/dx = full_convolution(g, rot180(w)), summed over output channels (stride 1) ``` The 180-degree rotation appears because the kernel index runs one way in the forward pass and the opposite way when you ask which outputs saw a given input. Padding inverts too: a forward with padding `p` needs `k - 1 - p` padding per side in the backward, so a `valid` forward becomes a `full` backward and a `same` forward stays `same`. If instead you define the forward as a true mathematical convolution — which already flips — then the backward is the unflipped correlation. The flip always lives on exactly one side of the pair; it never lives on both and never on neither. The practical reading: the input gradient is the same *kind* of operation as the forward, over tensors of the same size, so it costs roughly the same arithmetic. ### Gradient with respect to the kernel ``` dL/dw[o, c, u, v] = sum over batch and (i, j) of x[c, i*s + u - p, j*s + v - p] * g[o, i, j] ``` This is a correlation of the layer input with the upstream gradient. The structurally important part is the sum. One kernel weight is applied at every spatial position, so it receives one contribution from every position, added together. A 3x3 kernel sliding over a large satellite tile gathers on the order of hundreds of thousands of contributions per image per output channel, where a dense-layer weight would gather one per sample. That is the whole bargain of convolution: a tiny parameter count that still receives a dense, high-count gradient signal. It also means conv weight gradients scale with feature-map area. Change the input resolution and the effective step size on those weights changes with it — a real cause of a learning rate that suddenly stops working after a resolution change. The bias gradient is `dL/db[o] = sum over batch and (i, j) of g[o, i, j]`, because the bias was broadcast to every position and a broadcast's gradient sums over the broadcast axes. ### Stride Stride does not change the shape of the rule, only the spacing. The clean way to see it: a stride-`s` forward is a stride-1 forward followed by keeping every `s`-th output. The backward of keeping every `s`-th entry is inserting `s - 1` zeros between entries. So the strided backward is: dilate `g` with `s - 1` zeros, then run the stride-1 flipped convolution. The kernel gradient becomes a dilated correlation in the same way. Two consequences are worth stating precisely, because candidates usually get them backwards: - If `k < s`, some input positions were never covered by any window at all. They receive exactly zero gradient. The layer never read them, and no training signal reaches them through this path. - If `k` is not an integer multiple of `s`, different input positions are covered by different numbers of windows, so gradient magnitude varies periodically across the input grid. That uneven-overlap tiling does not break training, but it is why gradient maps out of strided layers show a regular texture rather than a smooth field. ### Shape checks that catch almost every error Two checks catch nearly every hand-derivation mistake. First, `dL/dx` must have exactly `x`'s shape and `dL/dw` exactly `w`'s shape — if a rotation or a padding choice is wrong, one of them comes out the wrong size. Second, every index that appears in the forward but not in the gradient you are computing must be summed over: output channels disappear from `dL/dx` (summed), and batch plus both spatial axes disappear from `dL/dw` (summed). A missing sum is the single most common bug, and it is silent — the shapes may still broadcast into something plausible. Because both backward operations are convolution-shaped over same-sized tensors, a full backward costs roughly twice a forward, which is why convolutional training time is usually modelled as three forward-equivalents per step.
- How does a stride of 2 change the backward pass?Insert one zero between every pair of upstream-gradient entries, then run the ordinary stride-1 flipped convolution; the kernel gradient becomes a dilated correlation the same way. Coverage also changes: with a kernel smaller than the stride, some input positions are covered by no window and get exactly zero gradient, and when the kernel is not a multiple of the stride, gradient magnitude varies periodically across positions.
- Why does the backward pass cost roughly twice the forward pass here?Both the input gradient and the kernel gradient are convolution-shaped operations over tensors the same size as the forward's, so each costs about the same multiply-accumulate count as the forward. Two of them plus the forward is why a training step is usually budgeted at roughly three forward-equivalents.
- What is the bias gradient for a convolution layer with one bias per output channel?The upstream gradient for that channel summed over the batch and both spatial axes. The bias was broadcast to every spatial position in the forward, and the gradient of a broadcast is a sum over the axes it was broadcast along. Its shape is therefore just the number of output channels.
saying these in an interview costs you the question
- Applies the forward kernel again without the 180-degree rotation
- Thinks each kernel weight gets one gradient contribution per image
- Claims the kernel gradient is an outer product like a dense layer
- Forgets to sum over output channels when routing gradient to the input
- Assumes stride plays no role in the backward pass
- Never checks that the input gradient matches the input's shape