skip to content

Why does VGG stack two 3x3 convolutions instead of using one 5x5 layer?

level: middleimportance: must knowfreq 70%

answer

  1. same window, fewer weights
  2. receptive field adds k minus 1
  3. 18 against 25 C-squared
  4. an extra nonlinearity between the layers

basics

~10 s

Two stacked 3x3 layers see the same 5x5 input patch as one 5x5 layer, but use 18C-squared weights instead of 25C-squared and apply two nonlinearities instead of one. Same reach, cheaper, more expressive.

solid answer

~50 s

For stride-1 layers the receptive field grows by `k - 1` per layer, so two 3x3 layers reach exactly the 5x5 patch that a single 5x5 layer reaches, and three 3x3 layers reach 7x7. At `C` input and `C` output channels the pair holds `2 * 9 * C^2 = 18C^2` weights against `25C^2` for the 5x5 layer, about 28% fewer; at the 7x7 field it is `27C^2` against `49C^2`, about 45% fewer. Because both run at the same spatial size, multiply-adds fall by the same ratio. The second win is nonlinearity: the pair applies a linear map, a nonlinearity, another linear map and another nonlinearity over that window, so it is a richer function class per unit of receptive field. That is VGG's whole design rule — 3x3, stride 1, everywhere — and it is why the family goes deep instead of wide-kerneled.

go deeper

for a junior

Recall the design rule and the headline fact: VGG uses 3x3 kernels at stride 1 everywhere, and two of them stacked cover the same 5x5 input patch as one 5x5 layer while using fewer weights.

for a middle

Derive it live. Receptive field grows by k minus 1 per stride-1 layer, and the weight counts are 2 times 9C-squared against 25C-squared. State clearly that the pair also inserts a second nonlinearity over the same window.

for a senior

Show where the saving fails to appear: activation memory during training and wall-clock latency, since two dependent layers at high resolution are memory-bound. Note that multiply-adds fall by the same ratio as parameters only when both run at equal spatial size.

for a principal

Own the limits of the rule. Receptive field grows only linearly with depth, trainable depth is bounded, and a uniform kernel rule is a default worth breaking when the binding budget is serving latency rather than parameter count.

## The two designs being compared A convolution layer with kernel size `k`, `C` input channels and `C` output channels holds `k*k*C*C` weights plus `C` biases, and applying it to an `H x W` feature map costs roughly `k*k*C*C*H*W` multiply-adds. VGG's design rule is that `k` is always 3 with stride 1 and padding 1, so the spatial size never changes inside a block and downsampling happens only at the pooling layers between blocks. Earlier convolutional designs mixed kernel sizes freely: LeNet-5, working on 32x32 grayscale handwritten-digit images, used 5x5 kernels in both of its convolution layers. ## Receptive field arithmetic The receptive field is the region of the input that can influence a single output unit. For stride-1 layers it grows additively: `r <- r + (k - 1)`. Starting from 1, one 3x3 layer gives 3, a second gives 5, a third gives 7. So two stacked 3x3 layers see exactly the 5x5 input patch a single 5x5 layer sees, and three of them reproduce a 7x7 field. Nothing about coverage is given up. ## Parameter and compute arithmetic At `C -> C` channels the comparison is: - one 5x5 layer: `25C^2` weights; two 3x3 layers: `18C^2` — about 28% fewer. - one 7x7 layer: `49C^2` weights; three 3x3 layers: `27C^2` — about 45% fewer. Because every layer in both stacks runs at the same spatial resolution, the multiply-add ratio equals the parameter ratio. The gap widens with kernel size for a simple reason: parameters grow as `k^2` while reach grows as `k`, so a big kernel pays quadratically for a linear amount of coverage. ## The nonlinearity argument The 5x5 layer applies one linear map over the window followed by one nonlinearity. The 3x3 pair applies a linear map, a nonlinearity, a second linear map and a second nonlinearity. Same reach, twice the nonlinear stages. A function that a single linear-plus-nonlinearity stage over the 5x5 window cannot represent may be easy for the two-stage composition. This is the half of the argument candidates most often forget, and it matters more than the parameter saving. ## The honest caveat about expressiveness Composing two 3x3 linear convolutions does yield a 5x5 kernel, but not an arbitrary one. `18C^2` free parameters cannot span the `25C^2`-dimensional space of 5x5 kernels, so the factorization is a restriction of the function class, not a generalization of it. The empirical claim VGG makes is that the restriction is benign and the extra nonlinearity more than pays for it. Saying this out loud in an interview shows you understand that matched receptive field is not the same as matched expressive power. ## What the swap does not buy Training memory. The pair produces two activation maps that must both be retained for the backward pass instead of one. At high resolution, activations, not weights, dominate training memory, so the two-layer version usually costs more memory during training even though the saved model is smaller. Latency. Two layers are two sequential, dependent steps. At high spatial size and modest channel counts each one is bound by memory traffic rather than arithmetic, so measured wall-clock time can be worse than the multiply-add ratio suggests. ## Where the argument stops paying Two limits. First, a stride-1 3x3 stack grows the receptive field only linearly — two pixels per layer — so covering a large fraction of the image purely by stacking takes many layers, whereas downsampling grows the field geometrically for free. Second, the depth you can actually optimize is bounded: as plain stacks got deeper, training difficulty, not capacity, became the binding constraint, and later architectural work was aimed squarely at that. The 3x3-stacking rule is an excellent default for the layers inside a stage; it is not a licence to keep adding layers indefinitely.

  • Does the two-layer stack save training memory as well as parameters?
    No. Weights fall from 25C-squared to 18C-squared, but the pair produces two activation maps that both have to be kept for the backward pass instead of one. At high resolution the activations, not the weights, dominate training memory, so the stack typically costs more memory to train even though the saved model is smaller.
  • Can two stacked 3x3 layers express any 5x5 kernel?
    No. Composing two 3x3 linear convolutions produces a 5x5 kernel, but only from a restricted family — 18C-squared parameters cannot span the 25C-squared-dimensional space of 5x5 kernels. The factorization narrows the function class; the nonlinearity inserted between the layers is what makes it a good trade rather than a strict downgrade.
  • How many 3x3 layers match a 7x7 receptive field, and what does that cost?
    Three, because a stride-1 layer adds k minus 1, which is 2, each time. Three 3x3 layers hold 27C-squared weights against 49C-squared for the single 7x7 layer, roughly 45% fewer, with three nonlinearities instead of one. The saving grows with kernel size because parameters scale as k squared while reach scales as k.

Two 3x3 layers are two small windows in a row: each unit peers through its little window at a map whose entries already peered through one, so the pair takes in a 5x5 patch of the original image.

saying these in an interview costs you the question

  • Says the stacked pair only sees a 3x3 region
  • Adds the kernel sizes and claims a 6x6 receptive field
  • Claims the pair is exactly equivalent to any 5x5 kernel
  • Mentions only parameters and forgets the extra nonlinearity
  • Assumes fewer parameters means less training memory

context