skip to content

In VGG-16, which layers hold most of the 138 million parameters, and why?

level: juniorimportance: should knowfreq 45%

answer

  1. parameters and compute live in different places
  2. look right after the flatten
  3. 7x7x512 into 4096 units
  4. about 90% in three dense layers

basics

~20 s

About 90% sit in the three fully connected layers at the end. The first alone, mapping the flattened 7x7x512 feature map to 4096 units, holds roughly 103 million weights; all 13 convolution layers together hold under 15 million.

solid answer

~50 s

At 224x224 input, five pooling stages leave a 7x7x512 feature map, which flattens to 25088 values. The first fully connected layer maps that to 4096 units, so it alone holds `25088 * 4096`, about 103 million weights — roughly three quarters of the model. Add `4096 * 4096` and `4096 * 1000` for the next two and the head is about 124 million of 138 million, near 90%. The entire 13-layer convolution stack is under 15 million, because a convolution shares one small kernel across every spatial position while a dense layer gives every input-output pair its own weight. The ranking flips for compute: convolutions reuse each weight at every position, so they account for essentially all the multiply-adds. Parameter count and compute cost rank the layers in opposite orders, which is why a headline parameter number is a poor proxy for serving cost.

code

python · 17 lines
python
def dense(nin, nout):
    return nin * nout + nout

TOTAL_VGG16 = 138_357_544
CONV_STACK  = 14_714_688          # all 13 convolution layers together

for side in (224, 448):
    grid = side // 32             # five 2x2 poolings, each halving the map
    flat = grid * grid * 512      # flattened last feature map
    print(side, grid, flat, dense(flat, 4096))

fc1 = dense(25088, 4096)          # 7x7x512 flatten -> 4096
fc2 = dense(4096, 4096)
fc3 = dense(4096, 1000)
head = fc1 + fc2 + fc3
print(head, CONV_STACK, head + CONV_STACK == TOTAL_VGG16)
print(round(100 * head / TOTAL_VGG16, 1))

go deeper

for a junior

Know that in a classic convolutional classifier the dense layers after the flatten hold most of the weights, and that convolution kernels are shared across positions so their count does not depend on input size.

for a middle

Do the arithmetic out loud: 7 times 7 times 512 is 25088 inputs, times 4096 units is about 103 million weights, and the three dense layers together are close to 90% of 138 million.

for a senior

Connect the count to consequences you have hit in practice: model file size and load time, the input resolution the flatten locks in, and the fact that latency is set by the early convolutions rather than by the head.

for a principal

Argue about which budget actually binds — parameter count, activation memory or multiply-adds — and be ready to explain why a headline parameter number misleads people about serving cost and hardware choice.

## Where the shape comes from VGG-16 takes a 224x224x3 image through five stages of 3x3 convolutions, each stage ending in a 2x2 pooling that halves height and width. Five halvings turn 224 into `224 / 32 = 7`, and the last stage carries 512 channels, so the final feature map is 7x7x512. Flattening it gives a vector of `7 * 7 * 512 = 25088` values, which is fed to the classifier head: dense layers of 4096, 4096 and 1000 units. ## The counts A dense layer from `n_in` to `n_out` holds `n_in * n_out + n_out` parameters. So: - first dense: `25088 * 4096 + 4096 = 102,764,544` - second dense: `4096 * 4096 + 4096 = 16,781,312` - third dense: `4096 * 1000 + 1000 = 4,097,000` That head totals about 123.6 million. The 13 convolution layers together come to about 14.7 million, and 123.6 plus 14.7 is the familiar 138.4 million. So roughly 89% of the model is the classifier head, and roughly 74% is one single layer — the first dense one. ## Why convolution is so much cheaper in weights A convolution layer with kernel `k`, `C_in` input and `C_out` output channels holds `k*k*C_in*C_out` weights no matter how large the feature map is: the same kernel slides over every position. VGG's most expensive convolution layers, at 512 to 512 channels with 3x3 kernels, hold about 2.36 million weights each. A dense layer has no sharing at all — every input element gets its own weight to every output unit — so a wide flattened input produces an enormous matrix. ## The resolution trap Because the flatten width depends on the input size, the head's parameter count scales with input area while the convolution stack does not. Feed 448x448 instead of 224x224 and five poolings leave 14x14x512, a flatten of 100352, so the first dense layer jumps to about 411 million weights — roughly four times as many — pushing the model past 440 million parameters without adding a single new feature detector. That is a pure consequence of the head's design, not of the features being learned. The same dependence is why such a network is locked to one input resolution: the stored weight matrix has a fixed input dimension, and a different image size produces a flattened vector of the wrong width. The convolution stack itself is happily size-agnostic. ## Parameters versus compute A dense layer costs about one multiply-add per weight per example, so the whole head is on the order of 124 million multiply-adds. A convolution reuses each weight at every spatial position, so a layer with 36 thousand weights running on a 224x224 map performs on the order of a billion multiply-adds. Summed over the network, the convolution stack accounts for essentially all of VGG-16's arithmetic — on the order of ten billion multiply-adds per image — while holding about a tenth of its weights. The layers that dominate the file size are not the layers that dominate the latency. ## What to take from it Three habits come out of this example. First, always locate parameters by shape rather than by intuition: multiply the flatten width by the first dense width before guessing. Second, distinguish the three budgets — parameter count, activation memory and multiply-adds — because a design decision usually improves one and worsens another. Third, treat a flatten-then-dense head as a resolution commitment, since everything about its size follows from the input dimensions you fixed at design time.

  • What happens to the parameter count if VGG-16 is fed 448x448 images instead of 224x224?
    The convolution stack is unchanged, since its kernels do not depend on spatial size. But five poolings now leave a 14x14x512 map, so the flatten is 100352 wide and the first dense layer jumps from about 103 million weights to about 411 million. The model goes from roughly 138 million parameters to roughly 446 million without gaining one new feature detector.
  • Why can a network with a flatten-then-dense head not accept arbitrary input sizes?
    The dense layer's weight matrix has a fixed input dimension equal to the flattened feature-map width. Change the input resolution and that width changes, so the stored matrix no longer matches the vector it is given. The convolution stack is size-agnostic because it slides shared kernels; the fixed input requirement comes entirely from the head.
  • If the dense head holds most of the weights, does it also dominate the compute?
    No, the opposite. A dense layer performs about one multiply-add per weight, so the whole head is around 124 million. A convolution reuses each weight at every spatial position, so early layers with tens of thousands of weights run on the order of a billion multiply-adds each. The convolution stack holds about a tenth of the weights and does essentially all of the arithmetic.

saying these in an interview costs you the question

  • Assumes the convolution layers hold most of the parameters
  • Thinks convolution weights grow with input resolution
  • Treats parameter count as a proxy for compute cost
  • Cannot explain where the 7x7x512 shape comes from
  • Says the 1000-way output layer is the biggest one

context