skip to content

Why does ResNet-50 use a 1x1-3x3-1x1 bottleneck block instead of two 3x3 layers?

level: middleimportance: should knowfreq 52%

answer

  1. 3x3 cost scales with channels squared
  2. squeeze, convolve, restore
  3. the 3x3 runs at a quarter width
  4. 1x1 layers mix channels, no spatial extent
  5. expansion factor four at the output

basics

~20 s

A 3x3 convolution's cost grows with input channels times output channels, so it is ruinous at wide layers. The bottleneck squeezes width down with a 1x1, runs the 3x3 cheaply, then restores width — roughly an order of magnitude fewer weights.

solid answer

~50 s

The two-3x3 basic block used in the shallower ResNets does not scale to wide layers: a 3x3 convolution needs `9 * C_in * C_out` weights, so at 256 channels one layer is already about 590k weights and a block is about 1.2M. The bottleneck block reorders the work — a 1x1 reduces 256 channels to 64, the 3x3 runs at that narrow width, and a second 1x1 expands back to 256. That is roughly 70k weights for the same input and output shape, about seventeen times cheaper, and the saved budget buys depth: ResNet-50, -101 and -152 all use bottlenecks while -18 and -34 use basic blocks. The convention is an **expansion factor of 4**: the block's output width is four times the internal 3x3 width. The spatial modelling still happens only in the 3x3; the 1x1 layers mix channels.

code

python · 14 lines
python
def conv_weights(k, c_in, c_out):
    return k * k * c_in * c_out

width = 256          # channels entering and leaving the block
narrow = width // 4  # internal width of the bottleneck

basic = 2 * conv_weights(3, width, width)
bottleneck = (conv_weights(1, width, narrow)
              + conv_weights(3, narrow, narrow)
              + conv_weights(1, narrow, width))

print("basic     ", basic)       # 1179648
print("bottleneck", bottleneck)  # 69632
print("ratio     ", round(basic / bottleneck, 1))  # 16.9

go deeper

for a junior

Know the shape of the block — 1x1 down, 3x3 at the narrow width, 1x1 back up — and that its purpose is to make wide layers affordable rather than to change what the block can see.

for a middle

Do the arithmetic out loud: nine times input channels times output channels per 3x3, quadratic in width, so squeezing to a quarter width cuts the 3x3 by sixteen. State the 4x expansion convention and where basic blocks are used instead.

for a senior

Show the second-order knowledge: parameters fall much more than activation memory, latency does not track parameter count, and an over-aggressive reduction becomes an information choke point.

for a principal

Own the budget argument. Depth versus width versus input resolution is an empirical tradeoff under a fixed cost target, and wide-shallow residual networks are a serious competitor to deep-narrow ones on real hardware.

## The cost model you need first A convolution with kernel size `k`, `C_in` input channels and `C_out` output channels holds `k * k * C_in * C_out` weights, and its multiply-accumulate count is that number times the output's spatial positions. Two things follow. Cost is **quadratic in width** — double the channels and you quadruple the layer. And it is proportional to `k*k`, so a 3x3 costs nine times a 1x1 at the same widths. Backbones widen as they go deeper, precisely because they also downsample: later stages run at 256, 512, 1024, 2048 channels. Quadratic growth in width is what makes a naive deep network unaffordable. ## Basic block versus bottleneck The **basic block** used in ResNet-18 and ResNet-34 is two 3x3 convolutions at the same width, with the shortcut added at the end. At a width of 256: ``` 3x3, 256 -> 256 : 9 * 256 * 256 = 589,824 3x3, 256 -> 256 : 589,824 total : 1,179,648 ``` The **bottleneck block** used in ResNet-50 and deeper replaces that with three convolutions: ``` 1x1, 256 -> 64 : 1 * 256 * 64 = 16,384 (reduce) 3x3, 64 -> 64 : 9 * 64 * 64 = 36,864 (transform, cheaply) 1x1, 64 -> 256 : 1 * 64 * 256 = 16,384 (restore) total : 69,632 ``` Same input shape, same output shape, about seventeen times fewer weights, and the same story holds for multiply-accumulates. The narrow 3x3 is where all the spatial pattern-matching happens; the 1x1 layers only mix channels, which is cheap because they have no spatial extent. ## The expansion factor The standard convention fixes the ratio: **output width = 4 x internal width**. If the 3x3 runs at 64 channels, the block emits 256; at 128 it emits 512; at 512 it emits 2048. This is why bottleneck ResNets end up so wide at the last stage while ResNet-34 tops out much narrower, and why the block's first appearance in a stage always needs a projection shortcut — its input width does not yet match the expanded output. It also explains layer counting. A ResNet-50 has fewer *blocks* than you might guess from the name: three convolutions per block, times sixteen blocks, plus the stem convolution and the classifier, gives fifty weighted layers. ResNet-34 has two per block. ## Why this particular shape Three design instincts are at work. **Spend the expensive operator where it is cheapest.** The 3x3 is the only layer that models spatial structure, and you want it, but you want it at the narrowest width you can tolerate. Reduce first, then convolve. **Keep the representation wide between blocks.** The residual stream carries the wide, 256-channel representation from block to block; the narrow width exists only inside the block. Later blocks still see rich features. **Trade width for depth under a fixed budget.** Given a fixed parameter or FLOP budget, the empirical answer through this era was that more blocks beat wider blocks, up to a point. Bottlenecks are the mechanism that makes the trade available. (The counter-position — that wider, shallower residual networks train faster and match accuracy — is a real and defensible one; "deeper is always better" is not a fact.) ## Consequences and gotchas - **Nonlinearity budget.** The bottleneck has three weighted layers but the middle one is narrow, so a block has less capacity than its parameter count in a basic block would suggest; you compensate with more blocks. - **Memory versus parameters.** Bottlenecks cut parameters far more than they cut *activation* memory, because the wide tensors between blocks still have to be stored for the backward pass. Teams surprised by memory usage after switching to a deeper bottleneck model are usually seeing this. - **Latency does not track parameters.** Three small layers can be slower per image than two larger ones on hardware that likes big dense operations; parameter count is a proxy for cost, not a measurement of it. - **A 1x1 that reduces width is a real bottleneck.** If you shrink too aggressively, the reduction layer becomes an information choke point and accuracy suffers. The 4x ratio is a tuned default, not a law. ## What to say in an interview Give the arithmetic. Say that a 3x3 costs `9 * C_in * C_out`, that this is quadratic in width, that the bottleneck squeezes to a quarter of the width so the 3x3 costs sixteen times less, and that the two 1x1 layers are cheap enough not to give the saving back. Then name where each block type is used and note the 4x expansion. That answer is short, quantitative, and shows you can size a layer.

  • Does the bottleneck block cut activation memory as much as it cuts parameters?
    No. The wide tensors flowing between blocks are unchanged, and they are what dominates the activations you must keep for the backward pass. The bottleneck only narrows the tensor inside the block. Teams that swap a basic-block model for a deeper bottleneck one often see parameters fall and training memory rise, because depth added more stored activations.
  • Why do ResNet-18 and ResNet-34 not use bottleneck blocks?
    They are narrow enough that the quadratic term never bites — their widest stage is 512 channels and they have few blocks, so two 3x3 layers are affordable and give more capacity per block. Bottlenecks pay off once you want fifty or more layers and widths in the thousands.
  • What breaks if you shrink the internal width far below a quarter of the output width?
    The reduction 1x1 becomes a genuine information choke point: everything the 3x3 can model has to fit through that narrow representation. Accuracy degrades even though parameters keep falling. The 4x expansion ratio is an empirically tuned default, and pushing it is a real accuracy-versus-cost knob, not free savings.
  • Is depth always the better use of a fixed parameter budget?
    No, and it is worth saying so. Widening a residual network while keeping it shallower can match the accuracy of a much deeper one and train faster, because wide layers use hardware better and shorter stacks optimize more easily. The honest answer is that depth-versus-width is a budget experiment, not a settled rule.

saying these in an interview costs you the question

  • Says the 1x1 layers enlarge the receptive field
  • Thinks the bottleneck also cuts activation memory proportionally
  • Cannot state that a 3x3 costs nine times C_in times C_out
  • Believes fewer parameters always means lower latency
  • Forgets the four-fold expansion at the block output

context