skip to content

Backbone Architectures

The families you get by stacking the operator: VGG's uniform small kernels, ResNet's identity shortcuts, and the cheap 1x1, grouped and depthwise blocks behind MobileNet.

on this pageshow

explore

questions

11

What does a 1x1 convolution compute in a CNN, and why would you add one?

level: juniorimportance: must knowfreq 72%

answer

  1. it only ever touches one pixel
  2. works across channels, not across space
  3. C_in times C_out weights, nothing more
  4. shrink the depth before an expensive filter

basics

~10 s

A 1x1 convolution mixes channels at one spatial position: each output channel is a learned linear combination of all input channels at that pixel. It changes channel count cheaply, leaving height and width untouched.

solid answer

~50 s

A 1x1 convolution slides a kernel of spatial size one over the feature map, so it never looks at a neighbouring pixel. At each position it takes the `C_in`-long vector of channel values and projects it to a `C_out`-long vector with a learned matrix, which is why it is often called a pointwise convolution or channel mixing. It costs `C_in * C_out` weights and `H * W * C_in * C_out` multiply-accumulates, nine times less than a 3x3 layer with the same channel counts. The three reasons to add one are: change the channel count (reduce or expand), add cross-channel mixing plus a nonlinearity for free in spatial terms, and cut the cost of an expensive spatial layer by shrinking its input depth first. What it cannot do is enlarge the receptive field — spatially, each output unit still sees exactly one input position.

go deeper

for a junior

Be ready to say exactly what a 1x1 kernel touches: one spatial position, all channels. Know that it changes the channel count and leaves height and width alone.

for a middle

Expect to compute its cost out loud: C_in times C_out weights, and H times W times C_in times C_out multiply-accumulates, which is nine times less than a 3x3 with the same channel counts.

for a senior

Show where you place these layers in a real design: reducing depth before an expensive spatial branch, restoring depth after a cheap one, and knowing what share of the model's weights ends up in them.

for a principal

Own the tradeoff between spending capacity on channel mixing versus spatial extent, and be able to argue when a tall stack of cheap 1x1 layers is the wrong shape for the accelerator you are targeting.

## The operator A convolution layer is defined by a kernel of spatial size `k x k`, an input channel count `C_in`, and an output channel count `C_out`. Its weight tensor has shape `C_out x C_in x k x k`. Setting `k = 1` collapses the spatial part entirely: each filter is just a vector of length `C_in`, and the output at channel `o`, position `(h, w)` is ``` y[o, h, w] = sum over c of W[o, c] * x[c, h, w] + b[o] ``` No neighbouring pixel enters the sum. The layer is a per-position linear map from `C_in` numbers to `C_out` numbers, with the *same* matrix reused at every one of the `H * W` positions. Equivalently, if you flatten the feature map into a `C_in x (H*W)` matrix, a 1x1 convolution is one matrix product with a `C_out x C_in` weight matrix. ## Cost - Parameters: `C_in * C_out` (plus `C_out` biases). - Multiply-accumulates: `H * W * C_in * C_out` at stride 1. Against a `k x k` layer with the same channel counts, that is a factor of `k * k` fewer — nine times fewer for a 3x3. This is why 1x1 layers are treated as the cheap way to move channels around, even though in a channel-heavy network they can still hold most of the parameters. ## Why you add one **1. Change the channel count.** This is the plain answer. Feature maps arrive with whatever depth the previous stage produced; a 1x1 projects that to whatever depth the next stage wants, without resampling space. **2. Cross-channel mixing with a nonlinearity.** Because a nonlinearity normally follows, a 1x1 layer is not merely a reshaping — it lets the network learn combinations of feature channels ("this edge detector AND that colour channel") and then bend them. Stacking a 1x1 on top of a spatial layer adds representational depth for a fraction of the cost of another spatial layer. **3. Dimension reduction in front of an expensive layer.** This is the classic use. Consider an Inception-style branch that wants a 5x5 convolution on a 256-channel input, producing 64 channels. Done directly it costs ``` 5 * 5 * 256 * 64 = 409,600 weights ``` Insert a 1x1 that first reduces 256 channels to 64: ``` 1 * 1 * 256 * 64 = 16,384 5 * 5 * 64 * 64 = 102,400 total = 118,784 weights ``` about 3.4x cheaper, and the multiply-accumulates fall by the same ratio because both paths run at the same spatial resolution. The saving is bounded by how aggressively you reduce: cutting to 16 channels instead of 64 gives `256*16 + 25*16*64 = 29,696`, close to 14x. The reduction layer is doing the work a human would call "summarise the 256 channels into a compact description, then spend the expensive spatial filter on that". ## What it does not do A 1x1 convolution does **not** change the spatial receptive field. If the units feeding it saw a 7x7 window of the input, the units after it still see a 7x7 window. It does not change height or width either (at stride 1). It cannot correct spatial misalignment or aggregate context. Anything requiring a wider view still needs a spatial kernel, a stride, a pooling step, or dilation. ## Relation to a fully connected layer A 1x1 convolution *is* a fully connected layer over the channel axis, applied independently at every position with shared weights. The distinction from a dense layer on a flattened feature map matters: the dense layer has `C_in * H * W * C_out` weights, ties the model to one input resolution, and destroys the spatial map. The 1x1 has `C_in * C_out` weights, is resolution-agnostic, and preserves the map for later layers. ## Where you see them Channel-reduction stems in multi-branch modules; the pointwise half of a separable block, where a `C_in x C_out` 1x1 typically holds well over ninety percent of the block's weights; projection shortcuts that need to match channel counts before an addition; and the final classification head that maps a pooled feature vector to logits. A common interview trap is to call a 1x1 layer "pointless because it does nothing spatial" — the whole point is that it operates on the axis a spatial kernel treats as fixed.

  • How is a 1x1 convolution different from a fully connected layer on the feature map?
    A 1x1 convolution is a fully connected layer over the channel axis only, with the same weights reused at every spatial position. That gives it `C_in * C_out` weights, independence from input resolution, and a preserved spatial map. A dense layer on the flattened map has `C_in * H * W * C_out` weights, is locked to one resolution, and throws the spatial structure away.
  • Does inserting a 1x1 layer change what an output unit can see in the input image?
    No. The spatial receptive field is unchanged, because the kernel spans a single position. The layer adds a learned channel mix and, with the activation after it, extra nonlinearity — but if you need a wider view of the image you still need a larger kernel, a stride, pooling, or dilation.
  • If 1x1 layers are so cheap, why can they still dominate a model's parameter count?
    Cheap is relative to `k * k`, not absolute. A 1x1 layer still costs `C_in * C_out` weights, which grows quadratically with width. In a network whose spatial filters are per-channel and whose widths are large, the 1x1 layers can carry the large majority of the weights and of the multiply-accumulates despite the trivial kernel.

A 1x1 convolution is a colour-correction matrix applied identically to every pixel: it re-mixes the channels at each location and never once looks at a neighbour.

saying these in an interview costs you the question

  • Says a 1x1 convolution is an identity or does nothing useful
  • Claims it enlarges the receptive field
  • Thinks it changes the spatial height and width
  • Confuses it with a dense layer on the flattened feature map
  • Assumes it is always negligible in parameter count

context

open as a page

Why does a residual block add its input back to its output instead of just stacking layers?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A residual block computes F(x) + x, so its layers only learn the change to make to the input. Behaving like an identity then just means pushing F toward zero, which plain stacked layers struggle to fit.

open as a page

Why does VGG stack two 3x3 convolutions instead of using one 5x5 layer?

level: middleimportance: must knowfreq 70%

basics

~10 s

Two stacked 3x3 layers see the same 5x5 input patch as one 5x5 layer, but use 18C-squared weights instead of 25C-squared and apply two nonlinearities instead of one. Same reach, cheaper, more expressive.

open as a page

Why is a depthwise-separable convolution roughly nine times cheaper than a dense 3x3 layer?

level: middleimportance: must knowfreq 66%

basics

~20 s

It factorises a dense layer into a spatial stage with one k x k filter per input channel and a pointwise stage that mixes channels. Cost falls to 1/C_out + 1/(k*k) of dense, about 1/9 at k=3.

open as a page

In VGG-16, which layers hold most of the 138 million parameters, and why?

level: juniorimportance: should knowfreq 45%

basics

~20 s

About 90% sit in the three fully connected layers at the end. The first alone, mapping the flattened 7x7x512 feature map to 4096 units, holds roughly 103 million weights; all 13 convolution layers together hold under 15 million.

open as a page

How does splitting a convolution into g groups change its cost and its channel connectivity?

level: middleimportance: should knowfreq 50%

basics

~20 s

Grouping partitions input and output channels into g disjoint sets, each output group seeing only its own input group. Weights and multiply-accumulates both drop by a factor of g, but no channel combines across group boundaries.

open as a page

Why does ResNet-50 use a 1x1-3x3-1x1 bottleneck block instead of two 3x3 layers?

level: middleimportance: should knowfreq 52%

basics

~20 s

A 3x3 convolution's cost grows with input channels times output channels, so it is ruinous at wide layers. The bottleneck squeezes width down with a 1x1, runs the 3x3 cheaply, then restores width — roughly an order of magnitude fewer weights.

open as a page

In a ResNet, when can a shortcut be a plain identity and when must it be a projection?

level: middleimportance: should knowfreq 58%

basics

~20 s

A ResNet shortcut can stay a plain identity only while the branch preserves spatial size and channel count, since the two are added elementwise. Where a stage strides down and widens, it becomes a 1x1 convolution.

open as a page

What does AlexNet's 11x11 stride-4 first layer discard that a 3x3 stem keeps?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Fine spatial detail. A stride-4 first layer samples the image on a coarse grid, aliasing away structure finer than four pixels, and one wide linear filter plus one nonlinearity is a weaker map than a stride-1 3x3 stack.

open as a page

You swapped 3x3 convolutions for depthwise-separable blocks, cut FLOPs 8x, but latency only halved — why?

level: seniorimportance: nice to knowfreq 34%

basics

~10 s

A depthwise stage removes arithmetic but not data movement: it streams the full activation tensor while doing only k-squared multiply-accumulates per element. Runtime becomes memory-bound, and one layer becoming two adds further passes.

open as a page

When would you concatenate skip features as DenseNet does rather than add them as ResNet does?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Concatenation keeps every earlier feature map intact so later layers can select among them; addition merges them irreversibly but holds channel count fixed. Choose concatenation for feature reuse at moderate depth, addition for very deep throughput-sensitive backbones.

open as a page