skip to content

Grouped and Depthwise Convolutions

A 1x1 convolution mixes channels and cuts their count, while grouped and depthwise-separable convolutions split spatial from channel mixing for a fraction of the multiply-adds.

on this pageshow

questions

4

What does a 1x1 convolution compute in a CNN, and why would you add one?

level: juniorimportance: must knowfreq 72%

answer

  1. it only ever touches one pixel
  2. works across channels, not across space
  3. C_in times C_out weights, nothing more
  4. shrink the depth before an expensive filter

basics

~10 s

A 1x1 convolution mixes channels at one spatial position: each output channel is a learned linear combination of all input channels at that pixel. It changes channel count cheaply, leaving height and width untouched.

solid answer

~50 s

A 1x1 convolution slides a kernel of spatial size one over the feature map, so it never looks at a neighbouring pixel. At each position it takes the `C_in`-long vector of channel values and projects it to a `C_out`-long vector with a learned matrix, which is why it is often called a pointwise convolution or channel mixing. It costs `C_in * C_out` weights and `H * W * C_in * C_out` multiply-accumulates, nine times less than a 3x3 layer with the same channel counts. The three reasons to add one are: change the channel count (reduce or expand), add cross-channel mixing plus a nonlinearity for free in spatial terms, and cut the cost of an expensive spatial layer by shrinking its input depth first. What it cannot do is enlarge the receptive field — spatially, each output unit still sees exactly one input position.

go deeper

for a junior

Be ready to say exactly what a 1x1 kernel touches: one spatial position, all channels. Know that it changes the channel count and leaves height and width alone.

for a middle

Expect to compute its cost out loud: C_in times C_out weights, and H times W times C_in times C_out multiply-accumulates, which is nine times less than a 3x3 with the same channel counts.

for a senior

Show where you place these layers in a real design: reducing depth before an expensive spatial branch, restoring depth after a cheap one, and knowing what share of the model's weights ends up in them.

for a principal

Own the tradeoff between spending capacity on channel mixing versus spatial extent, and be able to argue when a tall stack of cheap 1x1 layers is the wrong shape for the accelerator you are targeting.

## The operator A convolution layer is defined by a kernel of spatial size `k x k`, an input channel count `C_in`, and an output channel count `C_out`. Its weight tensor has shape `C_out x C_in x k x k`. Setting `k = 1` collapses the spatial part entirely: each filter is just a vector of length `C_in`, and the output at channel `o`, position `(h, w)` is ``` y[o, h, w] = sum over c of W[o, c] * x[c, h, w] + b[o] ``` No neighbouring pixel enters the sum. The layer is a per-position linear map from `C_in` numbers to `C_out` numbers, with the *same* matrix reused at every one of the `H * W` positions. Equivalently, if you flatten the feature map into a `C_in x (H*W)` matrix, a 1x1 convolution is one matrix product with a `C_out x C_in` weight matrix. ## Cost - Parameters: `C_in * C_out` (plus `C_out` biases). - Multiply-accumulates: `H * W * C_in * C_out` at stride 1. Against a `k x k` layer with the same channel counts, that is a factor of `k * k` fewer — nine times fewer for a 3x3. This is why 1x1 layers are treated as the cheap way to move channels around, even though in a channel-heavy network they can still hold most of the parameters. ## Why you add one **1. Change the channel count.** This is the plain answer. Feature maps arrive with whatever depth the previous stage produced; a 1x1 projects that to whatever depth the next stage wants, without resampling space. **2. Cross-channel mixing with a nonlinearity.** Because a nonlinearity normally follows, a 1x1 layer is not merely a reshaping — it lets the network learn combinations of feature channels ("this edge detector AND that colour channel") and then bend them. Stacking a 1x1 on top of a spatial layer adds representational depth for a fraction of the cost of another spatial layer. **3. Dimension reduction in front of an expensive layer.** This is the classic use. Consider an Inception-style branch that wants a 5x5 convolution on a 256-channel input, producing 64 channels. Done directly it costs ``` 5 * 5 * 256 * 64 = 409,600 weights ``` Insert a 1x1 that first reduces 256 channels to 64: ``` 1 * 1 * 256 * 64 = 16,384 5 * 5 * 64 * 64 = 102,400 total = 118,784 weights ``` about 3.4x cheaper, and the multiply-accumulates fall by the same ratio because both paths run at the same spatial resolution. The saving is bounded by how aggressively you reduce: cutting to 16 channels instead of 64 gives `256*16 + 25*16*64 = 29,696`, close to 14x. The reduction layer is doing the work a human would call "summarise the 256 channels into a compact description, then spend the expensive spatial filter on that". ## What it does not do A 1x1 convolution does **not** change the spatial receptive field. If the units feeding it saw a 7x7 window of the input, the units after it still see a 7x7 window. It does not change height or width either (at stride 1). It cannot correct spatial misalignment or aggregate context. Anything requiring a wider view still needs a spatial kernel, a stride, a pooling step, or dilation. ## Relation to a fully connected layer A 1x1 convolution *is* a fully connected layer over the channel axis, applied independently at every position with shared weights. The distinction from a dense layer on a flattened feature map matters: the dense layer has `C_in * H * W * C_out` weights, ties the model to one input resolution, and destroys the spatial map. The 1x1 has `C_in * C_out` weights, is resolution-agnostic, and preserves the map for later layers. ## Where you see them Channel-reduction stems in multi-branch modules; the pointwise half of a separable block, where a `C_in x C_out` 1x1 typically holds well over ninety percent of the block's weights; projection shortcuts that need to match channel counts before an addition; and the final classification head that maps a pooled feature vector to logits. A common interview trap is to call a 1x1 layer "pointless because it does nothing spatial" — the whole point is that it operates on the axis a spatial kernel treats as fixed.

  • How is a 1x1 convolution different from a fully connected layer on the feature map?
    A 1x1 convolution is a fully connected layer over the channel axis only, with the same weights reused at every spatial position. That gives it `C_in * C_out` weights, independence from input resolution, and a preserved spatial map. A dense layer on the flattened map has `C_in * H * W * C_out` weights, is locked to one resolution, and throws the spatial structure away.
  • Does inserting a 1x1 layer change what an output unit can see in the input image?
    No. The spatial receptive field is unchanged, because the kernel spans a single position. The layer adds a learned channel mix and, with the activation after it, extra nonlinearity — but if you need a wider view of the image you still need a larger kernel, a stride, pooling, or dilation.
  • If 1x1 layers are so cheap, why can they still dominate a model's parameter count?
    Cheap is relative to `k * k`, not absolute. A 1x1 layer still costs `C_in * C_out` weights, which grows quadratically with width. In a network whose spatial filters are per-channel and whose widths are large, the 1x1 layers can carry the large majority of the weights and of the multiply-accumulates despite the trivial kernel.

A 1x1 convolution is a colour-correction matrix applied identically to every pixel: it re-mixes the channels at each location and never once looks at a neighbour.

saying these in an interview costs you the question

  • Says a 1x1 convolution is an identity or does nothing useful
  • Claims it enlarges the receptive field
  • Thinks it changes the spatial height and width
  • Confuses it with a dense layer on the flattened feature map
  • Assumes it is always negligible in parameter count

context

open as a page

Why is a depthwise-separable convolution roughly nine times cheaper than a dense 3x3 layer?

level: middleimportance: must knowfreq 66%

basics

~20 s

It factorises a dense layer into a spatial stage with one k x k filter per input channel and a pointwise stage that mixes channels. Cost falls to 1/C_out + 1/(k*k) of dense, about 1/9 at k=3.

open as a page

How does splitting a convolution into g groups change its cost and its channel connectivity?

level: middleimportance: should knowfreq 50%

basics

~20 s

Grouping partitions input and output channels into g disjoint sets, each output group seeing only its own input group. Weights and multiply-accumulates both drop by a factor of g, but no channel combines across group boundaries.

open as a page

You swapped 3x3 convolutions for depthwise-separable blocks, cut FLOPs 8x, but latency only halved — why?

level: seniorimportance: nice to knowfreq 34%

basics

~10 s

A depthwise stage removes arithmetic but not data movement: it streams the full activation tensor while doing only k-squared multiply-accumulates per element. Runtime becomes memory-bound, and one layer becoming two adds further passes.

open as a page