What does a 1x1 convolution compute in a CNN, and why would you add one?
answer
- it only ever touches one pixel
- works across channels, not across space
- C_in times C_out weights, nothing more
- shrink the depth before an expensive filter
basics
~10 sA 1x1 convolution mixes channels at one spatial position: each output channel is a learned linear combination of all input channels at that pixel. It changes channel count cheaply, leaving height and width untouched.
solid answer
~50 sA 1x1 convolution slides a kernel of spatial size one over the feature map, so it never looks at a neighbouring pixel. At each position it takes the `C_in`-long vector of channel values and projects it to a `C_out`-long vector with a learned matrix, which is why it is often called a pointwise convolution or channel mixing. It costs `C_in * C_out` weights and `H * W * C_in * C_out` multiply-accumulates, nine times less than a 3x3 layer with the same channel counts. The three reasons to add one are: change the channel count (reduce or expand), add cross-channel mixing plus a nonlinearity for free in spatial terms, and cut the cost of an expensive spatial layer by shrinking its input depth first. What it cannot do is enlarge the receptive field — spatially, each output unit still sees exactly one input position.
go deeper
Be ready to say exactly what a 1x1 kernel touches: one spatial position, all channels. Know that it changes the channel count and leaves height and width alone.
Expect to compute its cost out loud: C_in times C_out weights, and H times W times C_in times C_out multiply-accumulates, which is nine times less than a 3x3 with the same channel counts.
Show where you place these layers in a real design: reducing depth before an expensive spatial branch, restoring depth after a cheap one, and knowing what share of the model's weights ends up in them.
Own the tradeoff between spending capacity on channel mixing versus spatial extent, and be able to argue when a tall stack of cheap 1x1 layers is the wrong shape for the accelerator you are targeting.
## The operator A convolution layer is defined by a kernel of spatial size `k x k`, an input channel count `C_in`, and an output channel count `C_out`. Its weight tensor has shape `C_out x C_in x k x k`. Setting `k = 1` collapses the spatial part entirely: each filter is just a vector of length `C_in`, and the output at channel `o`, position `(h, w)` is ``` y[o, h, w] = sum over c of W[o, c] * x[c, h, w] + b[o] ``` No neighbouring pixel enters the sum. The layer is a per-position linear map from `C_in` numbers to `C_out` numbers, with the *same* matrix reused at every one of the `H * W` positions. Equivalently, if you flatten the feature map into a `C_in x (H*W)` matrix, a 1x1 convolution is one matrix product with a `C_out x C_in` weight matrix. ## Cost - Parameters: `C_in * C_out` (plus `C_out` biases). - Multiply-accumulates: `H * W * C_in * C_out` at stride 1. Against a `k x k` layer with the same channel counts, that is a factor of `k * k` fewer — nine times fewer for a 3x3. This is why 1x1 layers are treated as the cheap way to move channels around, even though in a channel-heavy network they can still hold most of the parameters. ## Why you add one **1. Change the channel count.** This is the plain answer. Feature maps arrive with whatever depth the previous stage produced; a 1x1 projects that to whatever depth the next stage wants, without resampling space. **2. Cross-channel mixing with a nonlinearity.** Because a nonlinearity normally follows, a 1x1 layer is not merely a reshaping — it lets the network learn combinations of feature channels ("this edge detector AND that colour channel") and then bend them. Stacking a 1x1 on top of a spatial layer adds representational depth for a fraction of the cost of another spatial layer. **3. Dimension reduction in front of an expensive layer.** This is the classic use. Consider an Inception-style branch that wants a 5x5 convolution on a 256-channel input, producing 64 channels. Done directly it costs ``` 5 * 5 * 256 * 64 = 409,600 weights ``` Insert a 1x1 that first reduces 256 channels to 64: ``` 1 * 1 * 256 * 64 = 16,384 5 * 5 * 64 * 64 = 102,400 total = 118,784 weights ``` about 3.4x cheaper, and the multiply-accumulates fall by the same ratio because both paths run at the same spatial resolution. The saving is bounded by how aggressively you reduce: cutting to 16 channels instead of 64 gives `256*16 + 25*16*64 = 29,696`, close to 14x. The reduction layer is doing the work a human would call "summarise the 256 channels into a compact description, then spend the expensive spatial filter on that". ## What it does not do A 1x1 convolution does **not** change the spatial receptive field. If the units feeding it saw a 7x7 window of the input, the units after it still see a 7x7 window. It does not change height or width either (at stride 1). It cannot correct spatial misalignment or aggregate context. Anything requiring a wider view still needs a spatial kernel, a stride, a pooling step, or dilation. ## Relation to a fully connected layer A 1x1 convolution *is* a fully connected layer over the channel axis, applied independently at every position with shared weights. The distinction from a dense layer on a flattened feature map matters: the dense layer has `C_in * H * W * C_out` weights, ties the model to one input resolution, and destroys the spatial map. The 1x1 has `C_in * C_out` weights, is resolution-agnostic, and preserves the map for later layers. ## Where you see them Channel-reduction stems in multi-branch modules; the pointwise half of a separable block, where a `C_in x C_out` 1x1 typically holds well over ninety percent of the block's weights; projection shortcuts that need to match channel counts before an addition; and the final classification head that maps a pooled feature vector to logits. A common interview trap is to call a 1x1 layer "pointless because it does nothing spatial" — the whole point is that it operates on the axis a spatial kernel treats as fixed.
- How is a 1x1 convolution different from a fully connected layer on the feature map?A 1x1 convolution is a fully connected layer over the channel axis only, with the same weights reused at every spatial position. That gives it `C_in * C_out` weights, independence from input resolution, and a preserved spatial map. A dense layer on the flattened map has `C_in * H * W * C_out` weights, is locked to one resolution, and throws the spatial structure away.
- Does inserting a 1x1 layer change what an output unit can see in the input image?No. The spatial receptive field is unchanged, because the kernel spans a single position. The layer adds a learned channel mix and, with the activation after it, extra nonlinearity — but if you need a wider view of the image you still need a larger kernel, a stride, pooling, or dilation.
- If 1x1 layers are so cheap, why can they still dominate a model's parameter count?Cheap is relative to `k * k`, not absolute. A 1x1 layer still costs `C_in * C_out` weights, which grows quadratically with width. In a network whose spatial filters are per-channel and whose widths are large, the 1x1 layers can carry the large majority of the weights and of the multiply-accumulates despite the trivial kernel.
A 1x1 convolution is a colour-correction matrix applied identically to every pixel: it re-mixes the channels at each location and never once looks at a neighbour.
saying these in an interview costs you the question
- Says a 1x1 convolution is an identity or does nothing useful
- Claims it enlarges the receptive field
- Thinks it changes the spatial height and width
- Confuses it with a dense layer on the flattened feature map
- Assumes it is always negligible in parameter count