skip to content

Convolutional Networks

You will learn why convolutions beat dense layers on images, how ResNet's skip connections made very deep nets trainable, and what gets built on a backbone — detection scored by mAP, segmentation by U-Net. Expect to compute output shapes by hand; interviewers use that arithmetic as a fast competence check.

on this pageshow

explore

questions

page 1 of 2

What does a 1x1 convolution compute in a CNN, and why would you add one?

level: juniorimportance: must knowfreq 72%

answer

  1. it only ever touches one pixel
  2. works across channels, not across space
  3. C_in times C_out weights, nothing more
  4. shrink the depth before an expensive filter

basics

~10 s

A 1x1 convolution mixes channels at one spatial position: each output channel is a learned linear combination of all input channels at that pixel. It changes channel count cheaply, leaving height and width untouched.

solid answer

~50 s

A 1x1 convolution slides a kernel of spatial size one over the feature map, so it never looks at a neighbouring pixel. At each position it takes the `C_in`-long vector of channel values and projects it to a `C_out`-long vector with a learned matrix, which is why it is often called a pointwise convolution or channel mixing. It costs `C_in * C_out` weights and `H * W * C_in * C_out` multiply-accumulates, nine times less than a 3x3 layer with the same channel counts. The three reasons to add one are: change the channel count (reduce or expand), add cross-channel mixing plus a nonlinearity for free in spatial terms, and cut the cost of an expensive spatial layer by shrinking its input depth first. What it cannot do is enlarge the receptive field — spatially, each output unit still sees exactly one input position.

go deeper

for a junior

Be ready to say exactly what a 1x1 kernel touches: one spatial position, all channels. Know that it changes the channel count and leaves height and width alone.

for a middle

Expect to compute its cost out loud: C_in times C_out weights, and H times W times C_in times C_out multiply-accumulates, which is nine times less than a 3x3 with the same channel counts.

for a senior

Show where you place these layers in a real design: reducing depth before an expensive spatial branch, restoring depth after a cheap one, and knowing what share of the model's weights ends up in them.

for a principal

Own the tradeoff between spending capacity on channel mixing versus spatial extent, and be able to argue when a tall stack of cheap 1x1 layers is the wrong shape for the accelerator you are targeting.

## The operator A convolution layer is defined by a kernel of spatial size `k x k`, an input channel count `C_in`, and an output channel count `C_out`. Its weight tensor has shape `C_out x C_in x k x k`. Setting `k = 1` collapses the spatial part entirely: each filter is just a vector of length `C_in`, and the output at channel `o`, position `(h, w)` is ``` y[o, h, w] = sum over c of W[o, c] * x[c, h, w] + b[o] ``` No neighbouring pixel enters the sum. The layer is a per-position linear map from `C_in` numbers to `C_out` numbers, with the *same* matrix reused at every one of the `H * W` positions. Equivalently, if you flatten the feature map into a `C_in x (H*W)` matrix, a 1x1 convolution is one matrix product with a `C_out x C_in` weight matrix. ## Cost - Parameters: `C_in * C_out` (plus `C_out` biases). - Multiply-accumulates: `H * W * C_in * C_out` at stride 1. Against a `k x k` layer with the same channel counts, that is a factor of `k * k` fewer — nine times fewer for a 3x3. This is why 1x1 layers are treated as the cheap way to move channels around, even though in a channel-heavy network they can still hold most of the parameters. ## Why you add one **1. Change the channel count.** This is the plain answer. Feature maps arrive with whatever depth the previous stage produced; a 1x1 projects that to whatever depth the next stage wants, without resampling space. **2. Cross-channel mixing with a nonlinearity.** Because a nonlinearity normally follows, a 1x1 layer is not merely a reshaping — it lets the network learn combinations of feature channels ("this edge detector AND that colour channel") and then bend them. Stacking a 1x1 on top of a spatial layer adds representational depth for a fraction of the cost of another spatial layer. **3. Dimension reduction in front of an expensive layer.** This is the classic use. Consider an Inception-style branch that wants a 5x5 convolution on a 256-channel input, producing 64 channels. Done directly it costs ``` 5 * 5 * 256 * 64 = 409,600 weights ``` Insert a 1x1 that first reduces 256 channels to 64: ``` 1 * 1 * 256 * 64 = 16,384 5 * 5 * 64 * 64 = 102,400 total = 118,784 weights ``` about 3.4x cheaper, and the multiply-accumulates fall by the same ratio because both paths run at the same spatial resolution. The saving is bounded by how aggressively you reduce: cutting to 16 channels instead of 64 gives `256*16 + 25*16*64 = 29,696`, close to 14x. The reduction layer is doing the work a human would call "summarise the 256 channels into a compact description, then spend the expensive spatial filter on that". ## What it does not do A 1x1 convolution does **not** change the spatial receptive field. If the units feeding it saw a 7x7 window of the input, the units after it still see a 7x7 window. It does not change height or width either (at stride 1). It cannot correct spatial misalignment or aggregate context. Anything requiring a wider view still needs a spatial kernel, a stride, a pooling step, or dilation. ## Relation to a fully connected layer A 1x1 convolution *is* a fully connected layer over the channel axis, applied independently at every position with shared weights. The distinction from a dense layer on a flattened feature map matters: the dense layer has `C_in * H * W * C_out` weights, ties the model to one input resolution, and destroys the spatial map. The 1x1 has `C_in * C_out` weights, is resolution-agnostic, and preserves the map for later layers. ## Where you see them Channel-reduction stems in multi-branch modules; the pointwise half of a separable block, where a `C_in x C_out` 1x1 typically holds well over ninety percent of the block's weights; projection shortcuts that need to match channel counts before an addition; and the final classification head that maps a pooled feature vector to logits. A common interview trap is to call a 1x1 layer "pointless because it does nothing spatial" — the whole point is that it operates on the axis a spatial kernel treats as fixed.

  • How is a 1x1 convolution different from a fully connected layer on the feature map?
    A 1x1 convolution is a fully connected layer over the channel axis only, with the same weights reused at every spatial position. That gives it `C_in * C_out` weights, independence from input resolution, and a preserved spatial map. A dense layer on the flattened map has `C_in * H * W * C_out` weights, is locked to one resolution, and throws the spatial structure away.
  • Does inserting a 1x1 layer change what an output unit can see in the input image?
    No. The spatial receptive field is unchanged, because the kernel spans a single position. The layer adds a learned channel mix and, with the activation after it, extra nonlinearity — but if you need a wider view of the image you still need a larger kernel, a stride, pooling, or dilation.
  • If 1x1 layers are so cheap, why can they still dominate a model's parameter count?
    Cheap is relative to `k * k`, not absolute. A 1x1 layer still costs `C_in * C_out` weights, which grows quadratically with width. In a network whose spatial filters are per-channel and whose widths are large, the 1x1 layers can carry the large majority of the weights and of the multiply-accumulates despite the trivial kernel.

A 1x1 convolution is a colour-correction matrix applied identically to every pixel: it re-mixes the channels at each location and never once looks at a neighbour.

saying these in an interview costs you the question

  • Says a 1x1 convolution is an identity or does nothing useful
  • Claims it enlarges the receptive field
  • Thinks it changes the spatial height and width
  • Confuses it with a dense layer on the flattened feature map
  • Assumes it is always negligible in parameter count

context

open as a page

How do semantic, instance, and panoptic segmentation differ in what they label each pixel?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Semantic segmentation labels each pixel with a class but no identity, so touching objects merge. Instance segmentation returns one mask per countable object. Panoptic gives every pixel exactly one class and, for countable classes, an instance id.

open as a page

In object detection, how does IoU between a predicted and ground-truth box decide a detection is correct?

level: juniorimportance: must knowfreq 78%

basics

~20 s

IoU is the overlap area of a predicted and a ground-truth box divided by the area they jointly cover. A prediction counts as correct when its IoU with an unclaimed ground-truth box of the same class clears a threshold, commonly 0.5.

open as a page

In a CNN, what does one 3x3 filter over a 4-channel input actually contain?

level: juniorimportance: must knowfreq 76%

basics

~10 s

A 3x3 filter over a 4-channel input is a 3x3x4 weight block plus one scalar bias. The kernel always spans every input channel, sums across them, and produces a single output channel.

open as a page

Why does a convolutional layer share one kernel across all positions instead of using per-position weights?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A convolution reuses one small kernel at every position, so it learns a feature detector once instead of relearning it at each pixel. That cuts the parameter count enormously and lets the same feature be found anywhere in the input.

open as a page

What is the difference between max pooling and average pooling over a CNN feature map?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Max pooling keeps the largest activation in each window, so a sparse, high-contrast response survives the downsample. Average pooling keeps the window mean, so it preserves overall texture and intensity but dilutes an isolated strong response.

open as a page

Why does a residual block add its input back to its output instead of just stacking layers?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A residual block computes F(x) + x, so its layers only learn the change to make to the input. Behaving like an identity then just means pushing F toward zero, which plain stacked layers struggle to fit.

open as a page

In semantic segmentation, how does a CNN classifier become a dense per-pixel predictor?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Drop the flatten-and-dense head and make every layer convolutional. The dense head is rewritten as an equivalent convolution, the last layer emits one score map per class, and a decoder upsamples those maps back to the input resolution.

open as a page

How do you compute a 2D convolution layer's output height and width?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Each spatial axis independently gives floor((n + 2p - k)/s) + 1: n is that axis's input size, p the padding per side, k the kernel size, s the stride. The floor drops any partial final window.

open as a page

In a 1D convolution over a 6-channel 50 Hz sensor stream, what does the kernel slide over?

level: juniorimportance: must knowfreq 65%

basics

~20 s

A 1D kernel slides only along time. Each filter carries weights for every input channel at every offset in its width, so one kernel position sums k times 6 numbers into a single output value.

open as a page

Why does VGG stack two 3x3 convolutions instead of using one 5x5 layer?

level: middleimportance: must knowfreq 70%

basics

~10 s

Two stacked 3x3 layers see the same 5x5 input patch as one 5x5 layer, but use 18C-squared weights instead of 25C-squared and apply two nonlinearities instead of one. Same reach, cheaper, more expressive.

open as a page

Why is a depthwise-separable convolution roughly nine times cheaper than a dense 3x3 layer?

level: middleimportance: must knowfreq 66%

basics

~20 s

It factorises a dense layer into a spatial stage with one k x k filter per input channel and a pointwise stage that mixes channels. Cost falls to 1/C_out + 1/(k*k) of dense, about 1/9 at k=3.

open as a page

Why does RoIAlign produce better instance masks than RoI pooling on a detection head?

level: middleimportance: must knowfreq 58%

basics

~20 s

RoI pooling rounds twice, the box onto the feature grid and then the bin edges, so features come from up to half a cell off. RoIAlign keeps float coordinates and samples bilinearly, preserving the alignment masks need.

open as a page

How does GIoU repair the 1 - IoU box loss when the predicted and ground-truth boxes do not overlap?

level: middleimportance: must knowfreq 60%

basics

~20 s

Two disjoint boxes have IoU 0 however far apart they sit, so a 1 - IoU loss is flat there and produces no gradient. GIoU subtracts the empty fraction of the smallest box enclosing both, which keeps shrinking as they approach, restoring a direction to move.

open as a page

What does raising a convolution's stride from 1 to 2 change, and what does it cost?

level: middleimportance: must knowfreq 64%

basics

~20 s

Stride is the step between successive kernel placements. Stride 2 evaluates the filter at every other position, so each spatial dimension comes out roughly halved and compute drops about fourfold. Parameters are unchanged; spatial precision is lost.

open as a page

In an object detector, what do anchor boxes do, and how does an anchor-free head replace them?

level: middleimportance: must knowfreq 72%

basics

~20 s

Anchors are fixed reference boxes tiled over every feature-map location; the head predicts a class score and offsets that nudge an anchor onto an object. Anchor-free heads drop them and regress the four side distances straight from each location.

open as a page

In a CNN, what is the difference between translation equivariance and translation invariance?

level: middleimportance: must knowfreq 67%

basics

~20 s

Equivariance means shifting the input shifts the feature map by a corresponding amount. Invariance means the output does not change at all. Stacked convolutions give you equivariance; invariance has to be added by the readout or learned from data.

open as a page

Why does U-Net concatenate encoder feature maps into its decoder instead of upsampling alone?

level: middleimportance: must knowfreq 70%

basics

~20 s

Downsampling in the encoder destroys the precise location of edges, and upsampling cannot invent it back. U-Net concatenates the matching high-resolution encoder maps into each decoder stage, so the decoder combines deep semantics with exact boundary detail.

open as a page

Why must a causal 1D convolution pad only on the left of the sequence?

level: middleimportance: must knowfreq 52%

basics

~20 s

Causality means output at step t may use only inputs up to t. Padding (k-1)*d zeros on the left, then a valid convolution, preserves length and enforces that. Symmetric padding centres the kernel on t, so it reads the future.

open as a page

Why does class-wise non-maximum suppression delete correct boxes on a shelf of densely packed cartons?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Suppression assumes heavily overlapping boxes of the same class are duplicates of one object. Identical cartons packed side by side genuinely overlap above the threshold, so the lower-scoring box is deleted even though it is a real, distinct object.

open as a page

A ten-layer stack of stride-1 3x3 convolutions misses a 300-pixel lesion; what is its receptive field and how do you grow it?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Each stride-1 3x3 layer widens the receptive field by two pixels, so ten of them see just 21 input pixels, far too few for a 300-pixel lesion. Grow it by downsampling, by adding depth, or with larger kernels.

open as a page

In VGG-16, which layers hold most of the 138 million parameters, and why?

level: juniorimportance: should knowfreq 45%

basics

~20 s

About 90% sit in the three fully connected layers at the end. The first alone, mapping the flattened 7x7x512 feature map to 4096 units, holds roughly 103 million weights; all 13 convolution layers together hold under 15 million.

open as a page

How does splitting a convolution into g groups change its cost and its channel connectivity?

level: middleimportance: should knowfreq 50%

basics

~20 s

Grouping partitions input and output channels into g disjoint sets, each output group seeing only its own input group. Weights and multiply-accumulates both drop by a factor of g, but no channel combines across group boundaries.

open as a page

Why does a per-RoI mask branch predict one binary mask per class instead of a per-pixel softmax over classes?

level: middleimportance: should knowfreq 48%

basics

~20 s

Decoupling. The classification branch already names the class, so each mask channel only has to answer whether a pixel is inside the object. A per-pixel softmax would make classes compete for pixels and tie mask quality to classification confidence.

open as a page

In a CNN, what is the real cost of choosing valid padding over same zero padding?

level: middleimportance: should knowfreq 52%

basics

~20 s

Valid padding computes outputs only where the kernel fits inside real data, so the map shrinks at every layer and a border ring gets no output position. Same zero padding preserves size but feeds invented zeros into border windows.

open as a page

Why do CNN classifiers use global average pooling instead of flatten plus a dense layer?

level: middleimportance: should knowfreq 60%

basics

~20 s

Global average pooling reduces each channel's spatial map to its mean, so the classifier sees one number per channel. It adds no parameters, accepts any input size, and removes the huge flatten-to-dense layer that held most of a network's weights.

open as a page

Why does ResNet-50 use a 1x1-3x3-1x1 bottleneck block instead of two 3x3 layers?

level: middleimportance: should knowfreq 52%

basics

~20 s

A 3x3 convolution's cost grows with input channels times output channels, so it is ruinous at wide layers. The bottleneck squeezes width down with a 1x1, runs the 3x3 cheaply, then restores width — roughly an order of magnitude fewer weights.

open as a page

In a ResNet, when can a shortcut be a plain identity and when must it be a projection?

level: middleimportance: should knowfreq 58%

basics

~20 s

A ResNet shortcut can stay a plain identity only while the branch preserves spatial size and channel count, since the two are added elementwise. Where a stage strides down and widens, it becomes a 1x1 convolution.

open as a page

For a 3x3 convolution with 256 input and 256 output channels, how many parameters and multiply-adds does it cost on a 56x56 map?

level: middleimportance: should knowfreq 58%

basics

~10 s

Parameters are 33256256 + 256 = 590,080, independent of the feature-map size. Multiply-adds are 589,824 weights evaluated at every one of the 5656 output positions, about 1.85 billion.

open as a page

How do you compute a dilated temporal convolution stack's receptive field in time steps?

level: middleimportance: should knowfreq 44%

basics

~20 s

Each layer adds (k-1)*d input steps, so the receptive field is 1 plus the sum of (k-1)*d over the layers. Width 3 with dilations 1, 2, 4, 8, 16 covers 63 steps - divide by the sampling rate for seconds.

open as a page

showing 1–30 of 46