skip to content

The Convolution Operator

The convolution itself: a small learned kernel slid across the input with a stride and padding, summing over input channels. Interviewers start here because shape errors surface immediately.

on this pageshow

explore

questions

18

In a CNN, what does one 3x3 filter over a 4-channel input actually contain?

level: juniorimportance: must knowfreq 76%

answer

  1. the 3x3 is only two of three dimensions
  2. kernel depth follows the input
  3. channels are summed, not kept
  4. one bias belongs to the whole filter
  5. filters counted equals output channels

basics

~10 s

A 3x3 filter over a 4-channel input is a 3x3x4 weight block plus one scalar bias. The kernel always spans every input channel, sums across them, and produces a single output channel.

solid answer

~40 s

The `3x3` in a layer spec names only the spatial window; the kernel's third dimension is fixed by the data, equal to the number of input channels. On a 4-channel Sentinel-2 tile (red, green, blue, near-infrared) each 3x3 filter is really 3x3x4 = 36 weights. At every position it multiplies elementwise with the aligned 3x3x4 sub-volume, sums all 36 products into one scalar, and adds **one** bias that belongs to the filter, not to each channel. That scalar is one pixel of one output channel, so the number of filters is the number of output channels. A layer with 64 filters of size 3x3 over a 32-channel input therefore holds 64 x 32 x 3 x 3 weights and 64 biases. Channels are collapsed by the sum, never carried through untouched.

code

python · 19 lines
python
import random
random.seed(0)

# 4-channel patch (R, G, B, near-infrared), 3x3 each
patch  = [[[random.random() for _ in range(3)] for _ in range(3)] for _ in range(4)]
# the filter is 3x3x4 -- its depth is fixed by the input, not chosen
kernel = [[[random.random() for _ in range(3)] for _ in range(3)] for _ in range(4)]
bias   = 0.5   # ONE bias for the whole filter

acc = 0.0
for c in range(4):          # sum ACROSS input channels
    for i in range(3):
        for j in range(3):
            acc += patch[c][i][j] * kernel[c][i][j]
out = acc + bias            # one scalar -> one pixel of one output channel

n_weights = len([w for plane in kernel for row in plane for w in row])
print(n_weights, "weights +", 1, "bias")
print(round(out, 4))

go deeper

for a junior

Be ready to say out loud that a kernel is three-dimensional: spatial size times input channels, plus one bias. Know that the number of filters is the number of output channels.

for a middle

Explain the mechanics: elementwise multiply over the whole sub-volume, one sum across channels and space, one bias, one scalar. Derive a layer's weight count from the spec without hesitating.

for a senior

Show you use this when debugging. A channel-mismatch error is a statement about the previous layer's output, and an unnormalized wide-range channel can dominate the sum at initialization.

for a principal

Own the width decision: filter count sets both this layer's parameter cost and the next layer's kernel depth, so widening compounds. Be able to argue where channel capacity actually earns its memory.

## What "3x3" actually names A convolution layer is usually written down with two numbers: a kernel size and a filter count, for example "3x3, 64". The kernel size describes only the **spatial** extent of the sliding window - how tall and how wide it is. The window's third dimension is never written because it is not a free choice. A kernel's depth always equals the number of channels in the tensor it is applied to. So if the layer's input is a 4-channel Sentinel-2 satellite tile - red, green, blue and near-infrared stacked on top of each other - then each "3x3" filter is really a 3x3x4 block of learned weights: 36 numbers, not 9. ## One filter, one bias, one output channel The operation at a single output position is: 1. Align the 3x3x4 filter with the 3x3x4 sub-volume of the input centred at that position. 2. Multiply elementwise, giving 36 products. 3. Sum **all 36** into a single scalar. The channel axis is summed over, exactly like the spatial axes. 4. Add the filter's bias - one scalar for the whole filter. The result is one number. Slide the filter over every valid position and you get one 2D map: one output channel. This is why the filter count and the output channel count are the same number. "3x3, 64" over a 32-channel input means 64 separate kernels, each of shape 32x3x3, so `64 * 32 * 3 * 3 = 18,432` weights, plus 64 biases - one per filter. ## Why the bias is per filter and not per channel A frequent guess is that each input channel gets its own bias. It cannot help. The bias is added *after* the channel sum, so if you gave each of the 4 input channels its own constant, their contributions would add up to a single constant anyway - identical in effect to one bias, but with redundant parameters that the optimizer cannot distinguish. One bias per filter is both sufficient and non-degenerate. A related point: bias terms are optional, and layers that are immediately followed by a normalization layer are often built without them, because the normalization subtracts a mean and makes the bias unidentifiable. When the bias is present, there is exactly one per output channel. ## Why the channels are summed rather than kept apart Summing across channels is what lets a single filter encode a **cross-channel** pattern. Vegetation on a satellite tile is "bright in near-infrared and comparatively dark in red". A filter can express that directly with positive weights on the near-infrared plane and negative weights on the red plane; the sum then peaks exactly where that signature holds. If channels were processed independently and stacked, no single unit could represent that conjunction, and you would need a separate mixing stage to recover it. The same logic explains why input channel order does not matter as long as it is consistent: the filter simply learns which plane is which. It also explains why the channels must be scaled sensibly before training. A near-infrared band with a numeric range ten times wider than the visible bands dominates the sum at initialization, so per-channel normalization matters more than people expect on non-photographic data. ## Reading a layer specification Three quantities in a spec are easy to conflate: - **Kernel size** - the spatial window, e.g. 3x3. It says nothing about depth. - **Input channels** - fixed by whatever produced the input. It sets the kernel depth, and it is the one number a mismatched shape error is usually complaining about. - **Filters (output channels)** - a design choice, the width of the layer. Doubling it doubles the layer's weights and its output channel count. A useful mental check when a layer will not build: the kernel depth is not yours to pick, so a channel mismatch is always a claim about what the previous layer emitted. ## What interviewers are testing This question separates candidates who have only read layer specifications from those who have reasoned about the operation. The tells are specific: saying a 3x3 filter has 9 weights regardless of input depth, expecting one bias per input channel, or describing the layer as "applying the filter to each channel and stacking the results" - which is a different operation entirely, not the standard convolution. Getting the 3x3x4-plus-one-bias picture right also makes every downstream shape and parameter discussion trivial.

  • Why does the filter carry one bias rather than one bias per input channel?
    The bias is added after the channel sum, so per-channel constants would simply add up to a single constant. They would be mathematically indistinguishable from one bias while adding redundant parameters the optimizer cannot pin down. One scalar per filter is exactly the right amount of freedom.
  • A layer applies 64 filters of size 3x3 to a 32-channel input. How many learnable numbers does it hold?
    Each filter is 32x3x3 = 288 weights, and there are 64 of them, so 18,432 weights, plus 64 biases - one per filter - for 18,496 learnable numbers in total. Note that the kernel size contributes only the 3x3; the 32 comes from the input, which is why widening the previous layer inflates this one.
  • Do the channels stacked into one input have to be the same kind of measurement?
    No - red, green, blue and near-infrared are different physical quantities and stacking them is exactly the point, since a filter can learn a signature that spans them. But they must be scaled comparably. A band with a much wider numeric range dominates the channel sum at initialization, so per-channel standardization matters more on scientific imagery than on ordinary photographs.

Think of the filter as a stencil cut through a stack of transparencies: it covers the same little square on every sheet at once, and you read a single number off the whole stack, not one per sheet.

saying these in an interview costs you the question

  • Says a 3x3 filter has nine weights regardless of input channels
  • Expects one bias per input channel instead of per filter
  • Describes the filter as applied to each channel separately
  • Confuses the filter count with the kernel's spatial size
  • Thinks kernel depth is a hyperparameter you choose

context

open as a page

Why does a convolutional layer share one kernel across all positions instead of using per-position weights?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A convolution reuses one small kernel at every position, so it learns a feature detector once instead of relearning it at each pixel. That cuts the parameter count enormously and lets the same feature be found anywhere in the input.

open as a page

What is the difference between max pooling and average pooling over a CNN feature map?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Max pooling keeps the largest activation in each window, so a sparse, high-contrast response survives the downsample. Average pooling keeps the window mean, so it preserves overall texture and intensity but dilutes an isolated strong response.

open as a page

How do you compute a 2D convolution layer's output height and width?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Each spatial axis independently gives floor((n + 2p - k)/s) + 1: n is that axis's input size, p the padding per side, k the kernel size, s the stride. The floor drops any partial final window.

open as a page

In a 1D convolution over a 6-channel 50 Hz sensor stream, what does the kernel slide over?

level: juniorimportance: must knowfreq 65%

basics

~20 s

A 1D kernel slides only along time. Each filter carries weights for every input channel at every offset in its width, so one kernel position sums k times 6 numbers into a single output value.

open as a page

What does raising a convolution's stride from 1 to 2 change, and what does it cost?

level: middleimportance: must knowfreq 64%

basics

~20 s

Stride is the step between successive kernel placements. Stride 2 evaluates the filter at every other position, so each spatial dimension comes out roughly halved and compute drops about fourfold. Parameters are unchanged; spatial precision is lost.

open as a page

In a CNN, what is the difference between translation equivariance and translation invariance?

level: middleimportance: must knowfreq 67%

basics

~20 s

Equivariance means shifting the input shifts the feature map by a corresponding amount. Invariance means the output does not change at all. Stacked convolutions give you equivariance; invariance has to be added by the readout or learned from data.

open as a page

Why must a causal 1D convolution pad only on the left of the sequence?

level: middleimportance: must knowfreq 52%

basics

~20 s

Causality means output at step t may use only inputs up to t. Padding (k-1)*d zeros on the left, then a valid convolution, preserves length and enforces that. Symmetric padding centres the kernel on t, so it reads the future.

open as a page

A ten-layer stack of stride-1 3x3 convolutions misses a 300-pixel lesion; what is its receptive field and how do you grow it?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Each stride-1 3x3 layer widens the receptive field by two pixels, so ten of them see just 21 input pixels, far too few for a 300-pixel lesion. Grow it by downsampling, by adding depth, or with larger kernels.

open as a page

In a CNN, what is the real cost of choosing valid padding over same zero padding?

level: middleimportance: should knowfreq 52%

basics

~20 s

Valid padding computes outputs only where the kernel fits inside real data, so the map shrinks at every layer and a border ring gets no output position. Same zero padding preserves size but feeds invented zeros into border windows.

open as a page

Why do CNN classifiers use global average pooling instead of flatten plus a dense layer?

level: middleimportance: should knowfreq 60%

basics

~20 s

Global average pooling reduces each channel's spatial map to its mean, so the classifier sees one number per channel. It adds no parameters, accepts any input size, and removes the huge flatten-to-dense layer that held most of a network's weights.

open as a page

For a 3x3 convolution with 256 input and 256 output channels, how many parameters and multiply-adds does it cost on a 56x56 map?

level: middleimportance: should knowfreq 58%

basics

~10 s

Parameters are 33256256 + 256 = 590,080, independent of the feature-map size. Multiply-adds are 589,824 weights evaluated at every one of the 5656 output positions, about 1.85 billion.

open as a page

How do you compute a dilated temporal convolution stack's receptive field in time steps?

level: middleimportance: should knowfreq 44%

basics

~20 s

Each layer adds (k-1)*d input steps, so the receptive field is 1 plus the sum of (k-1)*d over the layers. Width 3 with dilations 1, 2, 4, 8, 16 covers 63 steps - divide by the sampling rate for seconds.

open as a page

When is a convolution's shared-weight assumption wrong for the data you are modelling?

level: seniorimportance: should knowfreq 41%

basics

~10 s

Sharing asserts that the same local pattern means the same thing everywhere along an axis. It fails where position itself carries meaning: a spectrogram's frequency axis, absolute-coordinate targets, and tabular columns with no order.

open as a page

A 300x300 image through five stride-2 stages breaks a decoder's skip concatenation — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

300 is not divisible by 32. The floor in the output-shape formula drops a row at every odd-sized stage, so the encoder produces 37 where the decoder's doubling produces 36, and concatenation needs identical height and width.

open as a page

What does a dilated (atrous) convolution buy you, and what does it cost?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Dilation spreads a kernel's taps apart, covering a wider region with no extra weights and no downsampling. The costs are sparse sampling between the taps, full-resolution activation memory, and gridding when layers repeat one rate.

open as a page

When would you replace every 2x2 max pooling layer with a stride-2 convolution, and what does it cost?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Replace it when the downsample itself should be learned: a stride-2 convolution trains a weighted summary per window instead of applying a fixed maximum. It costs parameters, multiply-adds and a useful prior, and pays off most when data is plentiful.

open as a page

Why is a 3D convolution over a 16-frame video clip usually factorized into (2+1)D?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A 3D kernel spans time, height and width at once, so weights and compute scale with clip length too. The (2+1)D form splits it into a spatial convolution then a temporal one: fewer weights, an extra nonlinearity, easier optimization.

open as a page