In a CNN, what does one 3x3 filter over a 4-channel input actually contain?
answer
- the 3x3 is only two of three dimensions
- kernel depth follows the input
- channels are summed, not kept
- one bias belongs to the whole filter
- filters counted equals output channels
basics
~10 sA 3x3 filter over a 4-channel input is a 3x3x4 weight block plus one scalar bias. The kernel always spans every input channel, sums across them, and produces a single output channel.
solid answer
~40 sThe `3x3` in a layer spec names only the spatial window; the kernel's third dimension is fixed by the data, equal to the number of input channels. On a 4-channel Sentinel-2 tile (red, green, blue, near-infrared) each 3x3 filter is really 3x3x4 = 36 weights. At every position it multiplies elementwise with the aligned 3x3x4 sub-volume, sums all 36 products into one scalar, and adds **one** bias that belongs to the filter, not to each channel. That scalar is one pixel of one output channel, so the number of filters is the number of output channels. A layer with 64 filters of size 3x3 over a 32-channel input therefore holds 64 x 32 x 3 x 3 weights and 64 biases. Channels are collapsed by the sum, never carried through untouched.
code
python · 19 linesimport random
random.seed(0)
# 4-channel patch (R, G, B, near-infrared), 3x3 each
patch = [[[random.random() for _ in range(3)] for _ in range(3)] for _ in range(4)]
# the filter is 3x3x4 -- its depth is fixed by the input, not chosen
kernel = [[[random.random() for _ in range(3)] for _ in range(3)] for _ in range(4)]
bias = 0.5 # ONE bias for the whole filter
acc = 0.0
for c in range(4): # sum ACROSS input channels
for i in range(3):
for j in range(3):
acc += patch[c][i][j] * kernel[c][i][j]
out = acc + bias # one scalar -> one pixel of one output channel
n_weights = len([w for plane in kernel for row in plane for w in row])
print(n_weights, "weights +", 1, "bias")
print(round(out, 4))go deeper
Be ready to say out loud that a kernel is three-dimensional: spatial size times input channels, plus one bias. Know that the number of filters is the number of output channels.
Explain the mechanics: elementwise multiply over the whole sub-volume, one sum across channels and space, one bias, one scalar. Derive a layer's weight count from the spec without hesitating.
Show you use this when debugging. A channel-mismatch error is a statement about the previous layer's output, and an unnormalized wide-range channel can dominate the sum at initialization.
Own the width decision: filter count sets both this layer's parameter cost and the next layer's kernel depth, so widening compounds. Be able to argue where channel capacity actually earns its memory.
## What "3x3" actually names A convolution layer is usually written down with two numbers: a kernel size and a filter count, for example "3x3, 64". The kernel size describes only the **spatial** extent of the sliding window - how tall and how wide it is. The window's third dimension is never written because it is not a free choice. A kernel's depth always equals the number of channels in the tensor it is applied to. So if the layer's input is a 4-channel Sentinel-2 satellite tile - red, green, blue and near-infrared stacked on top of each other - then each "3x3" filter is really a 3x3x4 block of learned weights: 36 numbers, not 9. ## One filter, one bias, one output channel The operation at a single output position is: 1. Align the 3x3x4 filter with the 3x3x4 sub-volume of the input centred at that position. 2. Multiply elementwise, giving 36 products. 3. Sum **all 36** into a single scalar. The channel axis is summed over, exactly like the spatial axes. 4. Add the filter's bias - one scalar for the whole filter. The result is one number. Slide the filter over every valid position and you get one 2D map: one output channel. This is why the filter count and the output channel count are the same number. "3x3, 64" over a 32-channel input means 64 separate kernels, each of shape 32x3x3, so `64 * 32 * 3 * 3 = 18,432` weights, plus 64 biases - one per filter. ## Why the bias is per filter and not per channel A frequent guess is that each input channel gets its own bias. It cannot help. The bias is added *after* the channel sum, so if you gave each of the 4 input channels its own constant, their contributions would add up to a single constant anyway - identical in effect to one bias, but with redundant parameters that the optimizer cannot distinguish. One bias per filter is both sufficient and non-degenerate. A related point: bias terms are optional, and layers that are immediately followed by a normalization layer are often built without them, because the normalization subtracts a mean and makes the bias unidentifiable. When the bias is present, there is exactly one per output channel. ## Why the channels are summed rather than kept apart Summing across channels is what lets a single filter encode a **cross-channel** pattern. Vegetation on a satellite tile is "bright in near-infrared and comparatively dark in red". A filter can express that directly with positive weights on the near-infrared plane and negative weights on the red plane; the sum then peaks exactly where that signature holds. If channels were processed independently and stacked, no single unit could represent that conjunction, and you would need a separate mixing stage to recover it. The same logic explains why input channel order does not matter as long as it is consistent: the filter simply learns which plane is which. It also explains why the channels must be scaled sensibly before training. A near-infrared band with a numeric range ten times wider than the visible bands dominates the sum at initialization, so per-channel normalization matters more than people expect on non-photographic data. ## Reading a layer specification Three quantities in a spec are easy to conflate: - **Kernel size** - the spatial window, e.g. 3x3. It says nothing about depth. - **Input channels** - fixed by whatever produced the input. It sets the kernel depth, and it is the one number a mismatched shape error is usually complaining about. - **Filters (output channels)** - a design choice, the width of the layer. Doubling it doubles the layer's weights and its output channel count. A useful mental check when a layer will not build: the kernel depth is not yours to pick, so a channel mismatch is always a claim about what the previous layer emitted. ## What interviewers are testing This question separates candidates who have only read layer specifications from those who have reasoned about the operation. The tells are specific: saying a 3x3 filter has 9 weights regardless of input depth, expecting one bias per input channel, or describing the layer as "applying the filter to each channel and stacking the results" - which is a different operation entirely, not the standard convolution. Getting the 3x3x4-plus-one-bias picture right also makes every downstream shape and parameter discussion trivial.
- Why does the filter carry one bias rather than one bias per input channel?The bias is added after the channel sum, so per-channel constants would simply add up to a single constant. They would be mathematically indistinguishable from one bias while adding redundant parameters the optimizer cannot pin down. One scalar per filter is exactly the right amount of freedom.
- A layer applies 64 filters of size 3x3 to a 32-channel input. How many learnable numbers does it hold?Each filter is 32x3x3 = 288 weights, and there are 64 of them, so 18,432 weights, plus 64 biases - one per filter - for 18,496 learnable numbers in total. Note that the kernel size contributes only the 3x3; the 32 comes from the input, which is why widening the previous layer inflates this one.
- Do the channels stacked into one input have to be the same kind of measurement?No - red, green, blue and near-infrared are different physical quantities and stacking them is exactly the point, since a filter can learn a signature that spans them. But they must be scaled comparably. A band with a much wider numeric range dominates the channel sum at initialization, so per-channel standardization matters more on scientific imagery than on ordinary photographs.
Think of the filter as a stencil cut through a stack of transparencies: it covers the same little square on every sheet at once, and you read a single number off the whole stack, not one per sheet.
saying these in an interview costs you the question
- Says a 3x3 filter has nine weights regardless of input channels
- Expects one bias per input channel instead of per filter
- Describes the filter as applied to each channel separately
- Confuses the filter count with the kernel's spatial size
- Thinks kernel depth is a hyperparameter you choose