skip to content

In a CNN, what does one 3x3 filter over a 4-channel input actually contain?

level: juniorimportance: must knowfreq 76%

answer

  1. the 3x3 is only two of three dimensions
  2. kernel depth follows the input
  3. channels are summed, not kept
  4. one bias belongs to the whole filter
  5. filters counted equals output channels

basics

~10 s

A 3x3 filter over a 4-channel input is a 3x3x4 weight block plus one scalar bias. The kernel always spans every input channel, sums across them, and produces a single output channel.

solid answer

~40 s

The `3x3` in a layer spec names only the spatial window; the kernel's third dimension is fixed by the data, equal to the number of input channels. On a 4-channel Sentinel-2 tile (red, green, blue, near-infrared) each 3x3 filter is really 3x3x4 = 36 weights. At every position it multiplies elementwise with the aligned 3x3x4 sub-volume, sums all 36 products into one scalar, and adds **one** bias that belongs to the filter, not to each channel. That scalar is one pixel of one output channel, so the number of filters is the number of output channels. A layer with 64 filters of size 3x3 over a 32-channel input therefore holds 64 x 32 x 3 x 3 weights and 64 biases. Channels are collapsed by the sum, never carried through untouched.

code

python · 19 lines
python
import random
random.seed(0)

# 4-channel patch (R, G, B, near-infrared), 3x3 each
patch  = [[[random.random() for _ in range(3)] for _ in range(3)] for _ in range(4)]
# the filter is 3x3x4 -- its depth is fixed by the input, not chosen
kernel = [[[random.random() for _ in range(3)] for _ in range(3)] for _ in range(4)]
bias   = 0.5   # ONE bias for the whole filter

acc = 0.0
for c in range(4):          # sum ACROSS input channels
    for i in range(3):
        for j in range(3):
            acc += patch[c][i][j] * kernel[c][i][j]
out = acc + bias            # one scalar -> one pixel of one output channel

n_weights = len([w for plane in kernel for row in plane for w in row])
print(n_weights, "weights +", 1, "bias")
print(round(out, 4))

go deeper

for a junior

Be ready to say out loud that a kernel is three-dimensional: spatial size times input channels, plus one bias. Know that the number of filters is the number of output channels.

for a middle

Explain the mechanics: elementwise multiply over the whole sub-volume, one sum across channels and space, one bias, one scalar. Derive a layer's weight count from the spec without hesitating.

for a senior

Show you use this when debugging. A channel-mismatch error is a statement about the previous layer's output, and an unnormalized wide-range channel can dominate the sum at initialization.

for a principal

Own the width decision: filter count sets both this layer's parameter cost and the next layer's kernel depth, so widening compounds. Be able to argue where channel capacity actually earns its memory.

## What "3x3" actually names A convolution layer is usually written down with two numbers: a kernel size and a filter count, for example "3x3, 64". The kernel size describes only the **spatial** extent of the sliding window - how tall and how wide it is. The window's third dimension is never written because it is not a free choice. A kernel's depth always equals the number of channels in the tensor it is applied to. So if the layer's input is a 4-channel Sentinel-2 satellite tile - red, green, blue and near-infrared stacked on top of each other - then each "3x3" filter is really a 3x3x4 block of learned weights: 36 numbers, not 9. ## One filter, one bias, one output channel The operation at a single output position is: 1. Align the 3x3x4 filter with the 3x3x4 sub-volume of the input centred at that position. 2. Multiply elementwise, giving 36 products. 3. Sum **all 36** into a single scalar. The channel axis is summed over, exactly like the spatial axes. 4. Add the filter's bias - one scalar for the whole filter. The result is one number. Slide the filter over every valid position and you get one 2D map: one output channel. This is why the filter count and the output channel count are the same number. "3x3, 64" over a 32-channel input means 64 separate kernels, each of shape 32x3x3, so `64 * 32 * 3 * 3 = 18,432` weights, plus 64 biases - one per filter. ## Why the bias is per filter and not per channel A frequent guess is that each input channel gets its own bias. It cannot help. The bias is added *after* the channel sum, so if you gave each of the 4 input channels its own constant, their contributions would add up to a single constant anyway - identical in effect to one bias, but with redundant parameters that the optimizer cannot distinguish. One bias per filter is both sufficient and non-degenerate. A related point: bias terms are optional, and layers that are immediately followed by a normalization layer are often built without them, because the normalization subtracts a mean and makes the bias unidentifiable. When the bias is present, there is exactly one per output channel. ## Why the channels are summed rather than kept apart Summing across channels is what lets a single filter encode a **cross-channel** pattern. Vegetation on a satellite tile is "bright in near-infrared and comparatively dark in red". A filter can express that directly with positive weights on the near-infrared plane and negative weights on the red plane; the sum then peaks exactly where that signature holds. If channels were processed independently and stacked, no single unit could represent that conjunction, and you would need a separate mixing stage to recover it. The same logic explains why input channel order does not matter as long as it is consistent: the filter simply learns which plane is which. It also explains why the channels must be scaled sensibly before training. A near-infrared band with a numeric range ten times wider than the visible bands dominates the sum at initialization, so per-channel normalization matters more than people expect on non-photographic data. ## Reading a layer specification Three quantities in a spec are easy to conflate: - **Kernel size** - the spatial window, e.g. 3x3. It says nothing about depth. - **Input channels** - fixed by whatever produced the input. It sets the kernel depth, and it is the one number a mismatched shape error is usually complaining about. - **Filters (output channels)** - a design choice, the width of the layer. Doubling it doubles the layer's weights and its output channel count. A useful mental check when a layer will not build: the kernel depth is not yours to pick, so a channel mismatch is always a claim about what the previous layer emitted. ## What interviewers are testing This question separates candidates who have only read layer specifications from those who have reasoned about the operation. The tells are specific: saying a 3x3 filter has 9 weights regardless of input depth, expecting one bias per input channel, or describing the layer as "applying the filter to each channel and stacking the results" - which is a different operation entirely, not the standard convolution. Getting the 3x3x4-plus-one-bias picture right also makes every downstream shape and parameter discussion trivial.

  • Why does the filter carry one bias rather than one bias per input channel?
    The bias is added after the channel sum, so per-channel constants would simply add up to a single constant. They would be mathematically indistinguishable from one bias while adding redundant parameters the optimizer cannot pin down. One scalar per filter is exactly the right amount of freedom.
  • A layer applies 64 filters of size 3x3 to a 32-channel input. How many learnable numbers does it hold?
    Each filter is 32x3x3 = 288 weights, and there are 64 of them, so 18,432 weights, plus 64 biases - one per filter - for 18,496 learnable numbers in total. Note that the kernel size contributes only the 3x3; the 32 comes from the input, which is why widening the previous layer inflates this one.
  • Do the channels stacked into one input have to be the same kind of measurement?
    No - red, green, blue and near-infrared are different physical quantities and stacking them is exactly the point, since a filter can learn a signature that spans them. But they must be scaled comparably. A band with a much wider numeric range dominates the channel sum at initialization, so per-channel standardization matters more on scientific imagery than on ordinary photographs.

Think of the filter as a stencil cut through a stack of transparencies: it covers the same little square on every sheet at once, and you read a single number off the whole stack, not one per sheet.

saying these in an interview costs you the question

  • Says a 3x3 filter has nine weights regardless of input channels
  • Expects one bias per input channel instead of per filter
  • Describes the filter as applied to each channel separately
  • Confuses the filter count with the kernel's spatial size
  • Thinks kernel depth is a hyperparameter you choose

context