skip to content

Output Shape Arithmetic

Deriving output height and width from input size, kernel, stride, padding and dilation, and counting a layer's weights. Interviewers use this arithmetic as a fast competence check.

on this pageshow

questions

3

How do you compute a 2D convolution layer's output height and width?

level: juniorimportance: must knowfreq 80%

answer

  1. one axis at a time
  2. count window start positions
  3. padding is added on both sides
  4. floor, never round up
  5. steps of stride, plus one

basics

~20 s

Each spatial axis independently gives floor((n + 2p - k)/s) + 1: n is that axis's input size, p the padding per side, k the kernel size, s the stride. The floor drops any partial final window.

solid answer

~50 s

Treat the axes separately: for one axis, `out = floor((n + 2p - k)/s) + 1`. The padded length is `n + 2p`, the last window can start no later than `n + 2p - k`, window starts are spaced `s` apart, and the `+ 1` counts the first start — the floor throws away a trailing window that would run off the edge. Valid padding is `p = 0`, which at stride 1 shrinks the map by `k - 1`; same padding at stride 1 means `p = (k - 1)/2`, which is why odd kernels like 3x3 and 5x5 are the norm. Dilation enters through an effective kernel size `k_eff = k + (k - 1)(d - 1)`, so a dilated kernel shrinks the map more without holding more weights. The formula says nothing about channels: output channels are just the number of filters.

go deeper

for a junior

Be ready to compute the output height and width on the spot for a padded 3x3 stride-1 layer, and to state that padding is counted on both sides of each axis.

for a middle

Explain where the floor and the +1 come from by counting window start positions, and derive p = (k - 1)/2 as the same-padding condition at stride 1.

for a senior

Show that you trace a whole stage chain before training and know which input sizes make a downsampling stack round, since that is where shape bugs are born.

for a principal

Own the convention for the team: decide whether models accept only divisible input sizes or handle arbitrary ones by padding and cropping, and write that contract down.

## What the formula counts A 2D convolution slides a `k x k` window over the (optionally padded) input and emits one number per valid window position. So the output size on an axis is just **how many window positions fit**, and that is a counting argument, not a division. On one axis, let `n` be the input size, `p` the zero padding added to *each* side, `k` the kernel size and `s` the stride. The padded axis has length `n + 2p`. A window occupying positions `[i, i + k - 1]` fits as long as `i <= n + 2p - k`. Valid starts are `0, s, 2s, ...`, so the count is ``` out = floor((n + 2p - k) / s) + 1 ``` The `+ 1` is the fence-post term for the window starting at 0. The `floor` is what happens when `(n + 2p - k)` is not a multiple of `s`: the leftover input at the far edge is simply never covered by any window and is silently dropped. Height and width are computed independently with their own `n`, `p`, `k`, `s` — nothing couples them. The formula is also silent about channels. Each filter spans **all** input channels, and the number of output channels equals the number of filters. Depth is a property of the layer definition, not of the sliding-window arithmetic. ## Valid and same padding **Valid** means `p = 0`. At stride 1 this gives `out = n - k + 1`, so every layer shaves `k - 1` off each axis. Stack ten 3x3 valid convolutions and you have lost 20 pixels per axis, which is why deep valid-padded stacks need generous input margins. **Same** at stride 1 means you want `out = n`. Setting `n = n + 2p - k + 1` gives `p = (k - 1)/2`. That is an integer only for **odd** `k`, which is the practical reason 3x3 (p = 1), 5x5 (p = 2) and 7x7 (p = 3) dominate over 2x2 and 4x4. At stride greater than 1, "same" is redefined as `out = ceil(n / s)`. The total padding required is `max((out - 1) * s + k - n, 0)`, split between the two sides. When that total is odd — for instance a 4x4 kernel at stride 1 needs a total of 3 — it cannot be split symmetrically, so one side receives the extra row or column and the output grid sits half a pixel off centre. The same asymmetry appears with an even stride when `n` is not a multiple of `s`. ## Dilation Dilation `d` inserts `d - 1` gaps between kernel taps. The window still holds `k x k` weights, but it *spans* ``` k_eff = k + (k - 1) * (d - 1) ``` input positions. A 3x3 kernel at dilation 4 has `k_eff = 3 + 2 * 3 = 9`: it reads from a 9x9 footprint while still costing nine weights per input channel. Substituting `k_eff` for `k` gives the general form `out = floor((n + 2p - k_eff)/s) + 1`, and preserving the size at stride 1 now needs `p = 4`, not `p = 1`. Forgetting this is the usual cause of a map that unexpectedly shrinks by eight pixels. ## Asymmetric strides Because the axes are independent, the stride and kernel may differ per axis. Take an audio input of 128 mel bins by 401 frames, a 3x3 kernel with `p = 1` and a stride of `(2, 1)`. The frequency axis gives `floor((128 + 2 - 3)/2) + 1 = 63 + 1 = 64`, while the time axis gives `floor((401 + 2 - 3)/1) + 1 = 401`. Frequency halves at every stage and time is preserved exactly — a deliberate design when you need frame-level output alignment. ## Working it by hand The reliable habit is to write the chain out before training: input size, then one line per layer with `k`, `s`, `p`, `d` and the resulting size. Two checks catch nearly every mistake. First, is the numerator `n + 2p - k` non-negative? If not, the kernel is larger than the padded input and the layer is invalid. Second, is `(n + 2p - k)` divisible by `s`? If not, the floor is discarding input, which is legal but is exactly where a later shape mismatch is born. Common errors: rounding up instead of down; forgetting that `p` is added on both sides, so the numerator gains `2p` and not `p`; dropping the `+ 1`; applying the formula to the channel axis; and assuming a 3x3 with one pixel of padding preserves size at *any* stride, when it only does so at stride 1.

  • With a 3x3 kernel at dilation rate 4, what replaces k in the formula?
    The effective kernel size `k_eff = k + (k - 1)(d - 1) = 3 + 2 * 3 = 9`. The window reads from a 9x9 footprint while still holding only nine weights per input channel, so the layer shrinks the map as if it were a 9x9 kernel. To preserve the size at stride 1 you need `p = 4` rather than `p = 1`.
  • What happens to same padding when the kernel size is even?
    The total padding needed at stride 1 is `k - 1`, which is odd for an even kernel — a 4x4 kernel needs 3. It cannot be split evenly, so one side gets 1 and the other 2, and the output grid is offset by half a pixel relative to the input. That asymmetry is a real reason even kernels are avoided in practice.
  • How does an asymmetric stride such as (2, 1) change the two axes?
    Each axis uses its own stride in the same formula. On a 128-by-401 mel input with a 3x3 kernel and `p = 1`, the frequency axis becomes `floor((128 + 2 - 3)/2) + 1 = 64` while the time axis stays at 401. Frequency is downsampled and time is left intact, which keeps outputs aligned with input frames.
  • Does the formula tell you anything about the output channel count?
    No. It governs the spatial axes only. The number of output channels equals the number of filters in the layer, chosen by the designer, and each filter already sums over every input channel. Confusing the channel axis with a spatial axis is a frequent slip when reading a shape error.

Counting output positions is a fence-post problem, not a fence-panel one: the number of window starts is the number of stride steps plus one for the very first window.

saying these in an interview costs you the question

  • Says the output size is just n divided by the stride
  • Rounds the division up instead of flooring it
  • Forgets the +1 for the first window position
  • Thinks padding is added once rather than to both sides
  • Applies the spatial formula to the channel dimension
  • Assumes 3x3 with one pixel of padding preserves size at any stride

context

open as a page

For a 3x3 convolution with 256 input and 256 output channels, how many parameters and multiply-adds does it cost on a 56x56 map?

level: middleimportance: should knowfreq 58%

basics

~10 s

Parameters are 33256256 + 256 = 590,080, independent of the feature-map size. Multiply-adds are 589,824 weights evaluated at every one of the 5656 output positions, about 1.85 billion.

open as a page

A 300x300 image through five stride-2 stages breaks a decoder's skip concatenation — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

300 is not divisible by 32. The floor in the output-shape formula drops a row at every odd-sized stage, so the encoder produces 37 where the decoder's doubling produces 36, and concatenation needs identical height and width.

open as a page