skip to content

For a 3x3 convolution with 256 input and 256 output channels, how many parameters and multiply-adds does it cost on a 56x56 map?

level: middleimportance: should knowfreq 58%

answer

  1. storage and compute are different questions
  2. one filter spans every input channel
  3. one filter per output channel, plus a bias
  4. weights do not know the resolution
  5. cost repeats at every output position

basics

~10 s

Parameters are 33256256 + 256 = 590,080, independent of the feature-map size. Multiply-adds are 589,824 weights evaluated at every one of the 5656 output positions, about 1.85 billion.

solid answer

~50 s

Parameters come only from the kernel and the channel counts: `params = k_h * k_w * C_in * C_out + C_out`, so `3 * 3 * 256 * 256 = 589,824` weights plus 256 biases gives 590,080. Arithmetic cost adds the spatial axes: each output position performs `k_h * k_w * C_in` multiply-accumulates per output channel, so `MACs = k_h * k_w * C_in * C_out * H_out * W_out = 589,824 * 56 * 56`, roughly 1.85 billion MACs (about 3.7 billion FLOPs if you count the multiply and the add separately). The consequence interviewers are probing for: move the same layer to a 28x28 map and the MACs fall to about 0.46 billion — exactly a quarter, since output positions scale with area — while the parameter count stays at 590,080. Weight sharing is why storage is position-independent and compute is not.

go deeper

for a junior

Recall the shape of a single filter: kernel height times kernel width times input channels, with one such filter per output channel and one bias each.

for a middle

Explain why the parameter count is independent of feature-map size while the multiply-add count is not, and compute both cleanly for a stated layer.

for a senior

Use the arithmetic as a budget: identify which layers in a stack dominate storage and which dominate compute, and pick resolution or channel width accordingly.

for a principal

Set the channel and resolution budget for a family of models and be able to explain to non-specialists which knob trades accuracy against arithmetic cost.

## Two different quantities A convolution layer has two costs that people routinely conflate: how much it **stores** and how much it **computes**. The arithmetic separates them cleanly. ### Parameters One filter produces one output channel and spans the full input depth, so it is a `k_h x k_w x C_in` block of weights. With `C_out` filters and one bias per filter: ``` params = k_h * k_w * C_in * C_out + C_out ``` For the 3x3, 256-in, 256-out layer: `3 * 3 * 256 * 256 = 589,824` weights, plus 256 biases, giving **590,080**. Notice what is absent from the expression: height, width, batch size. Nothing about the input resolution appears, because the same filter is reused at every spatial position. That reuse — weight sharing — is precisely why a convolution's storage is resolution-independent. Biases are a rounding error here: 256 of 590,080 is under 0.05 percent. They are still worth naming when asked for an exact count, and worth knowing they are often omitted entirely when a normalization layer such as BatchNorm follows, because that layer supplies its own per-channel shift. ### Multiply-accumulates Every output value is a dot product over the filter block, which is `k_h * k_w * C_in` multiply-accumulates. There are `C_out * H_out * W_out` output values, so ``` MACs = k_h * k_w * C_in * C_out * H_out * W_out ``` On a 56x56 output map: `589,824 * 3,136 = 1,849,688,064`, about **1.85 billion MACs**. A MAC is one multiply plus one add, so if you report FLOPs with multiplies and adds counted separately the figure roughly doubles to about 3.7 billion. Always state which convention you are using; half the disagreements about "how many FLOPs" are convention mismatches rather than arithmetic errors. ## The resolution lever Now move the identical layer to a 28x28 map. Output positions fall from 3,136 to 784, so ``` MACs = 589,824 * 784 = 462,422,016 ``` about 0.46 billion — **exactly one quarter**, because halving both spatial axes quarters the area. The parameter count is untouched at 590,080. Nothing about the layer changed; only how many times its weights were applied. This gives a compact way to think about the two levers. A convenient identity is ``` MACs = (params - C_out) * H_out * W_out ``` so `H_out * W_out` is literally the **reuse factor**: how many times each weight is used in a forward pass. Doubling channel width roughly quadruples both parameters and MACs (both `C_in` and `C_out` grow); halving resolution quarters MACs and leaves parameters alone. ## Reading a whole stack Run the arithmetic across a typical downsampling backbone and a pattern falls out. Early layers work at high resolution with few channels: small `C_in * C_out`, huge `H_out * W_out`, so few parameters but a lot of arithmetic. Late layers work at low resolution with wide channels: the reverse — many parameters, comparatively little arithmetic. This is why an answer that says "the parameter count tells you how expensive the layer is" is wrong in both directions, and why a proposal to shrink a model must state *which* cost it is shrinking. ## Doing it under interview conditions Keep the two formulas separate and substitute carefully: 1. Kernel area times input channels gives one filter's weight count. 2. Multiply by output channels for the layer's weights; add `C_out` for the biases. 3. Multiply the weight count by `H_out * W_out` for MACs — and get `H_out`, `W_out` from the output-shape formula, not from the input size, since a strided layer has fewer output positions than input positions. Step 3 is where most candidates slip: for a stride-2 version of the same 3x3 layer on a 56x56 input, the output is 28x28, so the MACs are a quarter of the stride-1 figure even though the input map is identical. The misconceptions worth pre-empting: that parameter count grows with image size; that a filter is 2D rather than a `k x k x C_in` block; that there is one kernel per layer rather than one per output channel; and that quoting a weight count answers a question about arithmetic cost.

  • Why does moving that same layer to a 28x28 map leave the parameter count unchanged?
    Because the layer stores one filter per output channel and reuses it at every spatial position. Storage depends on kernel size and the two channel counts only. Resolution changes how many times the weights are applied — the multiply-add count — not how many weights exist. That is the definition of weight sharing.
  • How do you count the multiply-adds for a stride-2 version of the same layer?
    Take the output size from the shape formula rather than the input size. A stride-2 3x3 layer on a 56x56 input with padding 1 outputs 28x28, so the MACs are `589,824 * 784`, about a quarter of the stride-1 figure. The parameter count is identical, because stride is not stored.
  • How much do biases contribute to the total?
    One per output channel, so 256 out of 590,080 — under 0.05 percent. They matter for an exact count and for the extra `C_out * H_out * W_out` additions, but never for a budget decision. They are commonly dropped when a normalization layer follows, since that layer already applies a per-channel shift.
  • If you must halve the arithmetic cost of one layer, what are your options?
    Halve the output area — a stride or a smaller input — which quarters MACs and leaves parameters alone. Or cut the channel width by about 30 percent, since MACs scale with `C_in * C_out`; that also cuts parameters. Which you pick depends on whether the model is storage-bound or arithmetic-bound.

saying these in an interview costs you the question

  • Thinks parameter count grows with the feature-map size
  • Forgets to multiply by the input channel count
  • Counts one kernel per layer rather than one per output channel
  • Quotes a weight count when asked for arithmetic cost
  • Reports MACs without multiplying by output positions
  • Confuses MACs and FLOPs without stating the convention

context