In a convolutional network, what does a batch normalization layer compute its statistics over?
answer
- the axis defines the layer
- per channel, not per sample
- pixels count as extra samples
- two learned vectors of length C
basics
~20 sOne mean and one variance per channel, pooled over every element of that channel in the mini-batch: all samples and all spatial positions. A layer after a 64-channel convolution holds 64 means, 64 variances, 64 scales and 64 shifts.
solid answer
~50 sBatch normalization is per-feature, and in a convolutional layer the feature is the channel. Given activations shaped batch x channels x height x width, it pools every element of one channel — across all N samples and all H*W spatial positions — into a single mean and variance, so a 64-channel layer produces 64 pairs, not one pair per sample and not one per pixel. Each channel is then standardized as `xhat = (x - mu) / sqrt(var + eps)` and passed through its own learned scale and shift, `y = gamma * xhat + beta`. Pooling the spatial positions applies the same normalization at every location, which is what a shared convolutional filter needs. The scale and shift matter because pure standardization would pin every channel to zero mean and unit variance; the two learned vectors let the network pick a different operating point, including the identity.
go deeper
Recall that the statistics are per channel and are shared across the samples in the batch, and that the layer adds a learned scale and shift on top of the standardization.
Be able to state the axes precisely for a batch x channels x height x width tensor, write the standardize-then-scale-and-shift formula, and count the layer's parameters and buffers.
Explain the design consequences: why spatial pooling preserves translation consistency, why the preceding bias is redundant, and why zero-initializing the scale in a residual block starts the network near the identity.
Frame the axis choice as the real decision — what a network gains by tying samples together through shared statistics, and what that coupling costs in serving, evaluation and reproducibility.
## The axis question Every normalization layer is defined by one thing: which axis its statistics are averaged over. Batch normalization averages over the **batch axis**. What varies is which other axes come along for the ride. For a fully-connected layer with activations shaped `[N, F]` (N samples, F features), the layer computes F means and F variances, each averaged over the N samples. Feature 3 is normalized using how feature 3 behaved across the batch, and never mixed with feature 4. For a convolutional layer with activations shaped `[N, C, H, W]`, the feature is the channel, and the spatial positions are treated as extra samples of that same feature. So the layer computes C means and C variances, each averaged over `N * H * W` numbers. Channel 3's mean is taken over every sample and every pixel of channel 3. That spatial pooling is a deliberate design choice, not an implementation shortcut. A convolutional filter applies the same kernel at every location, so its output at every location is a sample of the same feature. Normalizing each pixel with its own statistics would break that shared-weight logic and make the transform depend on position. ## The full computation Per channel c, over the mini-batch: ``` mu_c = mean over N, H, W of x[:, c, :, :] var_c = mean over N, H, W of (x[:, c, :, :] - mu_c)^2 xhat = (x[:, c, :, :] - mu_c) / sqrt(var_c + eps) y = gamma_c * xhat + beta_c ``` `eps` is a small constant inside the square root. It exists purely for numerical safety: a channel that happens to be nearly constant across the batch has a variance near zero, and dividing by the square root of zero would blow the activations up or produce a non-finite value. It is not a tuning knob in any interesting sense. ## What the layer stores For C channels, the layer keeps four vectors of length C: - `gamma` and `beta` — trainable, updated by the optimizer, `2C` parameters in total. - running mean and running variance — buffers, updated by a moving average of the batch statistics, `2C` more numbers that ship with the model but are never touched by the optimizer. So a normalization layer following a 256-channel convolution stores 1024 numbers, half of them trainable. This is tiny compared with the convolution itself, which is why the layer is essentially free in parameter count while being far from free in memory traffic. ## Why gamma and beta exist If the layer only standardized, it would force every channel to zero mean and unit variance forever, which is a real constraint on what the network can represent. Consider a channel feeding a saturating nonlinearity: pinning it to unit variance fixes how much of that nonlinearity's saturating region it can reach. The learned scale and shift give the parameters back. Crucially, the pair can recover the identity — if setting `gamma_c` to the batch standard deviation and `beta_c` to the batch mean is what minimizes the loss, the optimizer can find it. Normalization is therefore a reparameterization of the activation, not a restriction on the function class. A useful consequence: the bias of the layer immediately before the normalization is redundant. Any constant added to a channel is added to that channel's mean and subtracted right back out by the centering step, so the bias has no effect on the output — `beta` already plays that role. Dropping it costs nothing and removes a parameter whose gradient is structurally zero in effect. ## Where it sits in a block The conventional convolutional block is convolution, then normalization, then the nonlinearity. Placing normalization before the nonlinearity is what controls the input distribution the nonlinearity sees, which is the whole point; putting it after normalizes the already-rectified activations, whose distribution is one-sided, and is used far less often. The block also drops the convolution's bias for the reason above. ## Common confusions worth naming **'It normalizes each sample.'** It does not. Batch normalization's statistics are shared across the batch, and each sample's output depends on its neighbours. Normalization schemes that work within a single sample exist and are a different family with a different axis. **'It normalizes over the training set.'** Not during training. The mean and variance come from the current batch. Only the buffers, used at evaluation time, are a whole-training average, and even they are an exponentially weighted one. **'It has no learnable parameters.'** It has `2C` of them, and they matter: `gamma` in particular is the handle people inspect when a channel is effectively dead, and initializing the final normalization layer of a residual block with `gamma` at zero is a known trick for making very deep networks start close to the identity. ## The one-line answer One mean and one variance per channel, taken over the batch and all spatial positions; then a per-channel learned scale and shift on top.
- Why does a convolution placed immediately before a batch normalization layer not need its bias term?Because the centering step removes it. A bias adds a constant to every element of a channel, which raises that channel's batch mean by exactly the same constant, and subtracting the mean cancels it out. The learned shift `beta` already provides a per-channel offset after normalization, so the convolution's bias is a parameter with no effect on the output.
- If the layer just standardizes the activations, what do the learned scale and shift buy you?They restore representational freedom. Pure standardization would pin every channel to zero mean and unit variance permanently, which constrains where a channel sits relative to its nonlinearity. The scale and shift let gradient descent choose any mean and spread, including exactly reproducing the unnormalized activation, so normalization becomes a reparameterization rather than a restriction on what the network can express.
- How many numbers does a batch normalization layer after a 256-channel convolution hold?1024. There are 256 scale values and 256 shift values, which are trainable parameters, plus a 256-entry running mean and a 256-entry running variance, which are buffers the optimizer never touches but which are saved with the model. Half the storage is trainable; all of it must survive a checkpoint round-trip or evaluation behaviour breaks.
saying these in an interview costs you the question
- Says it normalizes each sample independently of the batch
- Forgets that spatial positions are pooled into the same statistic
- Claims the layer has no learnable parameters
- Says the running statistics are the learned scale and shift
- Computes one mean for the whole tensor instead of one per channel