skip to content

Regularization and Normalization

The tools that keep deep nets from overfitting and training stable: dropout, weight decay, augmentation, and batch versus layer normalization. 'BatchNorm vs LayerNorm, and why?' is near-universal.

on this pageshow

explore

questions

28

Why does a batch-normalized network give different predictions in training mode than in evaluation mode?

level: juniorimportance: must knowfreq 82%

answer

  1. two sets of statistics, not one
  2. who supplies the mean at test time
  3. batch statistics versus running averages
  4. evaluation must not peek at neighbours

basics

~20 s

In training, batch normalization standardizes each channel with the current mini-batch's mean and variance. In evaluation it uses running averages collected during training, so predictions become deterministic and stop depending on which other samples share the batch.

solid answer

~50 s

A batch normalization layer draws on two different sources of statistics. In training mode it measures the mean and variance of the activations in the current mini-batch, normalizes with those, then applies its learned scale and shift. In the same pass it updates a running mean and running variance as an exponential moving average of those batch statistics. In evaluation mode it stops looking at the batch entirely and normalizes with the running values, which turns the layer into a fixed per-channel affine transform: one sample in, the same answer out, no matter what its neighbours are. The gap is a real bug source. A model that scores 92% with the network left in training mode and 61% in evaluation mode on the very same batch is almost always a running-statistics problem, not a weight problem.

go deeper

for a junior

Be ready to state plainly that training uses the current batch's statistics while evaluation uses stored running averages, and that forgetting to switch modes is a bug you have actually hit.

for a middle

Explain the moving average that fills the running buffers, why they are state rather than trainable parameters, and why evaluation reduces the layer to a fixed per-feature affine transform.

for a senior

Show your diagnosis: compare per-channel running statistics against freshly measured batch statistics, count how many updates the buffers received, and check the momentum coefficient against how fast the data distribution moved.

for a principal

Own the policy angle — how mode state is guarded across shared training and serving code paths, and when a model whose output depends on batch composition should be considered unsafe to ship at all.

## The layer holds two kinds of state A batch normalization layer stores four vectors, one entry per normalized feature (per channel, for a convolutional layer), and they fall into two very different categories. - **Learned parameters**: a scale `gamma` and a shift `beta`. These are trained by gradient descent exactly like weights. - **Buffers**: a running mean and a running variance. These are *not* trained. They are bookkeeping, updated by a moving average during training-mode forward passes. Which of the two sets of statistics is used to standardize the activations is precisely what the mode switch controls. ## Training mode For each feature the layer computes, over the samples currently in the mini-batch: ``` mu = mean(x) var = mean((x - mu)^2) xhat = (x - mu) / sqrt(var + eps) y = gamma * xhat + beta ``` The key property is that `mu` and `var` come from *this* batch. A given sample's output therefore depends on the other samples that happened to be drawn alongside it. Resample the batch and that sample's activations change slightly. This is intentional during training — the noise is part of why the layer behaves as a mild regularizer — but it is unacceptable at serving time, where the same input must always produce the same prediction. In the same forward pass, the buffers are updated with an exponential moving average of the batch statistics, controlled by a momentum coefficient: ``` running_mean <- (1 - m) * running_mean + m * mu running_var <- (1 - m) * running_var + m * var ``` A small momentum coefficient means the buffers move slowly and average over many batches; a large one means they track the most recent batches closely. Note that no gradient flows into these updates — they are a running summary, not something the optimizer tunes. ## Evaluation mode In evaluation mode the layer ignores the incoming batch's statistics and uses the buffers: ``` y = gamma * (x - running_mean) / sqrt(running_var + eps) + beta ``` Every term on the right except `x` is now a constant, so the whole layer collapses to a per-feature affine map `y = a * x + b`. Three things follow. Predictions are deterministic. A batch of one works fine (in training mode a fully-connected batch of one has zero variance, which is degenerate). And no information leaks between test samples, which matters because a layer whose output depends on the other samples in the request is a correctness hazard, not just a performance quirk. ## Why the gap appears in practice The two modes agree only when the running statistics are good estimates of the activation statistics the network actually produces. They fail to agree in several familiar ways. - **The buffers were never warmed up.** A model evaluated after a handful of updates has running values still dominated by their initialization, so evaluation-mode activations are badly scaled and accuracy collapses while training-mode accuracy looks healthy. - **The momentum coefficient is too slow for a change in the data.** If the input distribution shifts partway through training — a new augmentation policy, a new data source — the batch statistics move immediately but the buffers lag, and evaluation accuracy trails training accuracy for several epochs before catching up. - **The mode was simply never switched.** Evaluating a network still in training mode leaks the evaluation batch's own statistics into its predictions and usually flatters the number; switching it back to evaluation then produces the sudden drop that sends people hunting for a bug in the weights. ## How to diagnose it The diagnostic is cheap: run one batch through the network in both modes and compare. If the weights were fine and only the mode changed, the weights are not the problem. Then, layer by layer, compare each channel's running mean and running variance against the mean and variance freshly measured on a representative batch. Large disagreement in a particular channel or a particular block localizes the problem. Also check how many training-mode updates the buffers have actually received; a model assembled from pieces, or one whose normalization layers spent training in a frozen state, may have buffers that saw almost no data. The fix follows from the cause: train longer so the averages converge, raise the momentum coefficient so the buffers track a shifted distribution faster, or run a pass over training data with the weights held fixed purely to re-estimate the statistics before deploying. ## The mental model to keep Training mode asks 'what do the activations look like right now, in this batch?'. Evaluation mode asks 'what did they look like on average during training?'. The layer is doing the same arithmetic in both cases; only the source of `mu` and `var` changes. Almost every batch normalization surprise reduces to that one sentence.

  • How do the running mean and variance get their values?
    By an exponential moving average of each mini-batch's statistics, applied on training-mode forward passes and controlled by a momentum coefficient. No gradient is involved — they are buffers, not parameters, so an optimizer never touches them. A small coefficient averages over many batches and is stable but slow to react; a large one tracks recent batches and reacts fast but is noisier.
  • Why not simply use the evaluation batch's own statistics at test time?
    Because a prediction would then depend on which other samples were served alongside it, so the same input could get different answers in different batches, and a single-sample request would have no usable variance at all. It also leaks information between test samples, which breaks reproducibility and can quietly inflate offline metrics that never survive production.
  • A model scores 92% in training mode and 61% in evaluation mode on the same batch — what do you check first?
    The running statistics, not the weights. Compare each layer's running mean and variance against statistics measured fresh on that batch, and check how many training-mode updates the buffers actually received. Common causes are buffers that never warmed up, a momentum coefficient too slow to follow a mid-training distribution change, or normalization layers that were frozen while the weights trained.

saying these in an interview costs you the question

  • Says a batch normalization layer behaves identically in both modes
  • Thinks evaluation recomputes the statistics from the test batch
  • Believes the running averages are learned by backpropagation
  • Confuses the learned scale and shift with the running statistics
  • Blames the weights when only the mode was switched

context

open as a page

Why does applying random label-preserving transforms to training images reduce overfitting?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Random label-preserving transforms show the network a different version of each image every epoch, so it cannot memorise exact pixels. This effectively enlarges the training set and bakes in invariances such as small shifts and rotations, cutting variance.

open as a page

Why does dropout reduce overfitting in a fully connected network?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Dropout zeroes each hidden unit at random on every training step, so no unit can rely on a particular other unit being present. Training therefore fits many thinned, weight-sharing subnetworks whose behaviour is averaged at evaluation.

open as a page

Which axis does LayerNorm normalize over, and why does batch size not change its output?

level: juniorimportance: must knowfreq 72%

basics

~20 s

LayerNorm computes a mean and variance across the feature dimensions of each sample on its own, then rescales with a learned per-feature gain and bias. No other sample enters the statistic, so batch size and batch composition cannot change the output.

open as a page

What does weight decay do to a neural network's weights during training?

level: juniorimportance: must knowfreq 82%

basics

~20 s

Weight decay pulls every decayed weight slightly toward zero on each update, on top of the gradient step. The result is a smaller-norm solution: the network fits with less extreme weights, which usually improves held-out performance.

open as a page

In a convolutional network, what does a batch normalization layer compute its statistics over?

level: middleimportance: must knowfreq 72%

basics

~20 s

One mean and one variance per channel, pooled over every element of that channel in the mini-batch: all samples and all spatial positions. A layer after a 64-channel convolution holds 64 means, 64 variances, 64 scales and 64 shifts.

open as a page

What does inverted dropout divide surviving activations by during training?

level: middleimportance: must knowfreq 62%

basics

~20 s

It divides every surviving activation by the keep probability, one minus the drop rate. That makes the layer's expected output equal its unmasked value, so evaluation runs the plain forward pass with no mask and no rescaling.

open as a page

Why does early stopping restore the best checkpoint rather than the final weights, and what must be saved to do it?

level: middleimportance: must knowfreq 62%

basics

~20 s

Patience keeps the run going for several epochs after its best validation score, so the final weights are exactly the ones that already got worse. Restore the best epoch's weights, saved together with the normalization layers' running statistics.

open as a page

How does label smoothing change the target that a classifier's cross-entropy loss fits?

level: middleimportance: must knowfreq 58%

basics

~10 s

Label smoothing replaces the one-hot target with a blend of one-hot and uniform: the true class is asked for about 1 minus epsilon, every other class for epsilon divided by the number of classes.

open as a page

In pre-norm versus post-norm blocks, where does normalization sit relative to the residual add?

level: middleimportance: must knowfreq 62%

basics

~20 s

Post-norm normalizes after the residual addition: out = Norm(x + F(x)). Pre-norm normalizes the branch input instead: out = x + F(Norm(x)). Pre-norm leaves an un-normalized identity path running the whole depth of the stack, which is why deep stacks train far more easily.

open as a page

How does mixup combine two training images and their one-hot labels into one example?

level: middleimportance: should knowfreq 52%

basics

~20 s

mixup draws a mixing weight lam from a Beta distribution and emits one blended example: pixels lam*x_i + (1-lam)x_j, with the target lamy_i + (1-lam)*y_j. The label is interpolated by the same weight, never rounded back to a single class.

open as a page

Why are biases and normalization scale and shift parameters usually exempt from weight decay?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because shrinking them regularizes almost nothing and can actively hurt. Biases and normalization scale and shift are few in number, do not control how sharply the function bends, and pulling a normalization scale toward zero throttles that layer's output.

open as a page

Why freeze a backbone's batch normalization statistics when fine-tuning at 2 images per device?

level: seniorimportance: should knowfreq 46%

basics

~20 s

With two samples per batch the mean and variance are almost pure noise, and updating on them overwrites far better statistics inherited from large-batch pretraining. Freezing keeps the inherited values and makes each layer a stable fixed affine transform.

open as a page

Designing an augmentation policy for chest radiographs, how do you decide a transform is label-preserving?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Ask whether an expert labeller would give the transformed image the same label. Brightness and contrast jitter mimic exposure differences and are safe on chest radiographs; a horizontal flip is not, since mirroring destroys laterality.

open as a page

Why does a dropout layer placed just before batch norm hurt evaluation accuracy?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Inverted dropout inflates the variance of the activations it emits, but only during training. Batch norm's running variance is therefore estimated on inflated inputs and is too large for the mask-free activations at evaluation, so outputs come out shrunk.

open as a page

Does label smoothing fix an overconfident classifier's calibration?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It usually reduces miscalibration, because it removes the incentive to drive confidence toward 1. But it is a blunt global cap, not a calibration procedure: too much smoothing turns overconfidence into underconfidence, and it never guarantees confidence tracks accuracy.

open as a page

How do you choose the group count for GroupNorm in a batch-size-1 3D segmentation network?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The group count is a dial between whole-sample statistics at one group and per-channel statistics at one group per channel. Pick a middle value, or a fixed channels-per-group, and decide by what the statistic destroys rather than by estimator noise.

open as a page

In a pre-norm stack, how does the residual stream's scale change with depth?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It grows. Each pre-norm block adds a branch output whose size is set by its own normalized input, not by the stream, so contributions accumulate and the stream's standard deviation rises roughly with the square root of depth. That is why a final normalization must sit before the output head.

open as a page

Sweeping weight decay across four decades on a wearable-accelerometer activity classifier, how do you pick the value?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Read the held-out curve, not the training curve. Across a logarithmic sweep the held-out metric is U-shaped: too little decay changes nothing, too much underfits. Pick from the flat top of that curve, then re-sweep when data volume or schedule changes.

open as a page

What should early stopping monitor when validation loss and the shipped horizon-7 MAPE disagree?

level: principalimportance: should knowfreq 44%

basics

~20 s

Monitor the quantity the model is judged on, not the surrogate loss, whenever it can be computed every epoch. Track its own direction of improvement and its own noise scale, and keep the checkpoint that is best on it.

open as a page

What does RMSNorm drop compared with LayerNorm, and what does that assume?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

RMSNorm drops the mean subtraction and usually the bias, dividing each feature by the vector's root mean square and applying a learned gain. It assumes re-scaling invariance is what makes normalization work, and that re-centring can be given up.

open as a page

Does batch normalization help by reducing internal covariate shift, or by something else?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

The original internal-covariate-shift story did not survive testing. Later work showed the layer mainly smooths the loss landscape, so gradients stay predictive and larger learning rates remain stable, with the batch-sampled noise adding a mild regularizing effect.

open as a page

When is ten-crop plus flip test-time augmentation worth its ten-fold inference cost?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Only when the accuracy gain is worth ten forward passes per input. Averaging predictions over ten label-preserving views cuts prediction variance for a fixed model, which suits offline or batch scoring far better than a latency-bound online request path.

open as a page

Why does elementwise dropout barely regularize a convolutional feature map?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Neighbouring positions in a feature map are strongly correlated because their receptive fields overlap, so zeroing scattered individual activations leaves the same evidence present next door. Channel-wise dropout, which zeroes whole feature maps, removes a detector outright and actually regularizes.

open as a page

Patience of 5 stops your run right after every cosine warm restart. Why, and how do you fix it?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

A warm restart lifts the learning rate back up, so validation worsens for several epochs by construction before recovering past the previous cycle's best. A patience of 5 reads that dip as failure. Judge once per cycle instead.

open as a page

How does weight decay affect a layer whose output feeds straight into a normalization layer?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

For a layer whose output is immediately normalized, rescaling its weights leaves the function unchanged. Weight decay there is not capacity control: it keeps the weight norm from growing, which stabilizes the effective step size those weights receive.

open as a page

Should label smoothing be on for a shared backbone serving both classification and retrieval?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Not automatically. Smoothing tightens each class into an equidistant cluster in the penultimate layer and erases the similarity structure between classes, so it can raise top-1 accuracy while degrading nearest-neighbour retrieval built on the same features.

open as a page

How do you choose between pre-norm and post-norm for a model you intend to scale deeper?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Decide by the depth you plan to reach, not by a shallow benchmark. Post-norm has been reported slightly better at around six blocks; pre-norm is the placement that still trains at a hundred. Trainability at the target depth dominates, and switching later invalidates your tuned hyperparameters.

open as a page