Why does a batch-normalized network give different predictions in training mode than in evaluation mode?
answer
- two sets of statistics, not one
- who supplies the mean at test time
- batch statistics versus running averages
- evaluation must not peek at neighbours
basics
~20 sIn training, batch normalization standardizes each channel with the current mini-batch's mean and variance. In evaluation it uses running averages collected during training, so predictions become deterministic and stop depending on which other samples share the batch.
solid answer
~50 sA batch normalization layer draws on two different sources of statistics. In training mode it measures the mean and variance of the activations in the current mini-batch, normalizes with those, then applies its learned scale and shift. In the same pass it updates a running mean and running variance as an exponential moving average of those batch statistics. In evaluation mode it stops looking at the batch entirely and normalizes with the running values, which turns the layer into a fixed per-channel affine transform: one sample in, the same answer out, no matter what its neighbours are. The gap is a real bug source. A model that scores 92% with the network left in training mode and 61% in evaluation mode on the very same batch is almost always a running-statistics problem, not a weight problem.
go deeper
Be ready to state plainly that training uses the current batch's statistics while evaluation uses stored running averages, and that forgetting to switch modes is a bug you have actually hit.
Explain the moving average that fills the running buffers, why they are state rather than trainable parameters, and why evaluation reduces the layer to a fixed per-feature affine transform.
Show your diagnosis: compare per-channel running statistics against freshly measured batch statistics, count how many updates the buffers received, and check the momentum coefficient against how fast the data distribution moved.
Own the policy angle — how mode state is guarded across shared training and serving code paths, and when a model whose output depends on batch composition should be considered unsafe to ship at all.
## The layer holds two kinds of state A batch normalization layer stores four vectors, one entry per normalized feature (per channel, for a convolutional layer), and they fall into two very different categories. - **Learned parameters**: a scale `gamma` and a shift `beta`. These are trained by gradient descent exactly like weights. - **Buffers**: a running mean and a running variance. These are *not* trained. They are bookkeeping, updated by a moving average during training-mode forward passes. Which of the two sets of statistics is used to standardize the activations is precisely what the mode switch controls. ## Training mode For each feature the layer computes, over the samples currently in the mini-batch: ``` mu = mean(x) var = mean((x - mu)^2) xhat = (x - mu) / sqrt(var + eps) y = gamma * xhat + beta ``` The key property is that `mu` and `var` come from *this* batch. A given sample's output therefore depends on the other samples that happened to be drawn alongside it. Resample the batch and that sample's activations change slightly. This is intentional during training — the noise is part of why the layer behaves as a mild regularizer — but it is unacceptable at serving time, where the same input must always produce the same prediction. In the same forward pass, the buffers are updated with an exponential moving average of the batch statistics, controlled by a momentum coefficient: ``` running_mean <- (1 - m) * running_mean + m * mu running_var <- (1 - m) * running_var + m * var ``` A small momentum coefficient means the buffers move slowly and average over many batches; a large one means they track the most recent batches closely. Note that no gradient flows into these updates — they are a running summary, not something the optimizer tunes. ## Evaluation mode In evaluation mode the layer ignores the incoming batch's statistics and uses the buffers: ``` y = gamma * (x - running_mean) / sqrt(running_var + eps) + beta ``` Every term on the right except `x` is now a constant, so the whole layer collapses to a per-feature affine map `y = a * x + b`. Three things follow. Predictions are deterministic. A batch of one works fine (in training mode a fully-connected batch of one has zero variance, which is degenerate). And no information leaks between test samples, which matters because a layer whose output depends on the other samples in the request is a correctness hazard, not just a performance quirk. ## Why the gap appears in practice The two modes agree only when the running statistics are good estimates of the activation statistics the network actually produces. They fail to agree in several familiar ways. - **The buffers were never warmed up.** A model evaluated after a handful of updates has running values still dominated by their initialization, so evaluation-mode activations are badly scaled and accuracy collapses while training-mode accuracy looks healthy. - **The momentum coefficient is too slow for a change in the data.** If the input distribution shifts partway through training — a new augmentation policy, a new data source — the batch statistics move immediately but the buffers lag, and evaluation accuracy trails training accuracy for several epochs before catching up. - **The mode was simply never switched.** Evaluating a network still in training mode leaks the evaluation batch's own statistics into its predictions and usually flatters the number; switching it back to evaluation then produces the sudden drop that sends people hunting for a bug in the weights. ## How to diagnose it The diagnostic is cheap: run one batch through the network in both modes and compare. If the weights were fine and only the mode changed, the weights are not the problem. Then, layer by layer, compare each channel's running mean and running variance against the mean and variance freshly measured on a representative batch. Large disagreement in a particular channel or a particular block localizes the problem. Also check how many training-mode updates the buffers have actually received; a model assembled from pieces, or one whose normalization layers spent training in a frozen state, may have buffers that saw almost no data. The fix follows from the cause: train longer so the averages converge, raise the momentum coefficient so the buffers track a shifted distribution faster, or run a pass over training data with the weights held fixed purely to re-estimate the statistics before deploying. ## The mental model to keep Training mode asks 'what do the activations look like right now, in this batch?'. Evaluation mode asks 'what did they look like on average during training?'. The layer is doing the same arithmetic in both cases; only the source of `mu` and `var` changes. Almost every batch normalization surprise reduces to that one sentence.
- How do the running mean and variance get their values?By an exponential moving average of each mini-batch's statistics, applied on training-mode forward passes and controlled by a momentum coefficient. No gradient is involved — they are buffers, not parameters, so an optimizer never touches them. A small coefficient averages over many batches and is stable but slow to react; a large one tracks recent batches and reacts fast but is noisier.
- Why not simply use the evaluation batch's own statistics at test time?Because a prediction would then depend on which other samples were served alongside it, so the same input could get different answers in different batches, and a single-sample request would have no usable variance at all. It also leaks information between test samples, which breaks reproducibility and can quietly inflate offline metrics that never survive production.
- A model scores 92% in training mode and 61% in evaluation mode on the same batch — what do you check first?The running statistics, not the weights. Compare each layer's running mean and variance against statistics measured fresh on that batch, and check how many training-mode updates the buffers actually received. Common causes are buffers that never warmed up, a momentum coefficient too slow to follow a mid-training distribution change, or normalization layers that were frozen while the weights trained.
saying these in an interview costs you the question
- Says a batch normalization layer behaves identically in both modes
- Thinks evaluation recomputes the statistics from the test batch
- Believes the running averages are learned by backpropagation
- Confuses the learned scale and shift with the running statistics
- Blames the weights when only the mode was switched