skip to content

What breaks when a validation pass is run in training mode instead of evaluation mode?

level: seniorimportance: should knowfreq 60%

answer

  1. measuring, not fitting
  2. two layer types behave differently
  3. randomly thinned network each batch
  4. prediction depends on batch neighbours
  5. held-out data written into the model

basics

~20 s

Dropout stays active, so every validation number is a noisy sample of a randomly thinned network, and normalization layers use the validation batch's own statistics and refresh their stored ones. The metric becomes irreproducible, batch-order dependent, and validation data leaks into the model.

solid answer

~50 s

An evaluation pass exists to answer one question: how does the *current* parameter set score on data it was not fitted to. It is therefore a forward pass only — no backward, no update — with gradient bookkeeping switched off, and every layer that behaves differently while training switched to its inference behaviour. Leave it in training mode and three things break. Dropout keeps randomly zeroing activations, so each validation number is a sample from a randomly thinned sub-network rather than the model you would ship. Normalization layers that use the current batch's statistics make an example's prediction depend on which other examples share its batch, so shuffling the validation set changes the score. Worst, those layers refresh their stored statistics from validation batches, so held-out data has now written into the model's state — a real leak, not just a noisy metric.

go deeper

for a junior

Know that a validation pass runs the forward pass only, with no backward and no weight update, and that the model must be switched to inference behaviour first.

for a middle

Explain concretely which layer behaviours differ between the two modes and why gradient bookkeeping is disabled even though no backward pass is called.

for a senior

Demonstrate that you can spot the leak, not just the noise: held-out data writing into stored statistics contaminates every later number and the shipped checkpoint.

for a principal

Own the guardrail. Require that evaluation be reproducible to the digit and that mode switching sit in shared, tested code rather than in each experiment script.

## What an evaluation pass is A training step and an evaluation step start the same way and then diverge sharply. | | Training pass | Evaluation pass | |---|---|---| | Forward pass | yes | yes | | Loss / metric computed | yes | yes | | Backward pass | yes | **no** | | Parameter update | yes | **no** | | Gradient bookkeeping | on | **off** | | Layer behaviour | training | **inference** | The first four rows are the ones candidates name. The last row is the one that causes the incident. ## Why forward-only, and why turn gradient bookkeeping off No backward pass is needed because you are measuring, not fitting. But merely *not calling* the backward pass is not enough: unless gradient bookkeeping is explicitly disabled, the forward pass still retains every intermediate activation in case a backward pass follows. On a validation set that is pure waste — often the dominant memory cost of the pass. Disabling it frees that memory, which in turn lets you evaluate with a much larger batch than you train with, and speeds the pass up. ## The three things that break in training mode **1. Dropout stays on.** During training, dropout randomly zeroes a fraction of activations to prevent co-adaptation. During inference it must be off so the full network is used. A validation pass with dropout still active evaluates a *different randomly thinned sub-network on every batch*. The number you get is a noisy sample, it changes on every re-run of the same pass over the same data with the same parameters, and it is systematically worse than the model's real quality. Teams have concluded a model was overfitting, or was worse than a baseline, on the strength of a number produced this way. **2. Normalization by the current batch's statistics.** Layers that normalize using statistics computed across the examples in the batch must switch, at inference, to fixed stored statistics. In training mode they instead use the validation batch's own statistics — so an example's prediction depends on which other examples happen to share its batch. Consequences: the reported metric moves when you reshuffle the validation set, moves when you change the evaluation batch size, and a single-example evaluation becomes meaningless or degenerate. A metric that depends on batch composition is not a property of the model. **3. Stored statistics are refreshed from validation data.** This is the serious one, because it is not merely a bad measurement — the model itself changes. Those layers update their stored statistics from every batch they see in training mode, so running validation in training mode writes information about the held-out set into the model's state. No parameter update was applied and no gradient was computed, yet the model that gets checkpointed has absorbed the validation distribution. Every subsequent validation number is contaminated, and if that checkpoint ships, the reported held-out score is not honest. And if the update was left in as well, you are no longer leaking — you are training on the validation set outright, and the split has ceased to exist. ## How you notice - **The metric is not reproducible.** Two evaluation passes over the same fixed parameters and the same data return different numbers. In a correct evaluation pass they must be identical. - **The metric moves with evaluation batch size or shuffling.** Both are strong evidence of batch-dependent normalization. - **Validation loss is implausibly noisy step to step** while the training loss is smooth — the mirror image of the usual expectation. - **The gap closes or reverses oddly** because the model has quietly seen the held-out distribution. ## The related bookkeeping subtlety Even a correct evaluation pass can misreport if you average per-batch means. If the validation set does not divide evenly by the evaluation batch size, the final ragged batch has fewer examples, and an unweighted average of batch means over-weights it. Accumulate the *sum* of per-example losses and divide by the total example count, or weight each batch mean by its size. This is a small error compared to the mode bug, but it is the reason two correct-looking implementations report slightly different validation losses on the same model. ## What to say in an interview Name all three failure modes, and be explicit that the third is a leak rather than a measurement error — that distinction is what separates a senior answer from a memorised one. Then name the check: fix the parameters, run the same evaluation pass twice, and require the identical number. A validation metric that is not reproducible to the last digit is telling you the pass is not an evaluation pass.

  • Why disable gradient bookkeeping during evaluation if you never run the backward pass anyway?
    Because the forward pass otherwise retains every intermediate activation in case a backward pass follows, and on a large validation set that retained graph is usually the dominant memory cost. Turning it off frees that memory, lets you evaluate at a much larger batch size than you train with, and makes the pass measurably faster.
  • Should two evaluation passes over the same data and parameters return the identical number?
    Yes, to the last digit. A correct evaluation pass is a deterministic function of the parameters and the data: dropout off, fixed normalization statistics, no augmentation. If the number moves between runs, the pass is still in training mode or randomness is leaking in from the data pipeline — treat it as a bug, not as noise.
  • Does reshuffling the validation set change the reported metric?
    In a correct evaluation pass, no — if you sum per-example losses and divide by the example count. It does shift slightly if you average per-batch means and the final batch is ragged, since the small batch gets equal weight. In training mode it changes the metric substantially, because batch composition then affects the predictions themselves.

saying these in an interview costs you the question

  • Says evaluation just skips the update and nothing else
  • Thinks dropout during validation is harmless
  • Believes a forward pass can never modify the model
  • Treats an irreproducible validation number as normal noise
  • Averages per-batch means over a ragged final batch

context