skip to content

Why does early stopping restore the best checkpoint rather than the final weights, and what must be saved to do it?

level: middleimportance: must knowfreq 62%

answer

  1. two rules, not one
  2. the exit condition tells you which epochs were worse
  3. patience epochs of measured decline at the end
  4. gradients are not the only tensors that change
  5. serve needs buffers, resume needs optimizer state

basics

~20 s

Patience keeps the run going for several epochs after its best validation score, so the final weights are exactly the ones that already got worse. Restore the best epoch's weights, saved together with the normalization layers' running statistics.

solid answer

~50 s

Early stopping is two mechanisms, not one: a rule that decides when to quit, and a rule that decides which weights you keep. The stopping rule only fires after `patience` epochs with no improvement, so by construction the last `patience` epochs are the ones where the monitored score got worse — shipping the final weights throws away the selection you just paid for. So every time the monitored score improves you write a checkpoint, and when the loop exits you load that file back before evaluating or exporting. To serve the model the checkpoint needs the learned parameters plus any non-gradient buffers, most importantly the running mean and variance a normalization layer accumulated. To *resume* training it also needs the optimizer's accumulated state, the epoch index, the schedule position and the best score seen so far, or the restarted run silently begins from a different place than it stopped.

go deeper

for a junior

Be ready to say what patience counts and that the weights you keep come from the best epoch, not the last one. Know that a checkpoint file is written when the monitored score improves.

for a middle

Explain why the trailing epochs are worse as a consequence of the exit condition, and list what a checkpoint must contain to serve versus to resume — buffers for one, optimizer and schedule state for the other.

for a senior

Show you have debugged the restore path: a model that scores worse after loading than it did in the loop, and the buffer or evaluation-mode mistake behind it. Own a two-file best/last convention rather than per-epoch dumps.

for a principal

Frame stopping and selection as separable policies and decide which one your team enables by default across every run. Own the reporting discipline that keeps the selected score out of the headline number.

## Two rules wearing one name "Early stopping" bundles two independent decisions, and interviewers probe whether you see them as separate. 1. **The stopping rule.** Evaluate a monitored quantity on a held-out validation split once per epoch. Track the best value seen so far. Count how many consecutive epochs have failed to beat it by more than a minimum-improvement threshold; when that count reaches `patience`, halt. 2. **The selection rule.** Of all the epochs you actually trained, which weights do you keep? You can run either without the other. Keeping the best checkpoint while training the full planned number of epochs is a perfectly good configuration — you get the selection benefit and give up only the saved compute. The reverse (stopping early, then shipping whatever was in memory) is the configuration that quietly costs you accuracy. ## Why the final weights are the wrong ones This follows mechanically from the counter. Suppose patience is 8 and the loop exits at epoch 40. The exit condition says: epoch 32 set the best score, and epochs 33 through 40 all failed to beat it. The final weights are therefore, by definition, drawn from a run of epochs that were *measured to be worse*. In the leaf's canonical case — a 60-epoch run whose best epoch was 12 — the gap is enormous: you would be exporting a model 48 epochs into overfitting. The gap is worse than it looks when the monitored curve is noisy. On a small training set (say a few thousand examples) the per-epoch validation score rattles up and down within a band, and the last epochs may sit anywhere in that band. "Final weights" is not just a slightly worse choice, it is an unselected draw. ## What a checkpoint has to contain Split this by what you intend to do with the file. **To evaluate or serve the model**, you need: - Every learned parameter — weights and biases, including the scale and shift of normalization layers. - Every **buffer** that is updated during training but not by gradients. The classic one is the running mean and variance a batch-normalizing layer accumulates during training and uses at evaluation time. These are not parameters; a save routine that only walks the gradient-carrying tensors drops them, and the restored model then evaluates with statistics that never got trained. The symptom is a restored model that scores far worse than the number your training loop printed for that same epoch. - Whatever the input pipeline learned, if anything — feature scaling constants, a vocabulary or label mapping. A checkpoint that cannot reproduce its own preprocessing is not restorable. **To resume an interrupted run**, add: - **Optimizer state.** A momentum method carries a velocity buffer per parameter; an adaptive method carries first- and second-moment estimates per parameter. Dropping these restarts the optimizer cold, which produces a visible bump in the loss at the resume point. - **The position in the learning-rate schedule** (step or epoch count), so the resumed run does not replay a warm-up or jump to the wrong rate. - **Bookkeeping for the stopping rule itself**: the best score so far and the epochs-since-best counter. Without them a resumed run either stops immediately or forgets that it already had a better model. - Random-number generator state, if you need bit-exact reproducibility of augmentation and shuffling. A practical convention is to keep exactly two files — `best` and `last` — rather than one per epoch. `best` is what you ship; `last` is what you resume from. Writing every epoch is mostly wasted disk on a large model, and it invites the mistake of picking the wrong file at export time. ## The score at the best checkpoint is optimistic One consequence people miss: the validation score you selected on is a *maximum over many noisy epochs*, so it is an optimistically biased estimate of generalization. You chose the epoch partly because that epoch got lucky on that particular split. The wider the noise band and the more epochs you scan, the larger the bias. The fix is not to stop selecting — it is to report the final number on a test split that played no part in the stopping or selection decision. ## When the final weights are acceptable If the run trained a full, fixed schedule that anneals the learning rate down to near zero, the last epochs are typically the best epochs, and best and last usually coincide. Even there, keeping the best checkpoint costs nothing and protects you from the epoch where a data-loading hiccup or a late instability spiked the score. Treat "export the final weights" as a decision to justify, never as the default.

  • The restored model scores much worse than the number your training loop printed at that epoch. What do you check first?
    Whether the checkpoint captured the non-gradient buffers, above all a normalization layer's running mean and variance. A save routine that walks only gradient-carrying tensors restores a model whose evaluation-time statistics were never trained, and the score collapses. Second check: that the restored model is in evaluation mode and the preprocessing constants match the ones used during training.
  • Can you trust the validation score at the restored best epoch as your reported result?
    No. That number is the maximum of a noisy per-epoch statistic over many epochs, so it is optimistically biased — you selected the epoch partly for its luck on that split. Report the final figure on a test split that took no part in stopping or checkpoint selection. The gap between the two grows with the number of epochs scanned and the noise of the validation set.
  • Do you need early stopping at all if you already restore the best checkpoint?
    Not necessarily. The two rules are separable: keeping the best checkpoint while training the full planned schedule gives you the entire selection benefit and gives up only saved compute. Early stopping adds value when runs are long, the budget is tight, or you sweep many configurations. Keeping the best checkpoint is close to free, so it is the one to enable unconditionally.

Patience is a stop-loss that only triggers after the price has already fallen for several days. You would not sell at the trigger price and call it your best exit — you record the peak as it happens.

saying these in an interview costs you the question

  • Says the weights at the stop are the best weights
  • Saves only gradient-carrying tensors and loses normalization running statistics
  • Thinks resuming needs weights alone, not optimizer state
  • Reports the selected best-epoch validation score as the final result
  • Writes a checkpoint every epoch and picks one by hand later

context