skip to content

Patience of 5 stops your run right after every cosine warm restart. Why, and how do you fix it?

level: seniorimportance: nice to knowfreq 31%

answer

  1. the counter assumes a settling run
  2. the dip is by construction, not by failure
  3. the schedule, not the data, sets the stop epoch
  4. measure once per cycle, not once per epoch
  5. the stopper must outlast the reducer

basics

~20 s

A warm restart lifts the learning rate back up, so validation worsens for several epochs by construction before recovering past the previous cycle's best. A patience of 5 reads that dip as failure. Judge once per cycle instead.

solid answer

~50 s

The stopping rule assumes the monitored curve improves roughly monotonically, and a restarting schedule breaks that assumption on purpose. Each restart raises the learning rate, which throws the weights out of the basin they had settled into; validation degrades for a few epochs and only then anneals back down, usually below the previous cycle's best. A patience of 5 counts those degraded epochs as consecutive non-improvements and halts mid-cycle — and mid-cycle is the worst possible place to stop, because the model never received the low-rate refinement the schedule was built to deliver. Three fixes, in order of preference: evaluate the stopping rule once per cycle, comparing cycle-best to cycle-best; or set patience comfortably longer than one cycle so a dip cannot exhaust it; or drop the stopping rule entirely, train the schedule to completion, and keep only the best-checkpoint half of the mechanism.

go deeper

for a junior

Know that the learning rate can change during a run and that early stopping compares validation scores between epochs. Recognise that the two mechanisms are watching the same curve.

for a middle

Explain why a restart makes validation worse for a few epochs before it improves, and why a fixed per-epoch patience counter reads that as failure. Be able to say what patience should be measured over instead.

for a senior

Diagnose it from the run's own evidence — a stop epoch locked to schedule events — and pick a fix you can defend, including simply dropping the stopping rule while keeping best-checkpoint selection. Know the plateau-reducer ordering rule.

for a principal

Own the convention across the team: which schedules may be paired with a stopping rule at all, at what granularity the rule is evaluated, and when a fixed-length run with best-checkpoint selection is simply the safer default.

## The assumption the stopping rule makes A patience counter encodes one belief: if the monitored score has not improved for `patience` epochs, it is not going to. That belief holds when the learning rate is constant or decaying, because the run is settling. It is false whenever the schedule deliberately perturbs the model — and a warm-restarting schedule does exactly that, periodically raising the learning rate back to a high value. The consequence is mechanical. Right after a restart, the larger steps knock the weights out of the basin they had annealed into, so validation gets worse for several epochs. Then the rate decays again and the score recovers, typically finishing the cycle below the previous cycle's best. A patience of 5 sitting on the epoch-level curve sees five consecutive non-improvements inside that dip and fires. The stop is not detecting overfitting; it is detecting the schedule. ## Why stopping mid-cycle is the expensive failure Stopping is bad here for a second, sharper reason. Under an annealing schedule, most of the model's final quality is bought in the low-rate epochs at the end of a cycle. Halt at epoch 30 of a 60-epoch anneal and you export a model that only ever saw large steps — it never got the refinement phase the schedule was sized to deliver. That model is not "the 60-epoch model, stopped early"; it is a different and worse model, and it will typically underperform a plain 30-epoch schedule that was allowed to anneal within its own budget. Schedules that are parameterized by total length carry a hidden assumption that you will actually run to the end. ## Fixes, in order **Judge at cycle boundaries.** Instead of running the counter on every epoch, take one measurement per cycle — the best score achieved within that cycle — and run patience over those. Now patience means "cycles without improvement", which is what you actually meant, and the intra-cycle dip is invisible to the rule. This is the cleanest fix: it keeps the mechanism honest and stops only at points where the model has just been annealed. **Lengthen patience past a cycle.** If you keep the per-epoch counter, patience must exceed the worst-case recovery window — for a restarting schedule, at least a full cycle, and more when cycles grow in length. This works but is fragile: it silently re-couples a stopping hyperparameter to a schedule hyperparameter, so lengthening the cycles later reintroduces the bug. **Stop stopping.** Since the stopping and selection rules are independent, you can drop the stopping rule, train the entire schedule as designed, and keep only the best-checkpoint machinery. You retain all of the selection benefit and give up only saved compute. On a run whose length is already fixed by the schedule, this is often the right answer, and it removes the interaction entirely. ## The same clash in the plateau-triggered case A related pairing bites just as often: a schedule that cuts the learning rate when the monitored score plateaus, combined with an early-stopping rule on the same score. Both watch the same signal with their own counters, so their patience values must be ordered. The stopper's patience has to be strictly and comfortably larger than the reducer's — large enough that after the rate drops, the run gets a genuine chance to show the improvement the drop usually produces. If they are equal, or the stopper is shorter, the run halts on the very plateau that would have triggered the cut, and you never find out that a lower rate had another point of accuracy in it. A sound rule of thumb is to give the stopper enough patience to survive at least two reductions. ## Diagnosing it in the wild The signature is easy to read once you have seen it: the run always halts a few epochs after a schedule event, the stop epoch tracks the schedule rather than the data, and the recorded best epoch sits at the tail of the previous cycle. Change the cycle length and the stop epoch moves with it — that is the confirmation. A run that stops at genuinely varying points across seeds is stopping on the data; a run whose stop epoch is locked to the schedule is stopping on the schedule. ## What not to conclude Do not respond by making patience enormous everywhere. That turns off the mechanism for every run, including the ones where the model really is overfitting and you are paying for epochs you do not want. Fix the mismatch at its source — measure at the granularity the schedule actually operates on — rather than by inflating a number until the symptom disappears.

  • How would you confirm the stop is caused by the schedule rather than by overfitting?
    Check whether the stop epoch tracks schedule events. If the halt always lands a few epochs after a restart, and changing the cycle length moves the stop epoch with it, the schedule is the cause. Genuine overfitting stops at epochs that vary across seeds and follow the train/validation gap, not the rate curve.
  • You pair a plateau-triggered learning-rate cut with early stopping on the same metric. How should the two patience values relate?
    The stopper's patience must be clearly larger than the reducer's — ideally large enough to survive two reductions. Both counters watch the same signal, so if the stopper is equal or shorter it fires on the very plateau that would have triggered the cut, and you never see the improvement the lower rate usually buys.
  • Why is stopping halfway through an annealing schedule worse than running a shorter schedule from the start?
    Because an annealing schedule spends its early epochs at large step sizes and buys most of its final quality in the low-rate epochs at the end. Halting halfway exports a model that never received that refinement. A schedule sized to the shorter budget anneals within it and typically ends up better, at the same cost.

It is a smoke alarm wired into a kitchen: it goes off every time you cook, exactly as designed, and the fix is where you put the sensor, not how loud it is.

saying these in an interview costs you the question

  • Blames overfitting for a stop that tracks the schedule
  • Thinks a warm restart resets the weights or the best checkpoint
  • Sets patience to a huge value on every run to hide the symptom
  • Runs stopper and plateau-reducer with the same patience
  • Treats a mid-anneal stop as equivalent to a shorter schedule

context