In a warm-restart schedule, why does the training loss get worse right after the rate jumps back up?
answer
- warm means weights are kept
- only the schedule resets
- peak step is huge for a settled region
- track the envelope, not the sawtooth
- always end on a completed cycle
basics
~20 sA restart raises the step back to the peak while keeping the current weights. Those large steps push the parameters out of the narrow low-loss region they had settled into, so the loss rises before the next anneal brings it down.
solid answer
~50 sWarm restarts run a sequence of annealing cycles — say a cosine over ten epochs, then twenty, then forty — and at each boundary the rate snaps back to the peak. Warm means only the schedule restarts: the weights and the optimizer's accumulated state carry over, nothing is re-initialised. The previous cycle ended at a tiny rate, so the parameters sat in a narrow region; the peak rate is far too large for it, so the first steps after the restart throw them out of it and the loss spikes. That is the mechanism, and it is intended — the spike is the exploration you are paying for. What matters is the envelope, the loss at the *end* of each cycle, which should be flat or improving. Two operational rules follow: always finish on a completed cycle, and if the envelope climbs across cycles the peak is too high.
go deeper
Be ready to state what restarts do at all: the learning rate goes back up to its peak on a fixed cycle, while the model keeps the weights it has learned so far.
Explain the mechanism of the spike — a peak-size step is far too large for the narrow region a fully annealed model sits in — and describe how geometric cycle lengths place the restart points.
Demonstrate that you operate these runs: track the end-of-cycle envelope rather than the sawtooth, plan the budget so it lands on a cycle boundary, and diagnose a rising envelope as a peak rate that is too high.
Own the call of whether restarts are worth their cost at all for a given budget, and set the convention that reported numbers come from completed cycles so results across a team mean the same thing.
## The schedule A warm-restart schedule chops training into cycles. Inside each cycle the rate anneals from a peak down to a small floor, typically along a cosine. At the end of a cycle the rate is set straight back to the peak and the next cycle begins. Cycle lengths usually grow geometrically: an initial cycle of ten epochs with a doubling multiplier gives cycles of 10, 20 and 40 epochs, so restarts land at epochs 10, 30 and 70. The word *warm* is the whole point of the design and the thing candidates most often get wrong. Nothing about the model restarts. The weights carry over exactly as they were at the end of the previous cycle, and the optimizer's accumulated state carries over with them. Only the learning-rate schedule is reset. A cold restart — re-initialising the weights — would be a fresh training run, which is a completely different and far more expensive thing. ## Why the loss rises At the end of a cycle the rate is near its floor, so the parameters have settled tightly into whatever low-loss region they were in. Then the rate snaps back to a value perhaps a hundred times larger. That step size is enormous relative to the width of the region the parameters are sitting in, so the first handful of updates carry them out of it. Loss rises, sometimes to a level well above where the cycle started. This is not a bug and it is not instability in the numerical sense. It is the mechanism: the run buys back the exploration that annealing spent, then anneals again from a different starting point. The reason a practitioner tolerates the spike is that the next descent often lands somewhere at least as good, and sometimes better, than the previous cycle — and it gets there in the same run, without re-paying the cost of early training. ## Reading a restart run The raw training curve of a restart schedule is a sawtooth and is nearly useless read literally. The quantity to track is the **envelope**: the loss (or validation metric) at the end of each cycle, when the rate has annealed to its floor and the model is in a comparable state to any other end-of-cycle point. - **Envelope flat or improving** — the schedule is doing its job. Later cycles are longer, so they get more annealing time and should usually be at least as good. - **Envelope climbing** — the restarts are destroying more than they find. The usual cause is a peak rate that is too high for the current stage of training; decaying the peak across cycles, or lowering it outright, is the fix. - **Envelope flat but the spikes get bigger** — cosmetic; the spike height depends on the peak and the sharpness of the region, not on progress. ## Operational consequences **Finish on a completed cycle.** This is the single most common practical mistake with restarts. If the budget runs out mid-cycle, you are holding a model whose rate is still high and which has not been annealed — strictly worse than the model you had at the end of the previous cycle. Because cycle lengths double, the last cycle is as long as everything before it combined, so a budget that overshoots by a little can leave a very large unfinished cycle. Plan the cycle structure so the cycle boundaries and the budget coincide. **A restart is not free.** Every spike costs steps spent re-descending ground already covered. On a short budget a single anneal is usually the better use of the compute; restarts earn their keep on longer runs, or when you genuinely want multiple annealed models out of one run. **The peak can be decayed too.** Nothing requires every cycle to return to the same peak. Multiplying the peak by a factor below one at each restart gives a schedule that explores hard early and increasingly refuses to leave later, which is often a better match to what the run needs. ## What it is not A restart is not a way to recover from divergence — if a run has blown up, raising the rate makes it worse. It is not early stopping, which is a rule about when to halt rather than how the rate moves. And it is not a plateau rule: restarts are scheduled in advance from the cycle structure, not triggered by a metric, so the run is reproducible from the configuration alone.
- With an initial cycle of ten epochs doubling each time, where do the restarts fall and how does that constrain your budget?Cycles of 10, 20 and 40 epochs put restarts at epochs 10, 30 and 70, with the run naturally finishing at 70. Because each cycle is as long as everything before it, budgets that are not one of those boundaries strand you mid-cycle with an un-annealed model. Either choose a budget that lands on a boundary or change the initial length or multiplier so that it does.
- How do you tell a healthy restart spike from a peak rate that is too high?Compare end-of-cycle values, not the peaks of the sawtooth. If each cycle finishes at least as low as the previous one, the spikes are the price of exploration and the schedule is working. If end-of-cycle loss drifts upward across cycles, the restarts are destroying progress the annealing cannot recover, and the peak should be lowered or decayed at each restart.
- Would you use warm restarts on a three-epoch budget?No. Each restart spends steps re-descending ground already covered, and on a very short budget there are not enough steps to amortise that. A single anneal over the whole budget will beat a chopped-up one. Restarts pay off when the run is long enough that the exploration bought by a spike has time to turn into a lower end-of-cycle loss.
saying these in an interview costs you the question
- Thinks a restart re-initialises the model weights
- Reads the sawtooth peaks instead of end-of-cycle values
- Stops the run mid-cycle and reports that model
- Calls the post-restart spike a divergence to be fixed
- Uses restarts to recover a run that has already blown up