When resuming training from a step-50,000 checkpoint, should the learning-rate warmup be replayed?
answer
- the schedule is a function of the step counter
- checkpoint the counter, not just the weights
- restarting the whole schedule is the real bug
- weights-only resume rebuilds the averages from zero
- spike that decays over a few hundred steps
basics
~20 sNormally no - resume the schedule at step 50,000, ramp already finished. The exception is a resume that restored only weights: with the optimizer's running averages starting from zero again, a short re-ramp is cheap insurance.
solid answer
~40 sThe default is to restore the step counter with the weights and continue the schedule where it left off, which means past the ramp. Replaying the ramp is mildly wasteful; the far more damaging bug is restarting the *whole* schedule, so a run that should be decaying near the end of its budget returns to a full-rate ramp and loses much of what it had. The exception is what the checkpoint contained. If it held only weights and the optimizer's moment estimates start from zero, the early-step problem is back - now on well-trained weights, where a mis-scaled step destroys more - so a few hundred steps of re-ramp is worth it. The same applies if the resume changed batch size, precision, or the data pipeline.
go deeper
Remember that the learning-rate schedule depends on the step number, so a resumed run needs to know which step it is on - not just which weights to load.
Explain what a checkpoint has to contain for a clean resume - weights, optimizer state, and the step counter - and what each one is for.
Show the operating judgment: continue the schedule by default, re-ramp briefly when optimizer state is missing or the input configuration changed, and read the first few hundred post-resume steps to confirm which case you are in.
Own the policy and the tooling around it - one written re-ramp rule, checkpoints that carry all three pieces as a unit, and a logged rate-and-step assertion on every resume so a broken restart is caught in minutes rather than at the next evaluation.
## The default answer A training schedule is a function of the step counter. If you checkpoint the counter along with the weights and the optimizer state, resuming is simply continuing that function: at step 50,000 the ramp finished long ago, and the resumed run picks up whatever rate the decay schedule specifies for step 50,000. No replay. The most valuable habit here is checkpointing the counter at all. A surprising number of resume bugs come down to a checkpoint that holds weights and optimizer state but no notion of where in the schedule the run was, so the resumed process starts its schedule from zero. ## The bug that actually costs you Replaying just the ramp is a minor inefficiency: a few hundred steps at a reduced rate on a well-trained model, and it costs almost nothing. The expensive mistake is restarting the *entire* schedule. A run that was 80 percent through its budget, sitting at a small decayed rate, is suddenly ramped back up to the full target rate and then re-decayed over another full budget's worth of steps. Well-annealed weights get shaken back out; the loss climbs and takes a long time to come back, if it does. When someone says "the resume made it worse," this is nearly always what happened, and the fix is to persist and restore the step counter rather than to add compensating hacks. ## When a re-ramp is genuinely the right call Ask what the checkpoint contained and what changed across the boundary. **Only weights were restored.** Momentum and the per-coordinate second-moment estimates start from zero again. That returns you to the opening-steps regime: an average over a handful of samples, an unstable denominator in the adaptive update, and a first step whose size is set by the learning rate rather than by the gradient. Worse, the weights are no longer random - they encode a lot of training, and a mis-scaled step now destroys real progress rather than shuffling noise. A short ramp, on the order of a few hundred steps up to the schedule's current rate, is cheap and removes the risk. **Something changed across the resume.** A different batch size, a change in numerical precision, a rebuilt input pipeline that alters the data order or the distribution of sequence lengths, a changed number of workers contributing gradients - each moves the gradient statistics the optimizer's averages were built on. A brief re-ramp lets the averages catch up to the new regime. **Nothing changed and full state was restored.** Continue. There is nothing to warm up; the statistics are exactly as valid as they were one step before the crash. ## Reading the resume The diagnostic is the first few hundred steps after the resume, logged unsmoothed. A loss spike that appears immediately and decays back over a few hundred steps is the signature of missing optimizer state - the run is re-accumulating its averages the hard way. A loss that jumps and stays high, or climbs steadily, points instead at a schedule restarted from zero or at a resumed rate that does not match where the run left off. A loss that continues its previous trend with no discontinuity at all is what a correct resume looks like, and it is worth asserting on: log the rate and the step counter on the first step after every resume, and compare against the last step before the crash. ## The practice this implies Checkpoint the step counter, the optimizer state, and the weights as one unit, and treat a checkpoint missing any of the three as a weights-only checkpoint with the re-ramp policy that implies. Make the resumed rate an explicit, logged value rather than something inferred from a restarted process. And decide the re-ramp rule once - "re-ramp for a few hundred steps whenever optimizer state was not restored or the input configuration changed, otherwise continue" - so nobody has to reason it out at 2 a.m. after a node failure.
- Why is a mis-scaled step more damaging at step 50,000 than at step 1?At initialization the weights carry no information, so an oversized step mostly rearranges noise and the run can still find its way. At step 50,000 the weights encode everything learned so far, and a step large enough to disrupt them throws away real work that took hours to accumulate. The same nominal step size therefore carries a much higher expected cost late in a run.
- What does an immediate loss spike after a resume that recovers within a few hundred steps tell you?It points at optimizer state that was not restored. The running averages are rebuilding from zero, so the first steps are effectively unwarmed, and the recovery time matches the averaging window. A spike that does not recover, or a loss that plateaus high, points somewhere else - usually a schedule restarted from step zero or a mismatched resumed rate.
- Should the re-ramp on a weights-only resume go up to the target rate or to the current scheduled rate?To the current scheduled rate for that step. The point of the re-ramp is to protect the opening steps while the optimizer's averages rebuild, not to undo the decay the run has already earned. Ramping back to the original target would be a restart of the schedule in disguise, which is the failure mode you are trying to avoid.
saying these in an interview costs you the question
- Replays the entire schedule from step zero on every resume
- Checkpoints weights and optimizer state but not the step counter
- Says a resume never needs a ramp because the model is already trained
- Ignores that a weights-only restore leaves the running averages at zero
- Re-ramps back up to the original target rate instead of the current scheduled rate