What do warmup and cosine decay do in a fine-tuning learning-rate schedule?
answer
- the rate is a path, not a number
- the first steps are the risky ones
- ramp up, then ease down
- 3–5% ramp, then a half-cosine tail
- no warmup shows as an early loss spike
basics
~20 sWarmup ramps the learning rate from near zero up to its peak over the first few percent of steps so early updates cannot destabilise the model. Cosine decay then lowers it smoothly back toward zero so late steps refine instead of overwrite.
solid answer
~50 sA schedule is the path the learning rate takes over the run; the rate people quote is only its **peak**. Warmup covers the first few percent of steps — often 3–5%, or a fixed few dozen steps on a short run — ramping linearly from near zero to the peak. It exists because the very first updates are the riskiest: the optimizer's running statistics are estimated from one or two batches, and the new task's data distribution can produce unusually large gradients, so a full-size step at step one is where runs blow up. **Cosine decay** then eases the rate down along a half-cosine curve to zero or to a small floor. Late in training you want small, refining steps; a rate still at its peak on the final batch leaves the model wherever the last few examples pushed it. Constant-with-warmup is the usual alternative when the total step count is not known in advance.
code
python · 11 linesimport math
def lr_at_step(step, total_steps, peak_lr, warmup_frac=0.03, min_lr=0.0):
warmup_steps = max(1, int(total_steps * warmup_frac))
if step < warmup_steps:
return peak_lr * (step + 1) / warmup_steps
progress = (step - warmup_steps) / max(1, total_steps - warmup_steps)
return min_lr + 0.5 * (peak_lr - min_lr) * (1 + math.cos(math.pi * progress))
for s in (0, 10, 30, 500, 999):
print(s, round(lr_at_step(s, 1000, 2e-4), 8))go deeper
Be able to say plainly that warmup ramps the rate up at the start and decay eases it down at the end, and that the quoted learning rate is the peak of that curve.
Explain the reasons: unreliable optimizer statistics and distribution shift make the first steps risky, and small late steps refine rather than overwrite. Give typical numbers such as a 3–5% warmup fraction.
Diagnose from the loss curve — an early spike as missing warmup, a mid-run rise as too high a peak — and pick the schedule to match the run, including constant-after-warmup when the step count is open-ended.
Standardise the schedule across the team's training recipes so runs are comparable, and be explicit about the interaction between schedule length, early stopping and checkpoint selection when budgets are cut mid-run.
## The schedule is the whole path, not one number When someone says a run used "2e-4", they mean the peak learning rate. What the optimizer actually sees is a curve: a ramp up, a plateau or immediate turn, and a decay. Two runs with the same peak but different schedules can behave completely differently — one stable, one diverging in the first hundred steps. This is why a reported learning rate is only comparable between runs when the schedule shape and the effective batch size are stated too. ## Warmup: why the first steps are the dangerous ones Adaptive optimizers such as AdamW keep running estimates of gradient magnitude. At step one those estimates are built from a single batch, so the per-parameter step sizes they imply are poorly grounded. At the same time, the fine-tuning data usually differs in distribution from what the model saw in pretraining, which produces larger-than-usual gradients on the first batches. Taking a full-size step under both conditions is the most likely moment for a run to destabilise. Warmup removes that risk cheaply by scaling the rate linearly from near zero to the peak over the first slice of the run. During the ramp the optimizer's statistics settle, the model absorbs the distribution shift gradually, and by the time full-size steps arrive nothing is being estimated from a single batch. The classic symptom of removing warmup from a run that otherwise worked is a **loss spike in the first few dozen steps** — loss jumps well above its starting value, then either recovers slowly at a worse level or diverges outright. If you see that shape, warmup length is the first thing to change. ## How long to warm up Common choices are a fraction of total steps (3–5% is typical) or a fixed count (tens of steps) on short runs. Longer warmup is a cheap insurance policy: with a few thousand total steps, spending a hundred on the ramp costs almost nothing. Very long warmup is not free either — it eats the budget at rates too small to make progress — but the failure it causes is mild underfitting rather than a wrecked run, which is why practitioners err long. ## Cosine decay and its alternatives After warmup, cosine decay follows a half-cosine from peak down to a floor, usually zero or around 10% of peak. The shape holds the rate near the peak for a while, then falls steeply through the middle, then flattens as it approaches the floor. That final flat tail is the useful part: the last stretch of training makes small refining updates instead of large ones, which typically yields a slightly better and much more reproducible final checkpoint than stopping while the rate is still high. Alternatives you will be asked to compare: - **Linear decay** — a straight line from peak to zero. Behaves similarly; slightly more time at low rates, slightly less near the peak. - **Constant after warmup** — the right choice when the total number of steps is unknown (open-ended training, or when you plan to stop on a validation signal), because both cosine and linear need the endpoint to compute the curve. - **Cosine with restarts** — the rate jumps back up periodically. It belongs to long pretraining-scale runs and is rarely useful for a short fine-tune. A schedule that decays to zero also interacts with early stopping: if you stop halfway through a cosine planned for 1,000 steps, you stop while the rate is still large, and the checkpoint is worse than a run that was *planned* to be that short. Set the schedule to the number of steps you actually intend to run. ## Reading the curve when it goes wrong - **Spike in the first tens of steps** → warmup too short or absent, or peak too high. - **Loss falls then rises through the middle** → peak too high for the effective batch size; the decay tail may partly hide it at the end. - **Loss descends and flattens far above where you expected** → peak too low, especially with adapter training at a full fine-tuning rate. - **Loss still dropping steeply when the run ends** → the schedule finished before the model did; add steps rather than raising the peak. ## What to say in an interview Warmup protects the beginning of the run; decay protects the end. Both are about *where in the run* a given step size is appropriate, not about the total amount of learning. The peak rate decides how aggressive the run is; the schedule decides when that aggression is applied.
- When would you use a constant schedule after warmup instead of cosine decay?When the total number of steps is not known up front. Both cosine and linear decay need an endpoint to compute the curve, so open-ended training, or a run you intend to cut on a validation signal, is better served by a constant rate after warmup. The cost is a slightly worse final checkpoint, since you lose the low-rate refining tail that decay provides.
- What goes wrong if you stop a run early while a long cosine schedule is still mid-decay?You stop while the learning rate is still large, so the last updates were coarse and the checkpoint is noisier than one from a run planned to end at that step. The fix is to set the schedule length to the number of steps you actually intend to run, and treat early stopping as a safety net rather than the normal exit path.
- Does a longer warmup let you push the peak learning rate higher?Somewhat, but not indefinitely. Warmup removes the specific instability of the first steps, so it lets a rate survive an opening that would otherwise diverge. It does not make an inherently too-large step size safe for the rest of the run — that shows up later as loss rising through the middle of training. Treat warmup as protection for the start, not as a licence to raise the peak.
Warmup is easing off the clutch instead of dropping it; cosine decay is lifting off the throttle as you approach the parking spot rather than braking at the last metre.
saying these in an interview costs you the question
- Thinks warmup increases the learning rate above the configured peak
- Says the schedule changes how much the model learns overall
- Uses cosine decay without knowing the total step count
- Skips warmup and blames the resulting loss spike on bad data
- Reports a learning rate without saying it is the peak