skip to content

Schedules and Batch Size

How the step size moves over a run, through warmup, decay and restarts, and how it must move when the batch size does. Interviewers probe it because the learning rate dominates every other knob.

on this pageshow

explore

questions

13

Why is a neural network's learning rate usually decayed over the course of training?

level: juniorimportance: must knowfreq 70%

answer

  1. the step scales signal and noise alike
  2. near a minimum, only noise is left
  3. constant rate implies a loss floor
  4. high rate early explores, low rate settles
  5. sum of steps infinite, squares finite

basics

~20 s

A large step keeps stochastic gradient descent bouncing around a minimum instead of settling in it. Shrinking the step later in training lets the noisy gradient estimates average out, so the parameters settle and the loss stops oscillating.

solid answer

~50 s

Each update is `w <- w - lr * g`, and `g` is a mini-batch estimate, so it carries noise that the learning rate scales along with the signal. Near a minimum the true gradient is nearly zero but the noise is not, so a constant rate leaves the parameters wandering in a region whose size grows with the step. That region is a loss floor: the run stops improving because the step is too big to settle, not because the model cannot fit better. Decaying shrinks the region and lowers the floor, which is why the training curve drops right after each cut. The early high-rate phase still earns its keep — it makes fast progress and behaves like a regularizer — so the goal is to spend time at a high rate first and then cool down, not to start small.

go deeper

for a junior

Be ready to say in one breath that a big step keeps the parameters bouncing near a minimum and a smaller step lets them settle, and to recognise the drop in a training curve right after a rate cut.

for a middle

Explain the mechanism: the step multiplies both the true gradient and the mini-batch noise, so a constant rate leaves a wandering region whose size grows with the rate. Name both failure modes, never decaying and decaying too soon.

for a senior

Show you use the rate cut as a diagnostic on real runs — cut and watch, to separate a step-size plateau from a capacity limit — and that you treat the long high-rate phase as regularisation you are deliberately buying.

for a principal

Own the framing that the schedule, not the optimizer choice, is usually the dominant knob in a training recipe, and that schedules must be specified against a budget so results across teams remain comparable.

## What the learning rate actually multiplies A gradient step is `w <- w - lr * g`. The quantity `g` is not the true gradient of the training loss; it is an estimate computed on one mini-batch, so it equals the full-dataset gradient plus a random error. The learning rate multiplies the whole thing, which means it scales the useful signal and the noise by exactly the same factor. Every argument about annealing follows from that one fact. ## The noise floor Far from a minimum, the signal dominates: the true gradient is large, the noise is a perturbation, and a big step buys fast progress. Close to a minimum the balance inverts. The true gradient goes to zero, the mini-batch noise does not, and each update injects a random displacement proportional to the step size while the curvature of the loss pulls the parameters back toward the bottom. The two forces reach a balance, and the parameters end up wandering inside a region around the minimum rather than sitting at it. The size of that region grows with the learning rate. Treating stochastic gradient descent as a discretisation of a noisy continuous process gives the same picture: for a fixed noise level, the excess loss the run hovers above the true minimum grows roughly in proportion to the step size. So a constant learning rate implies a loss floor. When a run plateaus with a constant rate, it usually has not run out of capacity or data; it has run out of resolution. Cutting the rate shrinks the region and lowers the floor. This is why a step-decayed run shows a visible cliff in the training loss immediately after each cut: the same weights, evaluated after a few steps at a tenth of the step size, sit closer to the bottom of the same basin. ## Two jobs, early and late The schedule is doing two different jobs at two different times. - **Early — travel and explore.** A large rate crosses a badly conditioned landscape quickly, and the noise it amplifies acts as a search mechanism, letting the run leave narrow regions it happened to land in. Runs that spend a long stretch at a high rate typically generalise better than runs that are decayed immediately, so the high-rate phase functions as a regularizer, not just as a hurry. - **Late — refine.** Once the trajectory is in the region it is going to stay in, the remaining job is to sit down in it. That needs small steps and nothing else. A schedule is simply a plan for switching between those two jobs over the budget you have. ## The classical condition Stochastic approximation theory states the tension precisely. For convergence you want the step sizes to satisfy two conditions at once: they must sum to infinity, so the run can still travel an arbitrary distance no matter where it starts, and their squares must sum to a finite value, so the accumulated noise is bounded. A constant rate satisfies the first and fails the second — hence the noise floor. A rate that collapses too fast satisfies the second and fails the first — the run freezes wherever it happens to be, and the final loss is worse than a slower decay would have reached. Deep networks are non-convex, so this is a guide rather than a guarantee, but it names the two failure modes correctly. ## The two failure modes in practice - **Never decaying.** Training and validation loss plateau above what the model can reach, and metrics jitter from epoch to epoch. Cutting the rate at that point produces an immediate improvement, which is the diagnostic. - **Decaying too early or too aggressively.** The run looks excellent for a few epochs, then flatlines well above where it should. The parameters were locked into place before they had finished travelling. This one is more expensive because it is invisible without a comparison run. ## A mental model Treat the learning rate as a temperature. High temperature explores and refuses to commit; low temperature crystallises whatever configuration it is handed. Annealing is a planned cooling path, and the reason it is planned rather than greedy is that the useful exploration happens before there is any evidence it was useful. ## Two things it drags along with it First, the noise in `g` depends on how many examples each estimate averages, so the same schedule can behave differently when that changes. Second, in optimizers that apply weight decay as a separate term scaled by the same schedule multiplier, annealing the rate also anneals the effective strength of that decay — the regularisation weakens exactly as the rate does. Neither is a reason to skip annealing; both are reasons to describe a schedule as part of the recipe rather than as a detail.

  • How would you tell from a training curve that a run is sitting on a noise floor rather than out of capacity?
    Cut the learning rate by a factor of ten and watch the next few hundred steps. If the loss drops sharply and then flattens at a lower level, the previous plateau was a step-size effect, not a capacity limit. If almost nothing happens, the model really has fitted what it can and the fix is elsewhere — more capacity, better features, or cleaner data.
  • Is starting at a small learning rate and never decaying a safe compromise?
    No. It gives up the exploratory high-rate phase that makes runs generalise well, converges far more slowly, and still leaves a noise floor — a smaller one, but a floor. You end up paying for the worst of both: slow travel and no clean settling. The cheap version of a schedule is a high rate held for most of the budget and then annealed, not a low rate everywhere.
  • Does the argument change if the gradients were exact rather than mini-batch estimates?
    Largely, yes. With exact gradients there is no noise floor, and a fixed step below the stability limit converges on its own. You would still reduce the step near the end for badly conditioned curvature or to satisfy a line-search condition, but the dominant reason to anneal in deep learning — that the step size sets how much sampling noise survives into the parameters — disappears.

Annealing metal works the same way: hold it hot so the atoms can rearrange, then cool it on a planned schedule so it settles into a good structure. Quench it immediately and you freeze in whatever mess was there.

saying these in an interview costs you the question

  • Says decay exists only to make training faster
  • Claims a plateau always means the model lacks capacity
  • Starts at a tiny rate to be safe and never decays
  • Thinks decay changes the loss surface rather than the step
  • Decays from step one, skipping the exploratory phase

context

open as a page

How does cosine annealing differ from step decay when both run over a fixed epoch budget?

level: middleimportance: must knowfreq 58%

basics

~20 s

Step decay holds the rate flat and multiplies it by a fixed factor at chosen epochs, leaving visible cliffs in the loss. Cosine annealing slides smoothly from the peak to near zero across the whole budget, whose length fixes the shape.

open as a page

What is the linear scaling rule for the learning rate when the mini-batch size grows?

level: middleimportance: must knowfreq 68%

basics

~10 s

Multiply the learning rate by the same factor you multiply the batch size by. A run at batch 256 with rate 0.1 moved to batch 8,192 scales both by 32, giving rate 3.2.

open as a page

What is linear learning-rate warmup, and what failure in the opening training steps does it prevent?

level: middleimportance: must knowfreq 72%

basics

~20 s

Warmup ramps the learning rate from near zero up to its target over the first few hundred to few thousand steps. It prevents the large, poorly informed updates a freshly initialized network takes at full rate, which can spike or destroy the run.

open as a page

When is square-root learning-rate scaling preferred over linear scaling as the batch grows?

level: middleimportance: should knowfreq 44%

basics

~20 s

Square-root scaling, multiplying the learning rate by the square root of the batch-size factor, is the usual starting point with adaptive per-coordinate optimizers, where full linear scaling overshoots. Linear scaling stays the default for momentum SGD.

open as a page

When does a reduce-on-plateau rule cut a learning rate that should have been left alone?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A reduce-on-plateau rule fires whenever its monitored metric is noisy enough to sit flat past the patience window by chance. Because the rule only ever reduces, that cut is permanent even if the metric improves next epoch.

open as a page

In a warm-restart schedule, why does the training loss get worse right after the rate jumps back up?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A restart raises the step back to the peak while keeping the current weights. Those large steps push the parameters out of the narrow low-loss region they had settled into, so the loss rises before the next anneal brings it down.

open as a page

Your team quadrupled the batch size, kept the epoch budget, and the run got worse — why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Two effects are being confused. A four-fold batch at the same epoch count means four times fewer updates, and an unscaled learning rate leaves each update the same size, so the run underfits. Scale the rate, then judge.

open as a page

Why does Adam still need learning-rate warmup if it already normalizes each coordinate's step?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Adam divides by a running estimate of each coordinate's squared gradient, and in the opening steps that estimate comes from a handful of samples. Its noise, not its average, is the problem: some coordinates get steps far larger than intended.

open as a page

How do you set warmup length - a fixed step count or a fraction of the run - when porting a recipe to a far larger dataset?

level: principalimportance: should knowfreq 38%

basics

~20 s

Warmup's job is finished after a certain number of optimizer steps, not after a share of the data, so a fixed step count is the defensible unit. A percentage silently stretches tenfold on a tenfold dataset.

open as a page

In the one-cycle schedule, why is the momentum coefficient scheduled opposite to the learning rate?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Momentum multiplies the effective step, which scales as the rate over one minus the coefficient. Lowering the coefficient at the peak rate keeps that effective step stable, and raising it again as the rate anneals keeps the low-rate tail moving.

open as a page

When resuming training from a step-50,000 checkpoint, should the learning-rate warmup be replayed?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Normally no - resume the schedule at step 50,000, ramp already finished. The exception is a resume that restored only weights: with the optimizer's running averages starting from zero again, a short re-ramp is cheap insurance.

open as a page

How large should a training batch get before extra parallel compute stops paying off?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

Up to a critical batch size, doubling the batch roughly halves the steps to a target loss, so parallelism becomes wall-clock savings. Past it, the step count flattens while compute per step keeps doubling.

open as a page