skip to content

Optimizers and Learning-Rate Schedules

What each optimizer adds over plain SGD, why AdamW decouples weight decay, and how warmup and decay steady a run. Interviewers probe it to see if you reason about hyperparameters or copy defaults.

on this pageshow

explore

questions

page 1 of 2

Why is a neural network's learning rate usually decayed over the course of training?

level: juniorimportance: must knowfreq 70%

answer

  1. the step scales signal and noise alike
  2. near a minimum, only noise is left
  3. constant rate implies a loss floor
  4. high rate early explores, low rate settles
  5. sum of steps infinite, squares finite

basics

~20 s

A large step keeps stochastic gradient descent bouncing around a minimum instead of settling in it. Shrinking the step later in training lets the noisy gradient estimates average out, so the parameters settle and the loss stops oscillating.

solid answer

~50 s

Each update is `w <- w - lr * g`, and `g` is a mini-batch estimate, so it carries noise that the learning rate scales along with the signal. Near a minimum the true gradient is nearly zero but the noise is not, so a constant rate leaves the parameters wandering in a region whose size grows with the step. That region is a loss floor: the run stops improving because the step is too big to settle, not because the model cannot fit better. Decaying shrinks the region and lowers the floor, which is why the training curve drops right after each cut. The early high-rate phase still earns its keep — it makes fast progress and behaves like a regularizer — so the goal is to spend time at a high rate first and then cool down, not to start small.

go deeper

for a junior

Be ready to say in one breath that a big step keeps the parameters bouncing near a minimum and a smaller step lets them settle, and to recognise the drop in a training curve right after a rate cut.

for a middle

Explain the mechanism: the step multiplies both the true gradient and the mini-batch noise, so a constant rate leaves a wandering region whose size grows with the rate. Name both failure modes, never decaying and decaying too soon.

for a senior

Show you use the rate cut as a diagnostic on real runs — cut and watch, to separate a step-size plateau from a capacity limit — and that you treat the long high-rate phase as regularisation you are deliberately buying.

for a principal

Own the framing that the schedule, not the optimizer choice, is usually the dominant knob in a training recipe, and that schedules must be specified against a budget so results across teams remain comparable.

## What the learning rate actually multiplies A gradient step is `w <- w - lr * g`. The quantity `g` is not the true gradient of the training loss; it is an estimate computed on one mini-batch, so it equals the full-dataset gradient plus a random error. The learning rate multiplies the whole thing, which means it scales the useful signal and the noise by exactly the same factor. Every argument about annealing follows from that one fact. ## The noise floor Far from a minimum, the signal dominates: the true gradient is large, the noise is a perturbation, and a big step buys fast progress. Close to a minimum the balance inverts. The true gradient goes to zero, the mini-batch noise does not, and each update injects a random displacement proportional to the step size while the curvature of the loss pulls the parameters back toward the bottom. The two forces reach a balance, and the parameters end up wandering inside a region around the minimum rather than sitting at it. The size of that region grows with the learning rate. Treating stochastic gradient descent as a discretisation of a noisy continuous process gives the same picture: for a fixed noise level, the excess loss the run hovers above the true minimum grows roughly in proportion to the step size. So a constant learning rate implies a loss floor. When a run plateaus with a constant rate, it usually has not run out of capacity or data; it has run out of resolution. Cutting the rate shrinks the region and lowers the floor. This is why a step-decayed run shows a visible cliff in the training loss immediately after each cut: the same weights, evaluated after a few steps at a tenth of the step size, sit closer to the bottom of the same basin. ## Two jobs, early and late The schedule is doing two different jobs at two different times. - **Early — travel and explore.** A large rate crosses a badly conditioned landscape quickly, and the noise it amplifies acts as a search mechanism, letting the run leave narrow regions it happened to land in. Runs that spend a long stretch at a high rate typically generalise better than runs that are decayed immediately, so the high-rate phase functions as a regularizer, not just as a hurry. - **Late — refine.** Once the trajectory is in the region it is going to stay in, the remaining job is to sit down in it. That needs small steps and nothing else. A schedule is simply a plan for switching between those two jobs over the budget you have. ## The classical condition Stochastic approximation theory states the tension precisely. For convergence you want the step sizes to satisfy two conditions at once: they must sum to infinity, so the run can still travel an arbitrary distance no matter where it starts, and their squares must sum to a finite value, so the accumulated noise is bounded. A constant rate satisfies the first and fails the second — hence the noise floor. A rate that collapses too fast satisfies the second and fails the first — the run freezes wherever it happens to be, and the final loss is worse than a slower decay would have reached. Deep networks are non-convex, so this is a guide rather than a guarantee, but it names the two failure modes correctly. ## The two failure modes in practice - **Never decaying.** Training and validation loss plateau above what the model can reach, and metrics jitter from epoch to epoch. Cutting the rate at that point produces an immediate improvement, which is the diagnostic. - **Decaying too early or too aggressively.** The run looks excellent for a few epochs, then flatlines well above where it should. The parameters were locked into place before they had finished travelling. This one is more expensive because it is invisible without a comparison run. ## A mental model Treat the learning rate as a temperature. High temperature explores and refuses to commit; low temperature crystallises whatever configuration it is handed. Annealing is a planned cooling path, and the reason it is planned rather than greedy is that the useful exploration happens before there is any evidence it was useful. ## Two things it drags along with it First, the noise in `g` depends on how many examples each estimate averages, so the same schedule can behave differently when that changes. Second, in optimizers that apply weight decay as a separate term scaled by the same schedule multiplier, annealing the rate also anneals the effective strength of that decay — the regularisation weakens exactly as the rate does. Neither is a reason to skip annealing; both are reasons to describe a schedule as part of the recipe rather than as a detail.

  • How would you tell from a training curve that a run is sitting on a noise floor rather than out of capacity?
    Cut the learning rate by a factor of ten and watch the next few hundred steps. If the loss drops sharply and then flattens at a lower level, the previous plateau was a step-size effect, not a capacity limit. If almost nothing happens, the model really has fitted what it can and the fix is elsewhere — more capacity, better features, or cleaner data.
  • Is starting at a small learning rate and never decaying a safe compromise?
    No. It gives up the exploratory high-rate phase that makes runs generalise well, converges far more slowly, and still leaves a noise floor — a smaller one, but a floor. You end up paying for the worst of both: slow travel and no clean settling. The cheap version of a schedule is a high rate held for most of the budget and then annealed, not a low rate everywhere.
  • Does the argument change if the gradients were exact rather than mini-batch estimates?
    Largely, yes. With exact gradients there is no noise floor, and a fixed step below the stability limit converges on its own. You would still reduce the step near the end for badly conditioned curvature or to satisfy a line-search condition, but the dominant reason to anneal in deep learning — that the step size sets how much sampling noise survives into the parameters — disappears.

Annealing metal works the same way: hold it hot so the atoms can rearrange, then cool it on a planned schedule so it settles into a good structure. Quench it immediately and you freeze in whatever mess was there.

saying these in an interview costs you the question

  • Says decay exists only to make training faster
  • Claims a plateau always means the model lacks capacity
  • Starts at a tiny rate to be safe and never decays
  • Thinks decay changes the loss surface rather than the step
  • Decays from step one, skipping the exploratory phase

context

open as a page

In SGD with momentum, what does the velocity term add to a plain gradient step?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Momentum keeps a running average of past gradients, the velocity, and steps along it rather than along the raw gradient. Consistent directions build up and move faster; components that flip sign each step cancel, so the path stops zig-zagging.

open as a page

In AdaGrad, how does the running sum of squared gradients set each parameter's step size?

level: middleimportance: must knowfreq 58%

basics

~20 s

AdaGrad divides the global learning rate by the square root of each parameter's own accumulated squared gradients. Consistently large gradients get small steps, quiet parameters keep large ones, and because the sum only grows, every step shrinks over time.

open as a page

Why does Adam divide its moment estimates by one minus the decay rate raised to the step count?

level: middleimportance: must knowfreq 64%

basics

~20 s

Both moving averages start at zero, so early on they read too small — the squared-gradient average far more so. Dividing each by one minus its decay rate to the step count removes that startup bias.

open as a page

What two running averages does Adam maintain, and how does its update rule combine them?

level: middleimportance: must knowfreq 82%

basics

~20 s

Adam keeps two exponential moving averages per parameter: one of the gradient, one of the squared gradient. The update is the first divided by the square root of the second, so each parameter gets its own step size.

open as a page

Why can Adam reach a lower training loss faster than SGD with momentum yet finish with worse validation accuracy?

level: middleimportance: must knowfreq 72%

basics

~20 s

Adam gives every parameter its own step size from its squared-gradient history, so training loss drops fast early. That rescaling also changes which solution you reach, and tuned momentum with a decay schedule often ends higher on validation.

open as a page

In Adam, what changes when weight decay is applied directly to the weights instead of added to the gradient?

level: middleimportance: must knowfreq 65%

basics

~20 s

Adding an L2 term to the gradient sends that term through the adaptive per-parameter denominator, so each weight ends up with a different effective shrinkage. Decoupled decay subtracts a fixed fraction of the weight itself, shrinking every weight at the same relative rate.

open as a page

How does cosine annealing differ from step decay when both run over a fixed epoch budget?

level: middleimportance: must knowfreq 58%

basics

~20 s

Step decay holds the rate flat and multiplies it by a fixed factor at chosen epochs, leaving visible cliffs in the loss. Cosine annealing slides smoothly from the peak to near zero across the whole budget, whose length fixes the shape.

open as a page

What is the linear scaling rule for the learning rate when the mini-batch size grows?

level: middleimportance: must knowfreq 68%

basics

~10 s

Multiply the learning rate by the same factor you multiply the batch size by. A run at batch 256 with rate 0.1 moved to batch 8,192 scales both by 32, giving rate 3.2.

open as a page

What is linear learning-rate warmup, and what failure in the opening training steps does it prevent?

level: middleimportance: must knowfreq 72%

basics

~20 s

Warmup ramps the learning rate from near zero up to its target over the first few hundred to few thousand steps. It prevents the large, poorly informed updates a freshly initialized network takes at full rate, which can spike or destroy the run.

open as a page

Why is a mini-batch gradient an unbiased estimate of the full-batch gradient, and what shrinks its noise?

level: middleimportance: must knowfreq 72%

basics

~20 s

A uniformly sampled mini-batch gradient has the same expected value as the full-batch gradient, so it is unbiased. Only its noise differs: the standard deviation of that noise falls roughly as one over the square root of the batch size.

open as a page

Why does an exponential moving average of a network's weights often evaluate better than the live weights?

level: middleimportance: must knowfreq 55%

basics

~20 s

Late in training the weights bounce inside a noise ball around a minimum rather than sitting in it. An exponential moving average cancels much of that mini-batch noise, so the averaged copy lands nearer the basin's centre and generalizes better.

open as a page

How does RMSProp's decaying average of squared gradients differ from AdaGrad's running sum?

level: juniorimportance: should knowfreq 66%

basics

~20 s

RMSProp keeps a decaying average of recent squared gradients instead of AdaGrad's ever-growing total. Old gradients fade out, so the denominator can fall again and the step size recovers rather than shrinking toward zero for the whole run.

open as a page

What breaks in the mini-batch gradient when a label-sorted file is streamed unshuffled?

level: juniorimportance: should knowfreq 52%

basics

~20 s

The mini-batch gradient stops being an unbiased estimate. Each batch holds one class, so its expected gradient is that class's gradient, not the dataset's. That is bias, not noise, and no batch size fixes it.

open as a page

When is square-root learning-rate scaling preferred over linear scaling as the batch grows?

level: middleimportance: should knowfreq 44%

basics

~20 s

Square-root scaling, multiplying the learning rate by the square root of the batch-size factor, is the usual starting point with adaptive per-coordinate optimizers, where full linear scaling overshoots. Linear scaling stays the default for momentum SGD.

open as a page

How does Nesterov's look-ahead gradient differ from heavy-ball momentum?

level: middleimportance: should knowfreq 52%

basics

~20 s

Heavy-ball measures the gradient at the current parameters. Nesterov measures it at the point the momentum part of the step alone would reach, so an impending overshoot is braked one iteration earlier. Same cost, same velocity recursion.

open as a page

Why would you raise the epsilon in Adam's denominator from a tiny default to a much larger value?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Epsilon puts a floor under Adam's denominator, capping how large a step a tiny second moment can produce. Raise it when gradients have become genuinely small and the optimizer is amplifying noise into full-size, jittery updates.

open as a page

A teammate says Adam beat SGD with momentum after tuning only Adam's rate. What do you challenge?

level: seniorimportance: should knowfreq 52%

basics

~10 s

Challenge the tuning asymmetry first: an optimizer given a rate search cannot be compared with one whose rate was inherited. Then demand schedule parity, per-optimizer regularization tuning, equal epoch budgets, and multiple seeds.

open as a page

Porting a recipe from Adam with L2 in the loss to AdamW, why re-tune the weight-decay value?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The coefficient does not carry over. With the L2 term inside the gradient, the adaptive denominator rescales the penalty before it lands, so the same number produces very different shrinkage once decay is decoupled — in practice the decoupled value usually has to be much larger.

open as a page

When does a reduce-on-plateau rule cut a learning rate that should have been left alone?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A reduce-on-plateau rule fires whenever its monitored metric is noisy enough to sit flat past the patience window by chance. Because the rule only ever reduces, that cut is permanent even if the metric improves next epoch.

open as a page

In a warm-restart schedule, why does the training loss get worse right after the rate jumps back up?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A restart raises the step back to the peak while keeping the current weights. Those large steps push the parameters out of the narrow low-loss region they had settled into, so the loss rises before the next anneal brings it down.

open as a page

Your team quadrupled the batch size, kept the epoch budget, and the run got worse — why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Two effects are being confused. A four-fold batch at the same epoch count means four times fewer updates, and an unscaled learning rate leaves each update the same size, so the run underfits. Scale the rate, then judge.

open as a page

Why does Adam still need learning-rate warmup if it already normalizes each coordinate's step?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Adam divides by a running estimate of each coordinate's squared gradient, and in the opening steps that estimate comes from a handful of samples. Its noise, not its average, is the problem: some coordinates get steps far larger than intended.

open as a page

Why can a batch-32 run walk off a loss plateau that a full-batch run sits on for thousands of steps?

level: seniorimportance: should knowfreq 44%

basics

~20 s

On a plateau the averaged gradient is nearly zero, so a deterministic step is nearly zero and the run creeps. Individual examples still disagree, so a small batch keeps producing non-zero steps that explore directions the average cancels out, and one of them leads downhill.

open as a page

Why can raising the momentum coefficient from 0.9 to 0.99 make a stable run diverge?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The momentum coefficient and the learning rate are coupled. Going from 0.9 to 0.99 multiplies the sustained step by ten and stretches the gradient-averaging window to about a hundred steps, so an untouched learning rate is now far too large.

open as a page

Why does averaging the weights of two independently initialized runs produce a broken model?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The two runs sit in unrelated regions of a non-convex loss surface, with hidden units learned in different orders and different arrangements. Coordinate-wise averaging mixes unrelated features, and the midpoint has high loss. Averaging only works along one trajectory.

open as a page

After averaging weights, why must you re-estimate a network's normalization running statistics?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The averaged weight vector never produced any activations during training, so the running mean and variance stored by normalization layers belong to different weights. One extra forward-only pass over training data recomputes them; skipping it can crater accuracy.

open as a page

Why are speech and translation transformer stacks trained adaptively while vision convnet recipes still ship plain momentum?

level: principalimportance: should knowfreq 42%

basics

~20 s

The split tracks architecture, not modality. Attention stacks mix parameter groups with wildly different gradient scales and rarely updated embeddings, which one global rate handles badly. Normalized convnets are better conditioned, so tuned momentum wins the last point.

open as a page

How do you set warmup length - a fixed step count or a fraction of the run - when porting a recipe to a far larger dataset?

level: principalimportance: should knowfreq 38%

basics

~20 s

Warmup's job is finished after a certain number of optimizer steps, not after a share of the data, so a fixed step count is the defensible unit. A percentage silently stretches tenfold on a tenfold dataset.

open as a page

In the one-cycle schedule, why is the momentum coefficient scheduled opposite to the learning rate?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Momentum multiplies the effective step, which scales as the rate over one minus the coefficient. Lowering the coefficient at the peak rate keeps that effective step stable, and raising it again as the rate anneals keeps the low-rate tail moving.

open as a page

showing 1–30 of 38