skip to content

Optimizers and Learning-Rate Schedules

What each optimizer adds over plain SGD, why AdamW decouples weight decay, and how warmup and decay steady a run. Interviewers probe it to see if you reason about hyperparameters or copy defaults.

on this pageshow

explore

questions

page 2 of 2

How do AdaGrad and RMSProp differ on rare parameters in a mostly-sparse embedding table?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

AdaGrad's sum only advances when a parameter actually receives a gradient, so rarely-seen rows keep large steps. RMSProp's average decays on every step, so a long-idle row's denominator collapses and its next update is far larger than intended.

open as a page

When switching from Adam to SGD with momentum at epoch 30, how do you choose the new learning rate?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Do not carry the adaptive rate across. Adam's steps are roughly the size of its learning rate whatever the gradient, so momentum using that rate barely moves. Pick a rate one to two orders larger.

open as a page

When resuming training from a step-50,000 checkpoint, should the learning-rate warmup be replayed?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Normally no - resume the schedule at step 50,000, ramp already finished. The exception is a resume that restored only weights: with the optimizer's running averages starting from zero again, a short re-ramp is cheap insurance.

open as a page

How should Adam's second-moment decay rate be set when rare batches carry enormous gradients?

level: principalimportance: nice to knowfreq 26%

basics

~10 s

The second-moment decay sets how long one batch influences the denominator: 0.999 remembers about a thousand steps, 0.98 about fifty. A lower rate spikes harder on an outlier but forgets it far sooner.

open as a page

Under AdamW with a cosine-annealed learning rate, should the decay term shrink along with the rate?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

In the common formulation the decay step is multiplied by the current learning rate, so annealing the rate quietly anneals regularization too. Decide that deliberately: either hold the decay-to-step ratio fixed on purpose, or apply a rate-independent decay and accept the different endgame.

open as a page

How large should a training batch get before extra parallel compute stops paying off?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

Up to a critical batch size, doubling the batch roughly halves the steps to a target loss, so parallelism becomes wall-clock savings. Past it, the step count flattens while compute per step keeps doubling.

open as a page

If you inject Gaussian noise into a full-batch gradient to imitate a small-batch run, what fails to transfer?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

You can match the magnitude of the noise but not its structure. Mini-batch noise is shaped by how examples disagree and shrinks as the model fits, while fixed isotropic Gaussian noise points everywhere equally and never fades, so the run keeps wandering and the loss floors out.

open as a page

Five checkpoints from one run: ensemble their predictions or collapse them into one averaged weight vector?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Prediction ensembling usually scores highest but multiplies inference cost and memory by five. Collapsing the same checkpoints into one averaged weight vector keeps single-model serving cost and captures part of the gain. Decide from the serving budget, then measure both.

open as a page

showing 31–38 of 38