Optimizers and Learning-Rate Schedules
What each optimizer adds over plain SGD, why AdamW decouples weight decay, and how warmup and decay steady a run. Interviewers probe it to see if you reason about hyperparameters or copy defaults.
on this pageshowhide
explore
- Stochastic Steps and Momentum11 questions
- Mini-Batch Gradient Noise4 questions
- Momentum and Nesterov3 questions
- Weight Averaging4 questions
- Adaptive Update Rules14 questions
- AdaGrad and RMSProp3 questions
- Adam Moment Estimates4 questions
- Decoupled Weight Decay3 questions
- Adam Versus SGD Momentum4 questions
- Schedules and Batch Size13 questions
- Learning-Rate Warmup4 questions
- Annealing and Restarts5 questions
- Batch-Size Scaling Rules4 questions
questions
page 2 of 2How do AdaGrad and RMSProp differ on rare parameters in a mostly-sparse embedding table?
basics
~20 sAdaGrad's sum only advances when a parameter actually receives a gradient, so rarely-seen rows keep large steps. RMSProp's average decays on every step, so a long-idle row's denominator collapses and its next update is far larger than intended.
When switching from Adam to SGD with momentum at epoch 30, how do you choose the new learning rate?
basics
~20 sDo not carry the adaptive rate across. Adam's steps are roughly the size of its learning rate whatever the gradient, so momentum using that rate barely moves. Pick a rate one to two orders larger.
When resuming training from a step-50,000 checkpoint, should the learning-rate warmup be replayed?
basics
~20 sNormally no - resume the schedule at step 50,000, ramp already finished. The exception is a resume that restored only weights: with the optimizer's running averages starting from zero again, a short re-ramp is cheap insurance.
How should Adam's second-moment decay rate be set when rare batches carry enormous gradients?
basics
~10 sThe second-moment decay sets how long one batch influences the denominator: 0.999 remembers about a thousand steps, 0.98 about fifty. A lower rate spikes harder on an outlier but forgets it far sooner.
Under AdamW with a cosine-annealed learning rate, should the decay term shrink along with the rate?
basics
~20 sIn the common formulation the decay step is multiplied by the current learning rate, so annealing the rate quietly anneals regularization too. Decide that deliberately: either hold the decay-to-step ratio fixed on purpose, or apply a rate-independent decay and accept the different endgame.
How large should a training batch get before extra parallel compute stops paying off?
basics
~10 sUp to a critical batch size, doubling the batch roughly halves the steps to a target loss, so parallelism becomes wall-clock savings. Past it, the step count flattens while compute per step keeps doubling.
If you inject Gaussian noise into a full-batch gradient to imitate a small-batch run, what fails to transfer?
basics
~20 sYou can match the magnitude of the noise but not its structure. Mini-batch noise is shaped by how examples disagree and shrinks as the model fits, while fixed isotropic Gaussian noise points everywhere equally and never fades, so the run keeps wandering and the loss floors out.
Five checkpoints from one run: ensemble their predictions or collapse them into one averaged weight vector?
basics
~20 sPrediction ensembling usually scores highest but multiplies inference cost and memory by five. Collapsing the same checkpoints into one averaged weight vector keeps single-model serving cost and captures part of the gain. Decide from the serving budget, then measure both.
showing 31–38 of 38