skip to content

Why can Adam reach a lower training loss faster than SGD with momentum yet finish with worse validation accuracy?

level: middleimportance: must knowfreq 72%

answer

  1. speed on train is not score on val
  2. per-parameter rescaling changes the destination
  3. the gap appears after the decay phase
  4. equal training loss, worse held-out number

basics

~20 s

Adam gives every parameter its own step size from its squared-gradient history, so training loss drops fast early. That rescaling also changes which solution you reach, and tuned momentum with a decay schedule often ends higher on validation.

solid answer

~50 s

Adam divides each parameter's averaged gradient by the root of that parameter's own averaged squared gradient, so directions with small or badly scaled gradients still move a full step. Early in training that is exactly what you want, and the training loss falls faster than a momentum run at the same wall-clock. But per-parameter rescaling does not just change the speed of the trip, it changes the destination: the weights you end up with are not the ones momentum would have reached. On a 100-class convnet image classifier it is a familiar pattern that Adam owns the first five epochs of training loss while a tuned SGD-with-momentum run, given a real decay schedule and the full epoch budget, finishes roughly a point and a half higher on validation. The lesson for the interview is that early training loss is not the metric. Compare final validation numbers, tuned equally, or the comparison says nothing.

go deeper

for a junior

Know that Adam adapts a step size per parameter and momentum uses one global rate, and that a faster-falling training curve is not by itself evidence of a better model.

for a middle

Be ready to explain why dividing by the root of the averaged squared gradient makes early steps uniform in size, and why that changes which minimizer the run settles on rather than only how fast it gets there.

for a senior

Show that you would not accept the claim without a full-schedule validation comparison, and name what usually explains a reported gap: no decay schedule, a rate inherited from the other optimizer, or an early stopping point.

for a principal

Own the framing that optimizer choice is a budget decision. Decide whether the last point of accuracy is worth the tuning cycles momentum demands across every model your team trains, and say so explicitly rather than defaulting either way.

## The two update rules, in one line each SGD with momentum keeps one running average of the gradient and takes a step proportional to it: the same scalar learning rate multiplies every parameter's gradient. A direction with a tiny gradient gets a tiny step; a direction with a large gradient gets a large one. The raw magnitude of the gradient is preserved in the step. Adam keeps two running averages per parameter: one of the gradient and one of the squared gradient. The step is the first average divided by the square root of the second. Because numerator and denominator both scale with the size of that parameter's gradients, the ratio is roughly order one, and the actual step size is roughly the learning rate regardless of whether the gradient was 1e-6 or 1e2. Adam therefore throws away most of the gradient's magnitude information and keeps mainly its direction and its consistency. ## Why that makes early progress fast A fresh deep network has wildly different gradient scales across its layers and parameter groups. Under one global learning rate, that rate must be small enough for the loudest parameters not to diverge, which leaves the quiet ones barely moving for many epochs. Adam removes that coupling: every parameter gets a step of comparable size in units of the learning rate, so the whole network starts making progress on step one. This is why an Adam run at a sensible default rate almost always beats an untuned momentum run on the first few epochs of training loss, and why it is so forgiving of a learning rate that was never really tuned. ## Why fast is not the same as good The crucial point, and the one candidates most often miss, is that this is not an optimization-speed story with a fixed endpoint. Both runs are minimizing a non-convex objective with many near-equivalent minimizers of the training loss. The update rule decides which one you drift into. Adam's rescaling systematically enlarges the steps taken in directions where gradients are small and rare, so the final weight vector differs in both direction and scale from what a momentum run produces. On several standard image-classification setups, published comparisons found that adaptive methods reached training loss as low as, or lower than, momentum while landing at a slightly worse held-out number. That is the signature of a generalization difference, not an optimization failure: equal or better training loss, worse validation. The mechanism behind that difference is still argued about in the literature, and you should say so rather than assert a settled cause. The defensible interview position is empirical: on some architectures and datasets the tuned momentum run ends higher, the gap is small (about a point, sometimes less), and it only shows up at the end of a full schedule. ## What actually shrinks the gap Most reported gaps are smaller than they look once the comparison is done properly: - **A decay schedule.** A momentum run held at a constant rate will lose. Momentum needs its rate decayed toward the end of the run; that final decay phase is where most of its validation advantage appears, which is precisely why an epoch-five comparison inverts the ranking. - **A real rate search for both.** Momentum's best rate is often one to two orders of magnitude larger than Adam's, so inheriting one optimizer's rate for the other guarantees a meaningless result. - **Long enough training.** The ordering can flip after the schedule's decay. A comparison stopped before both schedules complete measures the wrong thing. - **Regularization tuned per optimizer.** Weight decay strength interacts with the step-size rule, so a value copied across optimizers is a confound rather than a control. ## How to talk about it If someone tells you "Adam converged faster," the correct response is a question, not agreement: faster to what, measured on which split, against a momentum run tuned by whom? A convergence-speed claim about training loss and an accuracy claim about held-out data are two different claims, and only the second one decides what ships. State the phenomenon plainly, note that it is architecture- and budget-dependent, and refuse to convert a training curve into a model-quality verdict.

  • Is this an optimization failure or a generalization failure?
    A generalization failure. Adam typically reaches a training loss as low as, or lower than, the momentum run, so it is not failing to minimize the objective. The difference appears only on held-out data, which means the two update rules are settling on genuinely different solutions rather than one of them stalling.
  • How long must both runs train before the comparison means anything?
    Both must run their full schedule through the decay phase. Momentum earns most of its final validation number in the last portion of the run when its rate is decayed, so any comparison truncated before that systematically favours the adaptive run. Comparing at a fixed early epoch measures which optimizer starts faster, which nobody disputes.
  • Your Adam run is ahead on validation at epoch 5 and behind at epoch 90. What do you report?
    Report the epoch-90 numbers as the result and the epoch-5 numbers as a note about warm-up behaviour. The deliverable is the model you ship, which is the end of the schedule. The early crossover is still useful information: it tells you Adam is the better choice if your real budget is only a handful of epochs.

Adam is a driver who takes every corner at the same speed regardless of the road; momentum reads the road surface. The first driver is ahead for the first ten minutes and arrives at a slightly worse address.

saying these in an interview costs you the question

  • Says Adam is strictly better because the loss drops faster
  • Reports convergence speed with no validation number attached
  • Assumes the optimizer only changes speed, not the final solution
  • Compares Adam's default rate against an untuned momentum rate
  • Treats a lower training loss as proof of a better model

context