skip to content

When switching from Adam to SGD with momentum at epoch 30, how do you choose the new learning rate?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the rate is on a different scale after the switch
  2. adaptive step size is about the rate itself
  3. momentum step scales with the raw gradient
  4. match recent step magnitudes, or sweep briefly
  5. expect a recoverable bump, buffer starts empty

basics

~20 s

Do not carry the adaptive rate across. Adam's steps are roughly the size of its learning rate whatever the gradient, so momentum using that rate barely moves. Pick a rate one to two orders larger.

solid answer

~50 s

The two rules put the learning rate on different scales. Adam divides the averaged gradient by the root of its averaged square, so the per-parameter step is roughly the learning rate itself; momentum multiplies the raw gradient by the rate, and late-run gradients are small. Reusing Adam's rate therefore stalls the run. Two practical ways to set it: estimate the typical per-parameter step Adam was taking just before the switch and divide by the typical gradient magnitude, which lands you near the equivalent momentum rate; or run a short sweep of a few candidates for a few hundred steps each and keep the one whose loss resumes falling without a spike. Expect a visible loss bump for a few hundred steps because the preconditioner and momentum buffer are gone, and keep the remaining decay schedule running. Whether the switch beats simply training momentum from the start is an empirical question you should test.

go deeper

for a junior

Know that the learning rate is not portable between the two optimizers, and that a run which stalls right after an optimizer change usually has a rate that is far too small.

for a middle

Be able to explain why an adaptive step is roughly the size of the learning rate while a momentum step scales with the raw gradient, and what that implies for the new rate's order of magnitude.

for a senior

Show the operational plan: how you estimate or sweep the new rate, what transient you expect and how you tell a healthy bump from a bad rate, and how the remaining schedule continues.

for a principal

Weigh the extra hyperparameters and failure modes of a two-phase recipe against the point or so it buys, and decide whether a team maintaining many models should carry that complexity at all.

## Why the rate cannot be carried over The reason a mid-run switch is dangerous is a scale mismatch, and it is worth being precise about it. Under Adam, the update for a parameter is its averaged gradient divided by the square root of its averaged squared gradient, times the learning rate. Numerator and denominator carry the same units, so the ratio is roughly order one and the update magnitude is roughly the learning rate — a rate of 1e-3 means parameters move by about 1e-3 per step, whether their gradients are 1e-5 or 10. Under momentum the update is the learning rate times the momentum-averaged raw gradient. By epoch 30 gradients are typically much smaller than they were at initialization. If you hand that rule a rate of 1e-3, the resulting parameter movement is 1e-3 times a small number, which is far smaller than the steps the run was just taking. The loss stops improving and it looks like the switch "broke" training, when in fact the optimizer is barely moving. ## Two ways to pick the new rate **Match the step magnitude.** Just before the switch, record the typical magnitude of the parameter updates Adam is applying and the typical magnitude of the raw gradients. The momentum rate that reproduces the same movement is approximately the first divided by the second. This gives a principled starting point rather than a guess, and it is usually one to two orders of magnitude above the adaptive rate. A published proposal, SWATS, automates a version of this idea by estimating the equivalent momentum rate from the adaptive updates themselves, projecting the adaptive step onto the gradient direction. **Short empirical sweep.** Branch the checkpoint, run each of a handful of candidate rates for a few hundred steps, and keep the one whose training loss resumes a steady decline without an immediate spike. This costs little because the runs are short, and it is more robust than any formula when the model has unusual parameter groups. ## What to expect at the switch step Several things reset at once, so a transient is normal: - The per-parameter preconditioner disappears. Parameters that were being given inflated steps because their gradients were small now get proportionally small steps, which is the intended change but also a discontinuity. - The momentum buffer starts empty. The first steps after the switch use an under-accumulated average, so the effective step is smaller than steady state for the first few dozen updates. A brief ramp of the rate over a few hundred steps smooths this. - Training loss usually rises for a short window and then resumes falling. A bump that recovers within a few hundred steps is expected; a bump that does not recover means the rate is too high, and a flat line means it is too low. Keep the remaining decay schedule intact after the switch. The reason to switch at all is to get momentum's end-of-run decay behaviour, so switching and then leaving the rate constant throws away most of the point. ## When to switch, and whether to switch The usual trigger is the end of the fast-progress phase: switch once the adaptive run's validation curve has flattened and before the final decay phase, so that momentum owns the part of training where the final solution is decided. Treat both the switch epoch and the new rate as hyperparameters, because they are. And keep the baseline honest. The switch adds two hyperparameters and an operational failure mode, in exchange for a gap that is often around a point. Before adopting it, compare against the two simple alternatives on the same budget: adaptive all the way through, and momentum from the first step with a tuned rate and schedule. If momentum from scratch matches the switched run, ship the simpler recipe. The switch earns its complexity mainly when the early phase genuinely needs adaptivity — an unstable start, or parameter groups with very uneven gradient scales — and the final number still needs to be the best available.

  • What should the loss curve do in the first few hundred steps after the switch?
    A short rise that then resumes falling is normal, because the preconditioner is gone and the momentum buffer starts empty. A spike that does not recover means the new rate is too high; a flat curve at the pre-switch loss means it is far too low. Ramping the rate over the first few hundred steps softens the transient.
  • How do you decide the epoch at which to switch?
    Switch once the adaptive phase has stopped buying you much — the training curve's steep descent is over — and before the final decay phase, so momentum shapes the solution you keep. Treat the switch point as a tuned hyperparameter, and always compare against not switching at all, since the simpler recipe often matches it.

saying these in an interview costs you the question

  • Keeps the adaptive learning rate after switching to momentum
  • Calls the post-switch loss bump a bug and reverts
  • Leaves the rate constant after switching, dropping the decay
  • Assumes the momentum buffer carries over from the adaptive state
  • Adopts the switch without comparing against momentum from scratch

context