Why can raising the momentum coefficient from 0.9 to 0.99 make a stable run diverge?
answer
- two knobs that multiply, not two independent ones
- one over one-minus-beta changes by ten
- the averaging window stretches to a hundred steps
- the velocity is built from stale gradients
- cut learning rate by the same ratio
basics
~20 sThe momentum coefficient and the learning rate are coupled. Going from 0.9 to 0.99 multiplies the sustained step by ten and stretches the gradient-averaging window to about a hundred steps, so an untouched learning rate is now far too large.
solid answer
~50 sMomentum is not an independent knob. With `v <- beta * v + g` and `w <- w - lr * v`, a persistent gradient drives the velocity to `g/(1-beta)`, so the sustained displacement per step scales as `lr/(1-beta)`: the factor moves from 10 to 100. An untouched learning rate is now effectively ten times too large. The averaging horizon stretches with it, from about ten steps to about a hundred, so the direction of travel is built from far staler gradients. On a surface whose curvature changes over tens of steps, the run charges into a sharply curved region carrying velocity the current gradient cannot brake, and the loss spikes. The rule when raising the coefficient is to scale the learning rate by the ratio of the `(1-beta)` values, about tenfold down here, then re-tune, plus warmup and gradient clipping, since the worst overshoot comes early when gradients are largest.
go deeper
Remember the headline: momentum and learning rate are not independent, so raising one usually means lowering the other. Knowing the direction of the adjustment is enough here.
Derive the sustained step lr * g/(1-beta) and show it grows tenfold between the two coefficients. State which velocity convention your derivation assumes.
Diagnose a real blow-up: separate the step-amplification story from the stale-gradient story, name the signals you would plot, and give a concrete recovery recipe with warmup and clipping.
Own the tuning protocol a team follows: which quantity is held invariant when either knob moves, how sweeps are budgeted, and what a recipe must record so a coefficient change is never made blind.
## The arithmetic first With the undampened heavy-ball recursion ``` v <- beta * v + g w <- w - lr * v ``` a gradient that keeps pointing the same way drives the velocity towards `g/(1-beta)`. The sustained step is therefore `lr * g/(1-beta)`. Evaluate the factor at the two coefficients: `1/(1-0.9) = 10` and `1/(1-0.99) = 100`. Changing only the momentum coefficient multiplied the effective step along any persistent direction by ten. Nobody would raise a learning rate tenfold on an untouched recipe and expect it to survive, and this is the same change wearing a different name. The averaging horizon moves with it. The velocity weights a gradient `k` steps old by `beta^k`, and the effective window is about `1/(1-beta)` steps: ten at 0.9, a hundred at 0.99. So the direction the run is travelling in is an average over a hundred iterations of history. ## Why staleness, not per-step instability, is usually the killer It is tempting to say that a higher coefficient shrinks the stable learning rate, but that is not what the linear analysis says. On a quadratic with maximum curvature `L`, the heavy-ball iteration is stable for `0 <= beta < 1` and `lr < 2(1+beta)/L` — the bound on the learning rate actually rises slightly with the coefficient. What a fixed quadratic cannot show you is the transient: how far the iterate travels while the velocity is still adapting, and what happens when the surface it travels into is not the surface the old gradients were measured on. That is the real failure mode. With a hundred-step memory, the velocity encodes the loss surface as it was up to a hundred steps ago. Deep-network loss surfaces change character quickly — a run leaves a flat region, enters a sharply curved one, or the scale of the activations shifts — and the velocity keeps pushing in the old direction for tens of steps after the gradient has begun objecting. The iterate is carried into a region of much larger loss, the gradient there is enormous, the velocity absorbs it, and the next steps are catastrophic. In the logs this reads as a sudden vertical loss spike, frequently followed by non-finite values, rather than a gentle upward drift. Early training is the worst window, because the initial gradients are the largest of the whole run and the velocity has nothing to cancel against. This is exactly where an ill-conditioned surface — say a long ravine with a curvature ratio near a hundred to one — punishes over-large inertia: the transverse component that momentum is supposed to cancel is briefly amplified before the alternation has had time to average out. ## Which convention you are in matters The dampened convention `v <- beta * v + (1 - beta) * g` normalises the weights, so a persistent gradient drives the velocity to `g` and the sustained step stays `lr * g` regardless of the coefficient. Under that convention raising the coefficient does *not* multiply the step, and the tenfold learning-rate cut is the wrong reflex. The memory horizon still stretches to a hundred steps, so the staleness argument still applies and the run can still overshoot or stall — but the diagnosis and the fix differ. Establishing which convention the code in front of you uses is the first question to ask, before touching any number. ## How to make the change safely 1. **Scale the learning rate with `(1-beta)`.** Under the undampened convention, multiply the learning rate by `(1-beta_new)/(1-beta_old)` — a factor of about 0.1 for 0.9 to 0.99 — as the starting point, then re-tune around it. This keeps the sustained step roughly invariant. 2. **Warm up.** Ramp the learning rate from near zero over the first few hundred to few thousand steps so the velocity never builds on the largest, least informative gradients of the run. 3. **Clip gradients by global norm.** This bounds what any single bad batch can inject into a velocity that will then influence a hundred subsequent steps. 4. **Watch the right signals.** Track the gradient norm and the velocity norm, not just the loss. A velocity norm that climbs monotonically while the loss is flat is the early warning; by the time the loss spikes the run is already lost. 5. **Change one thing.** Because the two knobs multiply, sweeping them independently wastes runs. If you must sweep, sweep the pair `(lr, beta)` along the constant-`lr/(1-beta)` curve first, then perpendicular to it. ## What a strong answer sounds like A senior answer names the coupling explicitly, gives the `1/(1-beta)` arithmetic, distinguishes the amplification story from the staleness story, states which convention the argument assumes, and lands on a concrete remedy rather than 'lower the learning rate a bit'. A weak answer says only that 0.99 is 'too much momentum' without being able to say by what factor or in what units.
- If the implementation uses the (1-beta)-dampened velocity, does the tenfold learning-rate cut still apply?No. With `v <- beta * v + (1 - beta) * g` the weights are normalised, so a persistent gradient drives the velocity to `g` and the sustained step stays `lr * g` at any coefficient. There is nothing to compensate for. The averaging horizon still stretches to about a hundred steps, so staleness and overshoot are still live risks, but the remedy is warmup and clipping rather than a proportional learning-rate cut.
- What would you watch in the logs to confirm momentum, rather than the data, caused the blow-up?Plot the velocity norm alongside the gradient norm and the loss. Momentum-driven divergence shows the velocity norm climbing while the loss is still flat or improving, then a loss spike; a data problem shows a gradient-norm spike on a specific batch with a velocity that only follows. Re-running the same batch order at the old coefficient and seeing no spike settles it.
- Why is early training the most dangerous window for a high momentum coefficient?The gradients are largest and least consistent at initialisation, and the velocity starts at zero with nothing to cancel against, so it accumulates fast in whatever direction the first batches happen to point. With a hundred-step memory those early, uninformative gradients steer the run long after they stopped being relevant, which is exactly what a learning-rate warmup is there to prevent.
- How would you sweep the learning rate and the momentum coefficient together rather than separately?Treat `lr/(1-beta)` as the quantity that actually sets the sustained step and sweep along constant values of it first, which tells you the right step scale. Then vary the coefficient at fixed `lr/(1-beta)`, which isolates the effect of the averaging horizon. An independent grid over both wastes most of its runs on combinations whose effective step is wildly off.
saying these in an interview costs you the question
- Treats the momentum coefficient as independent of the learning rate
- Says higher momentum always converges faster
- Cannot state the one-over-one-minus-beta factor
- Ignores which velocity convention the code uses
- Blames the data for a loss spike without checking velocity norms