In SGD with momentum, what does the velocity term add to a plain gradient step?
answer
- the step remembers the previous steps
- a running average, not one gradient
- consistent directions add up
- alternating directions cancel out
- memory length is one over one-minus-beta
basics
~20 sMomentum keeps a running average of past gradients, the velocity, and steps along it rather than along the raw gradient. Consistent directions build up and move faster; components that flip sign each step cancel, so the path stops zig-zagging.
solid answer
~50 sPlain stochastic gradient descent uses only the current mini-batch gradient: `w <- w - lr * g`. Heavy-ball momentum adds a velocity that remembers the past: `v <- beta * v + g`, then `w <- w - lr * v`. Since `v` is an exponentially weighted sum of past gradients decaying by `beta`, a component pointing the same way step after step accumulates until the effective step along it approaches `1/(1-beta)` times the plain step, while a component that alternates sign cancels itself. That is the fix for an ill-conditioned ravine, a gently sloping floor between walls about a hundred times more curved, where plain descent bounces wall to wall and creeps along the floor. Momentum damps the bouncing, accelerates the useful direction, and carries the run across near-flat saddle regions where the raw gradient is too small to move. `beta` sets how much history counts: roughly `1/(1-beta)` steps.
code
python · 15 lines# f(x, y) = 0.5 * (x**2 + 100 * y**2): gentle along x, steep along y
def grad(p):
return (p[0], 100.0 * p[1])
def run(lr, beta, steps=60):
p, v = [10.0, 1.0], [0.0, 0.0]
for _ in range(steps):
g = grad(p)
for i in (0, 1):
v[i] = beta * v[i] + g[i] # velocity: EWA of past gradients
p[i] -= lr * v[i]
return p
print('plain ', run(0.0015, 0.0)) # x barely moves: about 9.14
print('momentum', run(0.0015, 0.9)) # x reaches about 4.20go deeper
Be ready to write the two lines from memory: velocity equals the momentum coefficient times the old velocity plus the gradient, then step along the velocity. Say in one sentence why that damps zig-zag.
Explain the unrolled exponentially weighted sum, that the weights total one over one-minus-beta, and that a persistent direction therefore gets a step up to that factor larger at a fixed learning rate.
Show you read it off real curves: which loss-curve shape says momentum is helping, which says it is too aggressive, and how you keep the learning rate and momentum coefficient tuned together rather than one at a time.
Own the framing that momentum is the cheapest correction for ill-conditioning available without curvature information, and be able to say when a project should invest in something stronger instead of stacking more inertia.
## The update Plain stochastic gradient descent takes one mini-batch gradient `g` and moves against it: ``` w <- w - lr * g ``` Each step is amnesiac. Whatever the previous gradients said is thrown away. Heavy-ball momentum adds one state vector, the **velocity** `v`, with the same shape as the parameters: ``` v <- beta * v + g w <- w - lr * v ``` `beta` is the momentum coefficient, a number in [0, 1); common values sit near 0.9. A second, equally common convention normalises the average, `v <- beta * v + (1 - beta) * g`, which changes the scale of the step but not the idea. Everything below uses the first, undampened form and says so where the distinction matters. ## Why it is an exponentially weighted average Unroll the recursion. After many steps, ``` v_t = g_t + beta * g_(t-1) + beta^2 * g_(t-2) + beta^3 * g_(t-3) + ... ``` The velocity is a geometric blend of every gradient seen so far, with the most recent one weighted 1 and a gradient `k` steps ago weighted `beta^k`. Two consequences follow directly. **Memory length.** The weights sum to `1/(1-beta)`. At `beta = 0.9` that is 10, and the first ten terms sum to `(1 - 0.9^10)/0.1 = 6.51`, so about 65 percent of the current velocity comes from the last ten gradients. The useful mental model is that momentum averages over roughly `1/(1-beta)` recent steps. **Step amplification.** If the gradient were constant at `g`, the velocity would converge to `g/(1-beta)`, so the sustained step becomes `lr * g/(1-beta)` rather than `lr * g` — ten times larger at `beta = 0.9`. Momentum is therefore not a free smoother: at a fixed learning rate it also makes steps along persistent directions bigger. ## What it fixes: the ravine Take a two-parameter quadratic bowl whose curvature ratio is about 100 to 1 — a long, narrow valley: gentle along one axis, steep across it. The gradient is dominated by the steep direction, so any learning rate small enough to be stable across the valley is far too small to make progress along it. Plain descent oscillates from wall to wall while inching down the floor. Now look at what the two components do inside the velocity. Along the floor, every gradient points the same way, so the terms add and the effective step grows. Across the valley, the gradient flips sign on alternate steps, so consecutive terms subtract and the transverse component shrinks. One knob suppresses the oscillation and accelerates the descent at the same time, without needing per-direction curvature information. The same accumulation helps in a different situation: a near-flat plateau or saddle region, such as a regression network mapping a six-degree-of-freedom robot arm's tip pose back to joint angles, where the loss surface has long stretches with tiny gradients. Plain descent nearly stalls there; a velocity built up before entering the flat region coasts the iterate across it. ## What it does not do Momentum does not change the loss, does not add a penalty, and is not a regulariser in the weight-decay sense. It does not change where the gradient is zero, so it does not change the set of stationary points — only which one you reach and how fast. It is not a guarantee of escaping local minima; it is extra inertia that helps with small barriers and flat regions, nothing stronger. And it is not free of risk: because the step along persistent directions grows by up to `1/(1-beta)`, raising `beta` without lowering the learning rate can destabilise a previously healthy run. ## Reading it in practice Symptoms that momentum is doing its job: a loss curve that falls faster and looks less serrated than the same recipe with the momentum coefficient at zero, at the same learning rate. Symptoms that it is too aggressive: loss spikes early in training, or a run that overshoots and recovers repeatedly. The two knobs are coupled — the learning rate and the momentum coefficient jointly set the distance travelled per unit of gradient — so tuning one without touching the other is the usual source of surprise.
- At a momentum coefficient of 0.9, how much of the current velocity comes from the last ten gradients?About 65 percent. The gradient from `k` steps ago carries weight `beta^k`, the full series sums to `1/(1-beta) = 10`, and the first ten terms sum to `(1 - 0.9^10)/0.1 = 6.51`. So the practical reading of the momentum coefficient is an averaging horizon of roughly `1/(1-beta)` steps, with the tail beyond that contributing little.
- Momentum also helps on nearly flat regions of the loss surface. Why?On a plateau or a saddle region the raw gradient is tiny, so plain descent takes tiny steps and can stall for thousands of iterations. Velocity accumulated before entering the flat region does not vanish when the gradient does; it decays only as `beta^k`, so the run coasts across. This is why a network stuck at a flat loss can start moving again once momentum is turned on.
- Does momentum change which minima exist, or only how the run travels?Only how it travels. The velocity is built from the same gradients, so the stationary points of the loss are untouched — momentum changes the trajectory, the speed, and therefore which basin you end up in, not the landscape. That is why it is an optimizer choice rather than a modelling choice, and why it never appears in the loss function.
A ball rolling down a rutted valley. It does not stop and restart at each new patch of slope; it carries speed forward along the valley floor and its side-to-side wobble against the walls damps out.
saying these in an interview costs you the question
- Says momentum changes the loss function or acts as a penalty
- Calls momentum a learning-rate schedule
- Claims momentum only smooths noise and never enlarges the step
- Says momentum guarantees escaping local minima
- Thinks the velocity stores only the previous single gradient