How does RMSProp's decaying average of squared gradients differ from AdaGrad's running sum?
answer
- Sum versus average
- One can go back down
- Old gradients decay geometrically
- Window is one over one minus the decay
basics
~20 sRMSProp keeps a decaying average of recent squared gradients instead of AdaGrad's ever-growing total. Old gradients fade out, so the denominator can fall again and the step size recovers rather than shrinking toward zero for the whole run.
solid answer
~50 sBoth divide each parameter's step by the square root of a per-parameter squared-gradient estimate; only the estimate differs. AdaGrad uses `acc <- acc + g*g`, a sum, so the denominator only grows and every parameter's effective step size is non-increasing from start to finish. RMSProp uses `v <- rho*v + (1-rho)*g*g`, an exponential moving average with a decay rate `rho` usually between 0.9 and 0.99, so `v` can go down as well as up and a parameter whose gradients have quietened gets its step size back. As a rule of thumb `v` reflects the last `1/(1-rho)` steps: 0.9 is about a ten-step memory, 0.99 about a hundred. That makes RMSProp better suited to non-stationary objectives where the useful step size changes during training, at the cost of giving up the convex regret guarantee that AdaGrad's monotone decay supports.
go deeper
Be ready to say the difference in one sentence: a sum that never forgets versus a decaying average that does, with the same per-parameter division by the square root in both.
Explain the mechanics: unroll the moving average to show the geometric weights, give the memory length as one over one minus the decay rate, and say why an estimate that can fall means a step size that can recover.
Show the operational read. Relate a flat loss with live gradients to a denominator that can only grow, and be able to justify a decay rate from how fast the gradient scale actually moves in your run.
Own the tradeoff between a self-annealing optimizer with a convex guarantee and a forgetful one tuned for non-stationary objectives, including what late-run behaviour you accept when nothing shrinks the step size for you.
## The one-line difference AdaGrad and RMSProp use the same update shape — divide each parameter's step by the square root of a per-parameter estimate of its squared gradient — and differ only in how that estimate is maintained. AdaGrad sums; RMSProp averages with forgetting. ``` AdaGrad: acc <- acc + g * g RMSProp: v <- rho * v + (1 - rho) * g * g both: w <- w - lr * g / (sqrt(estimate) + tiny) ``` `rho` is the decay rate, a number strictly between 0 and 1, usually chosen somewhere between 0.9 and 0.99. `tiny` is a small positive constant guarding the division. Everything is elementwise, so each parameter carries its own estimate. ## What an exponential moving average actually stores Unrolling the recursion shows what `v` is. After many steps, ``` v_t = (1 - rho) * ( g_t^2 + rho * g_(t-1)^2 + rho^2 * g_(t-2)^2 + ... ) ``` The weight on a squared gradient from `k` steps ago is proportional to `rho^k`, so influence decays geometrically rather than being carried forever. The weights sum to one, which is why `v` is an average — an estimate of the *typical* recent squared gradient — while AdaGrad's `acc` is a total that grows with the number of steps. `sqrt(v)` is therefore a root-mean-square of recent gradients, which is where the name comes from. The standard rule of thumb for the memory length is `1 / (1 - rho)` steps: `rho = 0.9` behaves like an average over roughly the last ten steps, `rho = 0.99` roughly the last hundred, `rho = 0.999` roughly the last thousand. Equivalently, the contribution of a gradient decays by a factor `rho` per step, so a value from a hundred steps back is worth `0.9^100`, which is negligible, at `rho = 0.9`, but still about a third of full weight at `rho = 0.99`. ## Why the difference matters The decisive consequence is that `v` can go **down**. AdaGrad's accumulator is a sum of non-negative terms, so each parameter's effective learning rate `lr / (sqrt(acc) + tiny)` is non-increasing for the entire run: once a parameter has been through a stretch of large gradients, it is permanently slowed, even if the run later reaches a region where its gradients are small and it ought to move further per step. Long runs therefore end with a flat loss curve and gradient norms that have not collapsed — the optimizer's steps died, not the slope. RMSProp forgets. A parameter whose gradients have quietened sees `v` shrink over the next `1 / (1 - rho)` steps and gets its step size back. There is no built-in decay toward zero: the step size tracks the current gradient scale instead of the whole history. That is what makes RMSProp the better fit for non-stationary objectives, meaning objectives whose local geometry changes as training proceeds — which is the norm in deep networks, where curvature and gradient scale differ wildly between the start of training and a later, sharper region. There is a second, more subtle property. If a parameter's gradients are steady, `sqrt(v)` is approximately `|g|` and the normalised update `g / sqrt(v)` has magnitude about one. RMSProp's step is then approximately `lr` in absolute size regardless of the gradient scale — the learning rate sets a distance per step rather than a multiplier on a gradient. This is why RMSProp is comparatively robust to badly scaled losses, and also why a loss multiplied by a constant barely changes its behaviour. ## What is given up AdaGrad's monotone decay is what makes its convex online-learning regret bound work; a forgetful accumulator abandons that argument, and RMSProp was introduced as a practical heuristic rather than with a matching guarantee. Two practical consequences follow. First, RMSProp does not anneal itself: the step size does not shrink just because training has gone on for a long time, so late-run gradient noise is not damped for free the way AdaGrad damps it. Second, if a parameter goes through a long stretch of near-zero gradients, `v` can collapse toward zero and the next real gradient produces a step much larger than intended — a genuine instability that AdaGrad, whose denominator can only grow, does not have. ## Choosing the decay rate Treat `rho` as "how many steps of gradient-scale history are relevant". If the gradient scale changes quickly — a training regime that shifts, or data whose distribution moves — a lower `rho` tracks it and a high one lags by roughly its window. If the gradient signal is noisy batch to batch, a low `rho` makes the denominator itself noisy, which shows up as erratic step sizes. Most tuning lands between 0.9 and 0.99 for exactly this reason, and the choice is worth revisiting only when you can point at drift or noise on the scale of the implied window.
- How do you translate a decay rate of 0.9 versus 0.99 into a concrete number of steps?The weight on a squared gradient from k steps ago is proportional to `rho^k`, and the effective memory is about `1/(1-rho)` steps: roughly ten at 0.9, a hundred at 0.99. So 0.99 gives a smoother denominator but lags a real change in gradient scale by about a hundred steps, while 0.9 tracks quickly and is noisier.
- In a recommender retrained continuously as the item catalogue turns over, how does the decay rate matter?It sets how fast last month's squared-gradient statistics are forgotten. If the catalogue and traffic mix have genuinely changed, a high decay keeps stale scale estimates alive and the steps are sized for a distribution that no longer exists. Pick it against how many optimizer steps a real shift takes, and back off if the denominator itself starts looking noisy.
- Does RMSProp inherit AdaGrad's convergence guarantee?No. AdaGrad's regret bound comes from the convex online-learning setting and leans on a denominator that never decreases. RMSProp's accumulator can shrink, so that argument does not carry over; it was proposed as a practical heuristic and is judged empirically. In exchange it avoids the step size decaying to nothing on long non-convex runs.
AdaGrad is a lifetime odometer and RMSProp is a rolling average of your speed over the last few minutes. Only one of them can tell you that you have slowed down.
saying these in an interview costs you the question
- Says RMSProp is AdaGrad with a different learning rate
- Claims RMSProp's denominator also only grows
- Cannot say what the decay rate means in steps
- Thinks AdaGrad's decay is applied globally rather than per parameter
- Assumes RMSProp anneals its own step size late in training