Why does Adam still need learning-rate warmup if it already normalizes each coordinate's step?
answer
- the normalizer is an estimate
- at step one the ratio is plus or minus one
- bias correction fixes the mean, not the spread
- small denominator, oversized step
- averaging window sets the ramp scale
basics
~20 sAdam divides by a running estimate of each coordinate's squared gradient, and in the opening steps that estimate comes from a handful of samples. Its noise, not its average, is the problem: some coordinates get steps far larger than intended.
solid answer
~50 sAdam's per-coordinate update is roughly `m_hat / (sqrt(v_hat) + eps)`, which makes the step scale-free - the size of the move is set by the learning rate, not by how big the gradient is. At the first step that is literally true: with one sample, the ratio is plus or minus one, so every parameter moves by the full learning rate no matter how small its gradient. Over the next few dozen steps the second-moment estimate is an average over very few samples, so it has high variance. The standard bias correction fixes the estimate's average but does nothing about its spread, and a coordinate whose gradients happened to be small in those few samples gets a small denominator and therefore an outsized effective step. Warmup keeps the rate small until enough samples have accumulated for the normalizer to be stable, which is why a sensible default ramp is on the order of the second-moment averaging window.
go deeper
Recall that Adam keeps running averages of the gradient and of the squared gradient, and that averages need several samples before they mean anything.
Be able to explain the update as a ratio of first moment to the square root of second moment, and why that makes the step size depend on the learning rate rather than on gradient magnitude.
Demonstrate the bias-versus-variance distinction and the one-sided risk of an underestimated denominator, and connect ramp length to the averaging window instead of quoting a fixed step count.
Own the reliability argument: an under-warmed adaptive run fails stochastically, so a shared recipe should carry a ramp sized with margin rather than tuned to the shortest value that happened to survive one launch.
## The intuition that makes this question interesting The natural objection is that an adaptive method should not need warmup at all. The whole point of dividing by a running root-mean-square of the gradient is that the step no longer depends on gradient magnitude: a coordinate with huge gradients and a coordinate with tiny gradients both move by something on the order of the learning rate. If gradient scale is already handled, why does the opening of a run still blow up? The answer is that the normalizer is an *estimate*, and early in the run it is estimated from almost no data. ## What the update actually does at step one Adam maintains two exponential moving averages: a first moment `m` of the gradient and a second moment `v` of the squared gradient. The parameter update is proportional to the ratio of the (bias-corrected) first moment to the square root of the (bias-corrected) second moment, plus a small constant in the denominator for numerical safety. At the very first step there is exactly one gradient sample `g`. Both averages are built from that one sample, so the ratio reduces to `g / |g|` - the sign of the gradient. Every single parameter in the network moves by the full learning rate, in the direction of its gradient's sign, regardless of whether that gradient was 1e-6 or 1e3. That is an enormous, uniform displacement across the whole parameter vector, applied at the exact moment when the weights are random and the gradient directions are least informative. ## Variance, not bias, is the problem After a few steps the ratio stops being exactly one, but the underlying issue persists in a subtler form. The second moment is an average of squared gradients over an effective window, and early in the run the window contains only as many samples as there have been steps. An average over five noisy samples is itself very noisy. There is a well-known correction that rescales these averages so their expected value is right from the first step; that correction addresses the *mean* of the estimate and leaves its *variance* untouched. High variance in the denominator is asymmetric in its consequences. If the estimate happens to come out too large, the step for that coordinate is merely too small - harmless. If it happens to come out too small, the step is too large, and because the estimate appears under a square root in a denominator, a modest underestimate produces a large step inflation. Across millions of coordinates, some fraction will draw a small denominator purely by chance in the opening steps, and those coordinates take badly oversized steps. The damage then enters the moving averages themselves and lingers for as long as those averages remember. This is exactly the motivation behind the rectified variant of Adam, which estimates how many samples the second moment has effectively accumulated and falls back to a non-adaptive, momentum-style update until the adaptive term's variance is tolerable. That method and an explicit warmup ramp are two answers to the same diagnosis; in practice most recipes just use the ramp. ## How long the accumulation takes An exponential moving average with decay `d` has an effective window of roughly `1 / (1 - d)` samples. With the common second-moment decay of 0.999 that is about a thousand steps, which sets a useful scale: a ramp of a few hundred to a couple of thousand steps is long enough for the estimate to be built from a real sample, and it is no coincidence that the ramp lengths practitioners converge on land in that range. If you deliberately shorten the second-moment window - a smaller decay coefficient - the estimate stabilizes sooner and the necessary ramp is correspondingly shorter, at the cost of a noisier normalizer for the rest of the run. ## What the interviewer is checking Three things. First, whether you know the update is scale-free by construction, so "the gradients are large at initialization" is not by itself an explanation for an adaptive method's early instability. Second, whether you distinguish bias from variance in an estimator - a candidate who says "the bias correction handles it" has repeated a fact without understanding what it corrects. Third, whether you connect the ramp length to the averaging window rather than treating it as an arbitrary magic number. ## The practical consequence Because the failure is stochastic - it depends on which coordinates happen to draw a small denominator, and on which batches arrive first - an under-warmed adaptive run is *seed-dependent*. It may train fine three times and diverge on the fourth. That is an important operating fact: a single successful run is not evidence that the ramp is long enough, and "it worked yesterday" is not a defence when the same configuration blows up on a rerun.
- Why is an underestimated second moment more dangerous than an overestimated one?The estimate sits under a square root in the denominator of the update. An overestimate only shrinks that coordinate's step, which costs a little progress. An underestimate inflates the step, and because the relationship is reciprocal, a coordinate whose few sampled gradients were unusually small can take a step many times the intended size. Damage is one-sided, so the risk is dominated by the low tail of the estimate.
- Why can an under-warmed adaptive run succeed on one seed and diverge on another?Which coordinates draw an unluckily small normalizer, and which batches arrive in the first few steps, are both random. A run whose ramp is marginally too short is sitting near a threshold, so the outcome depends on the draw. Treat a single green run as weak evidence: rerun with different seeds, or lengthen the ramp, before declaring the configuration safe.
- If you shorten the second-moment averaging window, what happens to the warmup you need?A shorter window means the estimate is built from fewer effective samples but stabilizes in fewer steps, so the ramp can be shorter in proportion. The tradeoff is that the normalizer stays noisier for the whole run, which shows up as jitter in the loss later. Most people leave the window alone and pay for it with a ramp of roughly its length.
saying these in an interview costs you the question
- Says the bias correction already fixes the early-step problem
- Claims adaptive methods never need warmup because they rescale gradients
- Argues large initial gradients are the reason, ignoring that the step is scale-free
- Treats one successful seed as proof the ramp is long enough
- Thinks the epsilon in the denominator is what keeps early steps bounded