When is square-root learning-rate scaling preferred over linear scaling as the batch grows?
answer
- ask what each rule holds constant
- noise per step versus noise per epoch
- adaptive updates already divide out magnitude
- root of the factor, not the factor
basics
~20 sSquare-root scaling, multiplying the learning rate by the square root of the batch-size factor, is the usual starting point with adaptive per-coordinate optimizers, where full linear scaling overshoots. Linear scaling stays the default for momentum SGD.
solid answer
~40 sThe two rules preserve different quantities. Take as given that the noise in a mini-batch gradient shrinks like `1/sqrt(B)`. Then the noise injected into the weights per *update* scales like `lr / sqrt(B)`, and holding that fixed gives `lr` proportional to `sqrt(B)`. Linear scaling instead holds the noise per unit of *data* fixed: in the continuous-time picture of SGD that level goes like `lr / B`, so `lr` proportional to `B` keeps it constant. In practice linear scaling is the convention with momentum SGD on convolutional models, while square-root scaling is the common recommendation with adaptive optimizers, whose updates already divide the gradient by a running root-mean-square of its own magnitude and so absorb part of the change. Both are heuristics: use one to set the search range, then sweep around it.
go deeper
Know that two recipes exist — multiply the rate by the batch factor, or by its square root — and that which one you use depends on the optimizer rather than on taste.
Be able to say what each rule keeps constant: per-step weight noise for the square-root rule, noise per unit of data for the linear rule. Explain why a gradient-normalizing update pushes you toward the square root.
Demonstrate that you treat the scaled rate as the centre of a short sweep and verify transfer by comparing loss against examples processed against the smaller-batch baseline.
Own the position that both rules are heuristics with a range of validity, and set the team's convention accordingly rather than letting each project rediscover it in a week of wasted sweeps.
### Two rules, two invariants When the batch grows by a factor `k`, the two candidate recipes are: - **Linear:** `lr -> k * lr` - **Square-root:** `lr -> sqrt(k) * lr` Neither is a theorem. Each falls out of asking *what should stay constant* as the batch grows, and they answer that question differently. Take as given the basic fact that averaging over more examples makes the mini-batch gradient a tighter estimate of the full-batch gradient, with the fluctuation around it shrinking like `1/sqrt(B)`. **Square-root scaling holds the per-update noise constant.** One update moves the weights by roughly `lr * (true gradient) + lr * (noise of size proportional to 1/sqrt(B))`. The random part of a single step therefore has magnitude proportional to `lr / sqrt(B)`. If you want each individual update to be as noisy as it was before — no more jittery, no more deterministic — you set `lr` proportional to `sqrt(B)`. **Linear scaling holds the noise per unit of data constant.** Modelling SGD as a continuous-time process, the amount of random exploration it performs is governed by a ratio that behaves like `lr / B`, sometimes called the temperature of the process. Fixing that ratio gives `lr` proportional to `B`. This matches the step-matching derivation of the linear rule: over a fixed number of examples the large-batch run takes `k` times fewer steps, and each is `k` times bigger, so both the signal and the accumulated randomness per epoch stay put. So linear scaling preserves the character of a *pass over the data*; square-root scaling preserves the character of a *step*. They cannot both be right, and which is closer to right depends on the update rule. ### Why the optimizer decides it Plain momentum SGD applies the gradient as-is: if the batch changes the gradient's typical magnitude or its noise level, the update inherits that change directly, and the linear rule's step-matching argument applies cleanly. This is why linear scaling is the convention in the vision setting where it was popularized — a convolutional classifier trained with momentum SGD on a fixed epoch budget. Adaptive per-coordinate optimizers behave differently. An Adam-family update is roughly `g / (sqrt(v) + eps)` where `v` is a running average of the squared gradient per coordinate. That division is a normalization: it already removes much of the change in the raw magnitude of the gradient, including part of the change caused by averaging over more examples. Multiplying the rate by the full factor `k` on top of an update that has already been rescaled therefore tends to overshoot, and `sqrt(k)` is the better-behaved default. Practitioners training large language models with adaptive optimizers and decoupled weight decay commonly report square-root-like scaling as the setting that transfers. ### How to use this in practice The useful posture is: *the rule sets the search range, the sweep sets the rate.* 1. Pick the rule that matches your update rule — full factor for momentum SGD, square root for adaptive methods. 2. Compute the scaled rate and treat it as the centre of a small grid, typically a factor of two either side. 3. Bring the rate up gradually at the start of the run rather than applying the scaled value at step zero; this is more important the bigger the jump. 4. Compare loss against **examples processed** with the baseline. If the scaled run tracks the baseline, the rule transferred; if it sits above it everywhere, either the rule is the wrong one for your optimizer or you have pushed the batch past the range where any rate scaling recovers the baseline. ### The honest interview answer If someone asks 'linear or square-root?', the answer that reads as experienced is not a single word. It is: linear is the default with momentum SGD, square-root is the default with adaptive methods, both are empirical heuristics that hold over a range and then stop, and in either case the scaled value is a starting point you verify rather than a number you trust. A candidate who insists one rule is universally correct is quoting a paper rather than describing a run they have done.
- What quantity does square-root scaling keep constant that linear scaling does not?The size of the random component of a single update. Per-step weight noise scales like `lr / sqrt(B)`, so holding `lr` proportional to `sqrt(B)` leaves each individual step as noisy as before. Linear scaling instead holds the noise accumulated per unit of data constant — the `lr / B` ratio — which is the right invariant if you care about the character of a full pass over the data rather than of one step.
- Why does an adaptive per-coordinate optimizer change the answer?Its update is roughly the gradient divided by a running root-mean-square of that same gradient per coordinate. That division normalizes away much of the change in raw gradient magnitude that a bigger batch produces, so applying the full linear factor on top of it double-counts and the resulting step is too large. Square-root scaling is the empirical middle ground that transfers better in that setting.
saying these in an interview costs you the question
- Claims one rule is universally correct
- Applies full linear scaling to an adaptive optimizer by reflex
- Says square-root scaling is derived, not empirical
- Skips the sweep around the scaled value