Why must you re-run a learning-rate range test after the batch size or initialization changes?
answer
- the curve belongs to the configuration
- gradient noise shrinks with batch size
- pretrained weights are already good
- the knee moves both directions
- two hundred steps to re-measure
basics
~20 sThe curve describes a configuration, not an architecture. A larger batch averages away gradient noise, so the divergence knee moves up; a pretrained encoder sits near a good solution, so rates that were fine from scratch wreck its features.
solid answer
~50 sA range test measures where one specific setup becomes unstable, and both changes move that point. Grow the per-device batch from 32 to 512 and each step's gradient averages sixteen times more examples, so its variance falls roughly proportionally; the stochastic overshoot that used to blow the run up no longer does, and the knee in the loss-versus-rate curve shifts toward larger rates. Keeping the old rate is then merely wasteful: fewer steps per epoch, each smaller than the setup tolerates. Initialization moves the knee the other way. A randomly initialized encoder has large, uninformative gradients and needs a big rate; a pretrained one already sits in a good region, so the same rate destroys learned features and the loss turns upward one or two decades earlier. Re-running the sweep costs a couple of hundred steps, so measure it rather than assume.
go deeper
Remember that the number a range test produces is tied to the setup it was measured on. If someone hands you a learning rate, ask which batch size and which starting weights it came from before reusing it.
Explain the mechanism in both directions: averaging more examples per step lowers gradient variance so larger batches stay stable at larger rates, while pretrained weights start near a good solution and are wrecked by rates that suited random initialization.
Show the operational habit of re-measuring after any recipe change and confirming on the first hundred steps of the real run. Be able to say which way the knee moved and why, not merely that you re-ran the sweep.
Decide what the team re-measures and what it inherits. Pin every recorded rate to a configuration fingerprint — batch size, initialization source, loss composition — so an inherited number is either valid by construction or visibly stale.
## What the curve is a measurement of A loss-versus-rate curve is not a property of the architecture. It is a property of the exact pairing of model, data, per-device batch size, starting weights, loss composition and optimizer that produced it. Change any of those and the sweep you ran earlier measured a different system. Two changes move the curve in opposite directions and come up constantly in real work. ## Batch size moves the knee up Each step's gradient is an average over the examples in the batch. For B roughly independent examples, the variance of that average falls about as `1/B`, so the noise standard deviation falls about as `1/sqrt(B)`. Going from 32 to 512 is a sixteen-fold variance reduction — roughly a factor of four in noise scale. Two consequences follow. The update direction is closer to the full-data gradient, so a given rate produces a step that is more reliably downhill. And the tail events — the occasional batch whose gradient is unusually large and points somewhere unhelpful — become rarer and smaller. Those tail events are precisely what turn a marginal rate into a divergent one, so the sharp rise on the curve moves toward larger rates. The keyword-spotting sweep that turned upward at some rate with a batch of 32 keeps improving past it at 512. This has a ceiling. Once the batch is large enough that its gradient is already close to the full-data gradient, adding more examples buys almost no further stability and the knee stops moving. There is also an accounting change worth stating out loud: at a fixed epoch budget, a batch of 512 takes one sixteenth as many steps as a batch of 32, so carrying over the old, now-too-conservative rate gives you both fewer steps and smaller ones. ## Initialization moves the knee down A randomly initialized encoder holds no information. Its loss starts near the chance level, its gradients are large, and the job of early training is to move a long way through weight space — large rates are both tolerated and needed. A pretrained encoder is the opposite. Its weights already encode useful features, the loss starts far lower, and the run's job is a comparatively small adjustment. On the curve, two things change: the whole trace sits lower, and the rise begins much earlier, because a large update does not improve the features, it erases them and drags the loss back toward the from-scratch level. It is common for the usable rate on a pretrained trunk to be one to two orders of magnitude smaller than on the same architecture trained from scratch. Run both sweeps on the same model and you get two curves of similar shape displaced along the rate axis; that displacement is the entire point of re-running the test. ## The mixed case A pretrained trunk with a freshly initialized head is the case that catches people. The head needs a large rate to move at all; the trunk is damaged by it. One curve reports one compromise number, usually dictated by whichever part destabilizes first, and it hides the conflict entirely. If you intend to use separate rates per parameter group, sweep per group — or sweep once with the trunk frozen and once with it unfrozen — and read two numbers instead of one. ## What else invalidates a previous curve Anything that changes the scale or noisiness of the gradients: a changed loss weighting or a new auxiliary term, adding or removing normalization layers, a different gradient-clipping threshold, a longer input sequence, a different data mix, a change of width or depth, or a change of optimizer. The safe rule is that the curve belongs to the exact recipe that produced it. ## The practical habit Re-measure. Two hundred steps is negligible next to the run it protects, and it turns an assumption into an observation. Record the rate together with the fingerprint that makes it valid — batch size per device, initialization source, loss composition — so that anyone inheriting the number can tell at a glance whether their setup is the one it was measured on. Then confirm on the first hundred steps of the real run, where a rate that is too hot or too cold shows up immediately.
- Your encoder is pretrained but the classification head is new. What does a single range-test curve hide?That the two parts want different rates. The fresh head needs a large rate to move at all, while the pretrained trunk is damaged by it. One curve reports a single compromise, usually set by whichever part destabilizes first. If you plan separate rates per parameter group, sweep per group — or sweep once with the trunk frozen and once unfrozen — and read two numbers.
- Besides batch size and initialization, what else invalidates a previous curve?Anything that changes the size or noisiness of the gradients: a changed loss weighting or an added auxiliary term, normalization layers added or removed, a different gradient-clipping threshold, a longer input sequence, a different data mix, altered width or depth, or a change of optimizer. The safe rule is that the curve belongs to the exact recipe that produced it.
- Does the divergence knee always move upward when the batch grows?Directionally yes, because averaging more samples reduces gradient variance, but not without limit. Once the batch gradient is already close to the full-data gradient, extra examples buy almost no additional stability and the knee stops moving. That is why the honest answer is to re-measure with a short sweep rather than to assume the knee tracks batch size indefinitely.
The same engine redlines at different revs depending on the fuel and the load. The tachometer reading you wrote down last time describes that day's setup, not the engine.
saying these in an interview costs you the question
- Treats the found rate as a property of the architecture
- Reuses the from-scratch rate when fine-tuning pretrained weights
- Assumes larger batches require smaller learning rates
- Says the knee cannot move because the model is unchanged
- Reads one curve for a frozen trunk and a fresh head