Why is a neural network's learning rate usually decayed over the course of training?
answer
- the step scales signal and noise alike
- near a minimum, only noise is left
- constant rate implies a loss floor
- high rate early explores, low rate settles
- sum of steps infinite, squares finite
basics
~20 sA large step keeps stochastic gradient descent bouncing around a minimum instead of settling in it. Shrinking the step later in training lets the noisy gradient estimates average out, so the parameters settle and the loss stops oscillating.
solid answer
~50 sEach update is `w <- w - lr * g`, and `g` is a mini-batch estimate, so it carries noise that the learning rate scales along with the signal. Near a minimum the true gradient is nearly zero but the noise is not, so a constant rate leaves the parameters wandering in a region whose size grows with the step. That region is a loss floor: the run stops improving because the step is too big to settle, not because the model cannot fit better. Decaying shrinks the region and lowers the floor, which is why the training curve drops right after each cut. The early high-rate phase still earns its keep — it makes fast progress and behaves like a regularizer — so the goal is to spend time at a high rate first and then cool down, not to start small.
go deeper
Be ready to say in one breath that a big step keeps the parameters bouncing near a minimum and a smaller step lets them settle, and to recognise the drop in a training curve right after a rate cut.
Explain the mechanism: the step multiplies both the true gradient and the mini-batch noise, so a constant rate leaves a wandering region whose size grows with the rate. Name both failure modes, never decaying and decaying too soon.
Show you use the rate cut as a diagnostic on real runs — cut and watch, to separate a step-size plateau from a capacity limit — and that you treat the long high-rate phase as regularisation you are deliberately buying.
Own the framing that the schedule, not the optimizer choice, is usually the dominant knob in a training recipe, and that schedules must be specified against a budget so results across teams remain comparable.
## What the learning rate actually multiplies A gradient step is `w <- w - lr * g`. The quantity `g` is not the true gradient of the training loss; it is an estimate computed on one mini-batch, so it equals the full-dataset gradient plus a random error. The learning rate multiplies the whole thing, which means it scales the useful signal and the noise by exactly the same factor. Every argument about annealing follows from that one fact. ## The noise floor Far from a minimum, the signal dominates: the true gradient is large, the noise is a perturbation, and a big step buys fast progress. Close to a minimum the balance inverts. The true gradient goes to zero, the mini-batch noise does not, and each update injects a random displacement proportional to the step size while the curvature of the loss pulls the parameters back toward the bottom. The two forces reach a balance, and the parameters end up wandering inside a region around the minimum rather than sitting at it. The size of that region grows with the learning rate. Treating stochastic gradient descent as a discretisation of a noisy continuous process gives the same picture: for a fixed noise level, the excess loss the run hovers above the true minimum grows roughly in proportion to the step size. So a constant learning rate implies a loss floor. When a run plateaus with a constant rate, it usually has not run out of capacity or data; it has run out of resolution. Cutting the rate shrinks the region and lowers the floor. This is why a step-decayed run shows a visible cliff in the training loss immediately after each cut: the same weights, evaluated after a few steps at a tenth of the step size, sit closer to the bottom of the same basin. ## Two jobs, early and late The schedule is doing two different jobs at two different times. - **Early — travel and explore.** A large rate crosses a badly conditioned landscape quickly, and the noise it amplifies acts as a search mechanism, letting the run leave narrow regions it happened to land in. Runs that spend a long stretch at a high rate typically generalise better than runs that are decayed immediately, so the high-rate phase functions as a regularizer, not just as a hurry. - **Late — refine.** Once the trajectory is in the region it is going to stay in, the remaining job is to sit down in it. That needs small steps and nothing else. A schedule is simply a plan for switching between those two jobs over the budget you have. ## The classical condition Stochastic approximation theory states the tension precisely. For convergence you want the step sizes to satisfy two conditions at once: they must sum to infinity, so the run can still travel an arbitrary distance no matter where it starts, and their squares must sum to a finite value, so the accumulated noise is bounded. A constant rate satisfies the first and fails the second — hence the noise floor. A rate that collapses too fast satisfies the second and fails the first — the run freezes wherever it happens to be, and the final loss is worse than a slower decay would have reached. Deep networks are non-convex, so this is a guide rather than a guarantee, but it names the two failure modes correctly. ## The two failure modes in practice - **Never decaying.** Training and validation loss plateau above what the model can reach, and metrics jitter from epoch to epoch. Cutting the rate at that point produces an immediate improvement, which is the diagnostic. - **Decaying too early or too aggressively.** The run looks excellent for a few epochs, then flatlines well above where it should. The parameters were locked into place before they had finished travelling. This one is more expensive because it is invisible without a comparison run. ## A mental model Treat the learning rate as a temperature. High temperature explores and refuses to commit; low temperature crystallises whatever configuration it is handed. Annealing is a planned cooling path, and the reason it is planned rather than greedy is that the useful exploration happens before there is any evidence it was useful. ## Two things it drags along with it First, the noise in `g` depends on how many examples each estimate averages, so the same schedule can behave differently when that changes. Second, in optimizers that apply weight decay as a separate term scaled by the same schedule multiplier, annealing the rate also anneals the effective strength of that decay — the regularisation weakens exactly as the rate does. Neither is a reason to skip annealing; both are reasons to describe a schedule as part of the recipe rather than as a detail.
- How would you tell from a training curve that a run is sitting on a noise floor rather than out of capacity?Cut the learning rate by a factor of ten and watch the next few hundred steps. If the loss drops sharply and then flattens at a lower level, the previous plateau was a step-size effect, not a capacity limit. If almost nothing happens, the model really has fitted what it can and the fix is elsewhere — more capacity, better features, or cleaner data.
- Is starting at a small learning rate and never decaying a safe compromise?No. It gives up the exploratory high-rate phase that makes runs generalise well, converges far more slowly, and still leaves a noise floor — a smaller one, but a floor. You end up paying for the worst of both: slow travel and no clean settling. The cheap version of a schedule is a high rate held for most of the budget and then annealed, not a low rate everywhere.
- Does the argument change if the gradients were exact rather than mini-batch estimates?Largely, yes. With exact gradients there is no noise floor, and a fixed step below the stability limit converges on its own. You would still reduce the step near the end for badly conditioned curvature or to satisfy a line-search condition, but the dominant reason to anneal in deep learning — that the step size sets how much sampling noise survives into the parameters — disappears.
Annealing metal works the same way: hold it hot so the atoms can rearrange, then cool it on a planned schedule so it settles into a good structure. Quench it immediately and you freeze in whatever mess was there.
saying these in an interview costs you the question
- Says decay exists only to make training faster
- Claims a plateau always means the model lacks capacity
- Starts at a tiny rate to be safe and never decays
- Thinks decay changes the loss surface rather than the step
- Decays from step one, skipping the exploratory phase