What stopping criteria end a gradient-descent fit when you watch only the training loss?
answer
- at a minimum the gradient vanishes
- relative, not absolute, tolerances
- compare epoch averages, not batches
- a budget guarantees termination
- record which criterion fired
basics
~20 sThree, used together: the full-data gradient norm falling below a relative tolerance, the relative improvement in the epoch-average loss falling below something like 1e-6, and a maximum-epoch budget as a backstop. Record which one fired.
solid answer
~50 sNumerical convergence has three practical tests. First, the gradient norm: at a minimum the gradient is zero, so stop when `||g||` is small — but relative to its starting value, since the norm is not scale-free, and computed on the full training set, because per-mini-batch gradients stay noisy at the optimum. Second, relative loss improvement: stop when `(L_prev - L_now) / max(1, |L_prev|)` falls below a tolerance such as 1e-6 for a couple of consecutive epochs. Compare epoch averages, never a single batch's loss — on a 10-epoch fit over 2 million rows the per-batch loss is visibly noisy while the epoch trend is clean. Third, a hard epoch budget, because neither tolerance fires if the fit is diverging or crawling. Log which criterion ended the run: stopping on the budget means the fit is unfinished, not converged.
go deeper
Know that a fit runs for a set number of epochs or until the loss stops improving, and that a fixed epoch cap exists so a run always terminates.
Explain why the gradient norm is the quantity that vanishes at a minimum, and why improvement should be measured relative to the current loss rather than as an absolute change.
Show the combination in practice: a budget, a relative-improvement test on epoch averages held for consecutive epochs, and a full-data gradient check. Explain why a per-batch quantity cannot be thresholded and what a hit budget signals.
Own the convention across the team's fits — tolerances defined relatively so they transfer between problems, the stopping reason logged with every run, and a clear separation between 'the optimisation converged' and any decision about how long to keep fitting.
## The question being asked "When do I stop?" has two different answers, and it is worth separating them out loud. One is **numerical convergence**: iterating further does not move the weights, so more compute is wasted. The other is **when further fitting starts to hurt generalisation**. This is about the first: given only the training loss and the gradient, how do you know the optimisation is finished? Three criteria answer it, and a robust fit uses all three together. ## 1. Gradient norm below a tolerance At the minimum of a differentiable loss the gradient is zero, so a small gradient norm means the weights are near a stationary point. Stop when `||g|| < tol`. Two cautions make this a senior answer rather than a textbook one. First, the norm is not scale-free: it grows with the number of features and with the scale of the loss, so an absolute threshold like 1e-6 that works on one problem is meaningless on the next. Use a relative form — compare `||g||` to its value at the first iteration, or normalise by the number of parameters. Second, and specific to mini-batch fitting, the gradient of a *single mini-batch* does not go to zero at the optimum: individual batches disagree with the full data, so per-batch gradients stay noisy forever. If you want to threshold a gradient norm you must compute it on the full training set, typically once at the end of each epoch. ## 2. Relative improvement in the loss Stop when the loss stops improving meaningfully between epochs: `(L_prev - L_now) / max(1, |L_prev|) < 1e-6`. This is the criterion most people actually use, because the loss is already being computed and needs no extra pass. The mini-batch caveat matters here too. A single batch's loss bounces around from batch to batch for reasons that have nothing to do with progress — it reflects which rows landed in that batch. On a 10-epoch mini-batch fit over 2 million sensor rows, the per-batch loss is visibly noisy while the *epoch-average* trend is clean and monotone. So compare epoch averages, or a loss evaluated on a full pass, and never threshold a single batch. It is also worth requiring the criterion to hold for two or three consecutive epochs, so one lucky epoch does not end the run. A close variant is to threshold the **parameter change**: stop when `||w_new - w_old||` relative to `||w||` is tiny. This is sometimes preferable, because it directly measures the thing you care about — the weights have settled — and it is insensitive to the loss's scale. ## 3. A maximum-epoch budget A hard cap — say 100 epochs — is the backstop. It exists because the other two criteria can *never* fire: a step size too large makes the loss grow rather than shrink, a step size too small makes progress real but glacial, and a badly conditioned problem can crawl for a very long time. Without a cap, either case runs forever. The cap is not a convergence criterion; it is a guarantee of termination, and hitting it is a *signal* — a run that stopped on the budget rather than on a tolerance should be treated as unfinished and investigated, not shipped as if it had converged. ## Putting them together The usual policy is: run at most `max_epochs`; after each epoch compute the epoch-average training loss (and, if cheap, the full-data gradient norm); stop when either the relative loss improvement or the gradient norm falls below its tolerance for a couple of consecutive epochs; log which criterion fired. That last detail is the one people skip and then regret — six weeks later, "did that model converge or did it just run out of epochs?" is unanswerable unless the run recorded the answer. ## Things that go wrong - **Thresholding absolute loss.** "Stop when loss < 0.1" is meaningless: the achievable loss depends on the data's irreducible noise, not on the optimiser. - **Thresholding an absolute improvement on an unscaled loss.** An improvement of 1e-6 is enormous if the loss is 1e-5 and negligible if the loss is 1e6. Use the relative form. - **Stopping on the first non-improving epoch.** With mini-batch noise, one flat epoch is normal; requiring a small number of consecutive flat epochs is more stable. - **Treating a hit budget as convergence.** It means the opposite. Note finally that all of this concerns the *training* loss and the health of the optimisation. Deciding to halt a fit that is still improving on training data because it has stopped improving elsewhere is a different decision with a different justification, and should not be confused with the convergence test.
- Why is a single mini-batch's loss a poor quantity to threshold?It reflects which rows landed in that batch as much as it reflects the state of the weights, so it fluctuates from batch to batch even when the fit is improving steadily. Thresholding it stops the run on noise. Compare the epoch-average loss, or a loss evaluated on a full pass, and require the criterion to hold for two or three consecutive epochs.
- Why should the gradient-norm test use the full-data gradient rather than a batch gradient?Individual batches disagree with the full data, so their gradients do not go to zero at the optimum — a per-batch norm stays noisy forever and never crosses a small threshold reliably. Computing the gradient over the whole training set once at the end of an epoch gives the quantity that actually vanishes at a minimum.
- What does it mean when a fit stops because it hit the maximum-epoch budget?That it did not converge. The budget is a termination guarantee, not a convergence criterion — it exists so a diverging or extremely slow fit ends at all. Hitting it is a signal to investigate: check whether the loss was still improving, whether the step size is too small, or whether the run was diverging all along.
saying these in an interview costs you the question
- Thresholds an absolute loss value like 0.1
- Stops on the first epoch that fails to improve
- Thresholds the loss of a single mini-batch
- Treats hitting the epoch budget as convergence
- Uses an absolute gradient tolerance across different problems