skip to content

How do you set the patience for early stopping when the validation curve is jagged?

level: seniorimportance: should knowfreq 48%

answer

  1. the curve is a sample statistic
  2. measure the noise before choosing
  3. the two errors are not symmetric
  4. counted in checks, not passes
  5. best iterate is restored either way

basics

~20 s

Set early-stopping patience from the noise scale of the monitored curve: it must outlast the runs of worsening checks that random variation alone produces. Patience of one halts still-improving fits; over-long patience costs only compute.

solid answer

~40 s

I read the curve before picking a number. Fit once past the point I would normally stop, plot the monitored metric per check, and see how long the false dips last — that length plus a margin is the patience. Consider a validation log-loss of 0.451, 0.455, 0.458, 0.452, 0.440: three worsening checks, then a genuine new low. Patience one would have halted before the improvement ever appeared. The two errors are asymmetric: too small and you ship a worse model and never notice, too large and you only burn compute, since the best iterate is snapshotted and restored. So err large. Two details bite people: patience counts validation *checks*, so with evaluation every fifth pass patience 3 means fifteen passes; and without a minimum improvement threshold, microscopic gains keep resetting the counter forever.

code

python · 20 lines
python
val_loss = [0.62, 0.55, 0.49, 0.46, 0.451, 0.455, 0.458, 0.452, 0.44, 0.437]

def stop_at(losses, patience):
    best, best_i, waited = float("inf"), -1, 0
    for i, v in enumerate(losses):
        if v < best:
            best, best_i, waited = v, i, 0
        else:
            waited += 1
            if waited >= patience:
                return i, best_i, best
    return len(losses) - 1, best_i, best

for p in (1, 4):
    halted, best_i, best = stop_at(val_loss, p)
    print("patience", p, "halted at", halted,
          "best iterate", best_i, "best loss", round(best, 3))

# patience 1 halted at 5 best iterate 4 best loss 0.451
# patience 4 halted at 9 best iterate 9 best loss 0.437

go deeper

for a junior

Know what the word means: patience is how many consecutive validation checks may fail to improve before the run halts, and it exists because the curve wobbles rather than descending smoothly.

for a middle

Explain the counting rule precisely — patience is in checks, the counter resets on any new best, and a minimum improvement threshold stops trivial gains from resetting it forever.

for a senior

Show the diagnostic habit: measure the curve's noise from a long run, size patience above the longest false dip, and argue the asymmetry that too-large patience costs only compute.

for a principal

Frame it as a budget call across many retrains. Decide the monitoring cadence, whether the team stops on a proxy loss or a business metric, and who owns the untouched split that keeps reported numbers honest.

### What patience actually buys Early stopping halts after some number of consecutive validation checks without improvement. That number is `patience`, and it exists for one reason: the validation curve is a *sample statistic*, not a smooth function. It wobbles because the validation slice is finite, because the optimizer path is stochastic, and because the metric itself may be coarse. Patience of one halts on the first wobble; a well-chosen patience halts only when the wobble has lasted long enough to be evidence. ### Reading the curve first Consider a validation log-loss that reads 0.462, 0.451, 0.455, 0.458, 0.452, 0.440, 0.437 over consecutive checks. Three consecutive checks are worse than the 0.451 at check two, and then the fit finds a genuinely better region. Patience one would have halted at 0.455 and restored the 0.451 iterate — leaving real accuracy on the table for no reason other than impatience. Patience four survives the plateau and lands on 0.437. So the calibration question is empirical: **how long a run of worsening checks does noise alone produce on this curve?** You can read that off a single long diagnostic run — fit past the point you would normally stop, plot the monitored metric per check, and look at the length of the false dips. Set patience above the longest of those runs. On a smooth, well-sampled curve that might be 3; on a small validation slice or a coarse business metric it might be 20. ### The asymmetry that should drive the decision The two errors are not symmetric. - **Patience too small**: you stop before the fit is done and ship a genuinely worse model. The loss is in accuracy, and you never see it — the run looks like it converged. - **Patience too large**: you spend extra compute after the best iterate has already been passed. Because the best iterate is snapshotted and restored, the shipped model is *unaffected*. The costs are compute versus quality, and compute is usually cheaper. When in doubt, err large. This also means the common ritual of copying patience 10 from the last project is defensible as a default but indefensible as an answer — it is only right if the curve's noise scale happens to match. A useful companion knob is a minimum improvement threshold: require a new best to beat the old one by some small margin before it resets the counter. Without it, a curve that improves by 0.00001 per check will never trigger a halt, and patience does nothing. ### Patience is counted in checks, not passes If you evaluate every fifth pass, patience 3 means fifteen passes of no improvement, not three. Mixing the two units is a real and common bug: someone moves the evaluation to every fifth pass to save time and unknowingly makes the stopping rule five times more lenient. State the cadence and the patience together, and reason about the product. Cadence itself is a tradeoff. Frequent checks give a fine-grained curve and a tight stop, but each check costs a pass over the validation slice. Infrequent checks are cheap but coarse — you may overshoot the true minimum by several passes, and each check is a noisier draw of the metric, which argues for a *larger* patience, not a smaller one. ### What you monitor changes everything Monitoring validation log-loss every pass gives a smooth, sensitive signal, and it is the right default for a probability model. Monitoring a business metric — captured revenue at a fixed impression budget, say — is more faithful to the decision the model serves but is typically a step function of the threshold, evaluated on fewer effective samples, and far noisier per check. If you stop on it, raise patience and check less often; if the metric is so noisy that its check-to-check variation swamps the real improvement, stopping on it is close to random and you should stop on the smooth loss while *reporting* the business metric. ### The slice you carved out A validation slice used purely to watch the run is rows removed from training. On a few million ad impressions that is a rounding error. On 2,000 rows it is a real cost, and the stopping point chosen on 400 held-out rows is itself high variance. The fix is to spend folds instead of a single split: run the fit on each fold, record the best iterate per fold, and refit on all rows for the averaged count. You pay in compute and get back both the rows and a more stable stopping point. Finally, the score at the halt is the minimum over every check you made — a selected value, and therefore optimistic. Never quote it as the expected production number; quote a split that had no part in the stopping decision.

  • You want to stop on a business metric that can only be computed every fifth pass. What changes?
    Two things. Patience is counted in checks, so the same number now buys five times as many passes — restate it in passes before comparing. And a coarse business metric is usually far noisier per check than a log-loss, so it needs a larger patience, not a smaller one. If its check-to-check wobble swamps the real improvement, stop on the smooth loss and report the business metric.
  • Is the validation score at the halt a fair estimate of production performance?
    No. It is the minimum over every check you made, so it is a selected value and biased optimistic — more checks make it more optimistic. The validation slice has been spent as a tuning set. Report a third split that took no part in the stopping decision, and expect it to read a little worse.
  • You only have 2,000 rows. Is carving out a validation slice purely to watch the run still worth it?
    Often not as a single split: those rows are lost to training and a stopping point chosen on 400 rows is itself high variance. Better to spend folds — fit on each fold, record the best iterate per fold, average the counts, then refit on all rows for that many iterations. You pay compute and get back both the rows and a stabler stop.
  • The validation curve never turns upward before your iteration budget runs out. What does that tell you?
    That the fit is not overfitting at this capacity, so early stopping has nothing to regularize. Stopping anyway just gives away accuracy. Either the model is underparameterised for the data, the learning rate is small enough that the run has not converged, or you need more iterations rather than a stopping rule.

saying these in an interview costs you the question

  • Copies patience ten without looking at any curve
  • Halts at the first check where validation worsens
  • Counts patience in passes while checking every fifth pass
  • Quotes the best validation score as expected production performance
  • Stops on a metric so noisy the halt is effectively random
  • Thinks a large patience risks shipping an overfitted model

context