skip to content

Your booster's held-out fold stops training at round 740 of 3,000 — how do you report its accuracy and ship that round count?

level: seniorimportance: should knowfreq 57%

answer

  1. the fold made a choice, not just a measurement
  2. 740 is one draw, not a constant
  3. tied to the rate that produced it
  4. more training data supports more rounds
  5. stop on the metric you ship on

basics

~20 s

Treat them separately. The stopping fold chose a hyperparameter, so its score is optimistic and the reported accuracy must come from data that played no part in stopping. Ship 740 as a fixed round count, valid only for the learning rate and subsampling it was found under.

solid answer

~50 s

Two separate things: the score and the round count. The fold that halted training selected a hyperparameter — the number of rounds — so its score is optimistically biased and must not be quoted as the model's accuracy; report on a test split that played no part in stopping, or nest the stopping fold inside an outer split. For the round count, 740 is one draw from a distribution: repeat the stopping across several folds and look at the spread before committing, since a single fold can be off by hundreds of rounds. Remember that 740 is bound to the learning rate and subsampling it was found under — change the rate and the number is meaningless. If you refit on training plus validation data, the extra data generally supports somewhat more rounds, so scale up modestly or simply ship the early-stopped model as it stands. And stop on the metric you will actually be judged on.

go deeper

for a junior

Know that boosting can pick its own number of rounds by watching a held-out fold, and that the chosen number belongs with the settings it was found under.

for a middle

Explain why the fold that chose the round count cannot also measure performance, and that the round count moves with the learning rate.

for a senior

Demonstrate the full protocol: a clean test split or a nested arrangement, the round count taken from several folds rather than one, and a deliberate decision about refitting on the stopping fold's data.

for a principal

Own the reporting standard. Decide what a model result must include before it is quoted to the business, so nobody ships a number selected on the same data that produced it.

## What early stopping in boosting actually is Boosting has one hyperparameter you never have to search over: the number of rounds. Because the ensemble is built incrementally and every prefix of it is a valid model, you can request a large budget — 3,000 rounds — score a held-out fold after each round, and simply keep the prefix that scored best. Here that prefix is 740 trees. The remaining 2,260 rounds were trained and thrown away, or training was cut short once the fold stopped improving. That convenience hides two decisions people routinely get wrong: what the fold's score means, and whether 740 is a number you can ship. ## The score is contaminated The held-out fold was used to **choose** the round count. That makes the round count a hyperparameter fitted on that fold, and the fold's score a fitted-on quantity, not a clean estimate. Concretely: out of 3,000 candidate models you selected the one that looked best on this fold, so you also selected whatever favourable noise that fold happened to contain. The bias is not enormous when the curve has a broad, flat minimum, and it can be substantial when the curve is jagged and the fold is small. The fix is structural, not cosmetic. Hold out a test split that is touched by nothing in training — not the stopping, not the rate choice, not the feature decisions — and quote that. If data is too scarce for a three-way split, put the stopping fold *inside* each outer training fold, so every outer fold's score comes from data its own early stopping never saw. What you must not do is quote the stopping fold's own best score as the model's performance; it is the single most common way a boosting result gets inflated on the way into a slide. ## Is 740 a stable number? Rarely, on its own. Run the same procedure on a different fold and you may get 600 or 900. The variability comes from three places: the fold is finite, so the curve it draws is noisy; subsampling makes the increments themselves noisy; and near the optimum the curve is usually flat, so the argmin moves a lot for a tiny change in the score. The practical response is to stop several times — once per cross-validation fold — and look at the distribution of best rounds rather than one value. Take the mean, or pick a conservative point on the flat part of the averaged curve. If the spread is huge relative to the value, that is itself information: the model is not sharply defined, and the difference between 600 and 900 rounds probably does not matter. ## The round count is not portable 740 rounds is meaningless without the settings that produced it. It is bound to: - **The learning rate.** Drop the rate from 0.03 to 0.015 and you need roughly twice as many rounds; 740 no longer refers to anything. - **The subsampling fractions and tree size.** Change how much each tree can absorb per round and the pace changes with it. - **The training-set size.** More data supports more rounds before the marginal tree starts fitting noise. So record the round count together with the configuration that produced it, and re-establish it whenever any of those move. Carrying a round count across a rate change is a real and frequently-made mistake. ## Refitting on train plus validation Once the round count is chosen, it is tempting to refit on all the data — training plus the stopping fold — to get the extra rows into the model. That is defensible, but 740 was calibrated on a smaller training set. With more data the same number of rounds is somewhat conservative, so a modest scale-up in proportion to the data increase is the usual adjustment. The alternative, entirely respectable, is to ship the early-stopped model exactly as it was fitted: you give up a slice of data and gain a model whose round count was actually validated on the data it was trained on. Choose deliberately; do not refit on everything and keep 740 without noticing you changed the problem. ## Stop on the metric you are judged on A ranking metric and a probabilistic loss do not peak at the same round. A booster's ranking quality can keep creeping up for hundreds of rounds after its log loss has bottomed out and started to rise, because later rounds sharpen the ordering while pushing probabilities toward the extremes. If the model outputs probabilities that feed a threshold, a bid or an expected-value calculation, stop on the probabilistic loss. If only the ordering is consumed, stop on the ranking metric. Stopping on one and shipping for the other is how a well-ranked, badly calibrated model reaches production. ## What to hand over The deliverable is a pair: the fixed round count and the full configuration it belongs to, plus a performance number from data that had no role in either. Anything less and the next person to retrain the model has no way of knowing whether 740 still means anything.

  • How much does the optimistic bias from stopping on that fold actually matter?
    It depends on the curve. With a broad flat minimum and a large fold, selecting the argmin out of 3,000 candidates buys you little extra noise and the bias is small. With a jagged curve and a small fold it can be material, because you are picking the luckiest wiggle. Either way I never quote it as the headline number — a clean test split costs nothing to keep.
  • Your cross-validation folds return best rounds of 610, 740, 905 and 660 — what do you ship?
    I look at the averaged validation curve rather than the four argmins. If it is flat between roughly 600 and 900, any value in there is fine and I take the mean, around 730, or a slightly conservative point. The spread itself tells me the round count is not a sharp choice, so I stop agonising over it and spend the effort elsewhere.
  • Can you reuse 740 rounds after halving the learning rate?
    No. Halving the rate roughly doubles the rounds needed, so 740 would leave the model badly underfitted. The round count is only meaningful alongside the rate, subsampling and tree size that produced it; change any of them and it has to be re-established.

saying these in an interview costs you the question

  • Reports the stopping fold's best score as the model's accuracy
  • Treats a single fold's best round as an exact constant
  • Carries the round count across a learning-rate change
  • Refits on all data and keeps the same round count unexamined
  • Stops on a ranking metric but ships calibrated probabilities

context