skip to content

Shrinkage and Early Stopping

A small learning rate needs many more trees but generalises better, row and column subsampling adds noise on purpose, and a validation fold says when to stop. Interviewers probe that trade.

on this pageshow

questions

4

In gradient boosting, what does the learning rate do, and how does it trade against the number of rounds?

level: juniorimportance: must knowfreq 74%

answer

  1. it scales something before it is added
  2. each round is a partial correction
  3. rate and rounds move together
  4. rate times rounds is roughly constant
  5. small rate, tiny budget means underfitting

basics

~20 s

The learning rate scales every tree's output before it is added to the ensemble. Shrinking it makes each round a smaller correction, so you need proportionally more rounds; roughly, learning rate times rounds sets how far the fit travels.

solid answer

~50 s

The learning rate, also called shrinkage, multiplies each newly fitted tree's output before it is added to the ensemble: `F_m = F_(m-1) + eta * h_m`. It is not a step size used inside the tree — the tree is fitted to the pseudo-residuals normally and then damped. Because each round closes only a fraction `eta` of the remaining error, total progress depends mostly on the product of rate and rounds: halving the rate roughly doubles the rounds needed for the same fit. Small rates are still preferred, because the trade is not exact — many small corrections average into a smoother, lower-variance model, and a finely spaced validation curve lets a held-out fold pin the stopping point precisely. The cost is linear: rate 0.3 with 200 rounds to rate 0.03 with 2,000 rounds is about ten times the training time and ten times the trees to score.

code

python · 10 lines
python
def boost(eta, rounds, y=1.0):
    pred = 0.0
    for _ in range(rounds):
        residual = y - pred      # negative gradient of squared error
        pred += eta * residual   # the shrunken update
    return pred

for eta, rounds in [(0.30, 10), (0.15, 20), (0.03, 100), (0.03, 10)]:
    print(eta, rounds, 'product', round(eta * rounds, 2),
          '-> fit', round(boost(eta, rounds), 4))

go deeper

for a junior

Be ready to state in one line that the learning rate damps each tree's contribution and that a smaller rate needs more rounds. Knowing the direction of the trade is the whole screening bar here.

for a middle

Explain the update rule and why the product of rate and rounds is what mostly determines the fit, including why the trade is only approximate rather than exact.

for a senior

Show that you treat the pair as one setting: quote the wall-clock and model-size cost of a low rate, and describe how you establish the round count for whichever rate you pick.

for a principal

Own the economics. Decide when a ten-fold training-cost increase is justified by the validation gain on a frequently retrained model, and set a default the whole team can train within its schedule.

## What the learning rate multiplies Gradient boosting builds its model additively. It starts from a constant prediction `F_0` — the mean of the target under squared error, the log-odds of the base rate under log loss. At each round `m` it computes pseudo-residuals (the negative gradient of the loss with respect to the current prediction, one value per training row), fits a shallow tree `h_m` to those residuals, and updates `F_m(x) = F_(m-1)(x) + eta * h_m(x)` The number `eta`, between 0 and 1, is the learning rate, or **shrinkage**. Note where it sits: the tree is grown by ordinary split search on the residuals, and only its finished output is multiplied by `eta` on the way into the ensemble. With `eta = 1` the model takes the full correction each tree proposes. With `eta = 0.03` it keeps 3% of it and deliberately leaves 97% of the error on the table for later rounds. ## The reciprocal trade Take the simplest possible case: one target value `y`, and a weak learner that can predict the current residual exactly. After `M` rounds the prediction is `pred_M = y * (1 - (1 - eta)^M)` Because `(1 - eta)^M` is approximately `exp(-eta * M)` for small `eta`, the fit depends mainly on the **product** `eta * M`. Numerically: rate 0.3 for 10 rounds, rate 0.15 for 20, and rate 0.03 for 100 all have product 3 and all land roughly 95-97% of the way to the target. Rate 0.03 for only 10 rounds has product 0.3 and reaches about 26% of the way — a badly underfitted model. Real boosting is messier: the trees are refitted each round against changing residuals, the loss is not quadratic, and subsampling adds noise. But the working rule survives — **halve the rate, double the rounds**. ## Why smaller rates are usually better If the trade were exact, the rate would be a free choice. It is not exact, and the asymmetry favours small rates. 1. **Resolution.** At rate 0.3 the ensemble moves in coarse jumps and the best stopping point may sit between two big steps; you can pass straight over the validation minimum between round 40 and round 41. At rate 0.03 the validation curve is smooth and a held-out fold can locate the optimum finely. 2. **Averaging.** Many small corrections behave like a smoother, lower-variance function than a few large ones. Each individual tree's idiosyncrasies get diluted rather than stamped into the model at full strength. This is why the classic advice is to pick the smallest rate you can afford and let the round count follow. ## What it costs Wall-clock, almost linearly. Retuning a display-ad click-through booster from rate 0.3 with 200 rounds to rate 0.03 with 2,000 rounds is roughly a ten-fold training-time increase, and the shipped model now carries ten times as many trees, so scoring latency and memory rise too. On a small table that is a trivial price. On a two-million-row impression log retrained on a schedule, ten times the training time is a real budget decision, and the honest question is how much validation gain the small rate actually buys. ## Where the rule breaks - **A small rate with a capped budget underfits.** Rate 0.01 with 300 rounds is not a regularised model, it is an unfinished one. The failure looks like high training *and* validation error together. - **Round counts do not transfer.** A stopping round found at rate 0.03 means nothing at rate 0.1. Change the rate and the round count must be re-established. - **Diminishing returns.** Past some point, lowering the rate further buys nothing measurable on validation while costing linearly more to train. - **Coupling.** Subsampling, tree size and the rate all interact; the pair (rate, rounds) is one setting, not two independent knobs. ## How to set it in practice Fix a rate you can afford to run, set a generous round budget, let a held-out fold tell you where the validation loss stops improving, and record the (rate, rounds) pair together. Tuning the two as though they were independent wastes most of the search: they mostly move along the same axis.

  • If the rate and rounds are interchangeable, why not just use a large rate and save the time?
    Because the interchange is only approximate. A large rate moves the ensemble in coarse jumps, so the validation minimum falls between two steps and you cannot stop precisely at it; and the model is a sum of a few big corrections rather than many small ones, which is measurably higher-variance. Small rates trade compute for a smoother, better-resolved fit.
  • What exactly does it cost to go from rate 0.3 with 200 rounds to rate 0.03 with 2,000 rounds?
    Roughly ten times the training wall-clock, because per-round cost is unchanged and you run ten times as many rounds. The served model also holds ten times as many trees, so inference latency and memory grow with it. On a small table that is free; on a large log retrained frequently it is a real budget call that has to be justified by validation gain.
  • Someone reports that a booster with rate 0.01 performs worse than one at 0.1 — what do you check first?
    The round budget. At one tenth the rate they need roughly ten times the rounds, so if both runs used the same number of trees the slow one is simply unfinished. The tell is that training error is high too, not just validation error — underfitting, not regularisation.

Shrinkage is walking to a destination in many short steps instead of a few long ones. You arrive at nearly the same place, but you can stop much closer to the exact spot — and it takes longer.

saying these in an interview costs you the question

  • Calls it the gradient step size used inside each tree
  • Says lowering the rate improves the model with the round count unchanged
  • Tunes rate and number of rounds as independent, unrelated knobs
  • Claims the smallest possible learning rate is always best
  • Reuses a round count found at one rate after changing the rate

context

open as a page

In gradient boosting, what do row subsampling and column subsampling per round buy you?

level: middleimportance: should knowfreq 46%

basics

~20 s

Fitting each round's tree on a random fraction of rows and columns decorrelates successive trees, adds a regularising noise to the gradient estimate, and cuts per-round cost. It usually validates better than the full-data fit, until the subsample is so small the fit turns unstable.

open as a page

Your booster's held-out fold stops training at round 740 of 3,000 — how do you report its accuracy and ship that round count?

level: seniorimportance: should knowfreq 57%

basics

~20 s

Treat them separately. The stopping fold chose a hyperparameter, so its score is optimistic and the reported accuracy must come from data that played no part in stopping. Ship 740 as a fixed round count, valid only for the learning rate and subsampling it was found under.

open as a page

Your booster's training loss reaches zero on 5%-mislabelled data while validation loss rises after round 300 — why?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Boosting fits whatever the ensemble still gets wrong, and a mislabelled row is permanently wrong, so later trees carve tiny regions around those rows. Training loss collapses because the noise is being memorised; validation loss turns up because those rounds add nothing real.

open as a page