Your booster's training loss reaches zero on 5%-mislabelled data while validation loss rises after round 300 — why?
answer
- boosting chases whatever is still wrong
- a wrong label never stops being wrong
- the training curve is not a signal
- later trees isolate individual rows
- there is a loss floor you cannot beat
basics
~20 sBoosting fits whatever the ensemble still gets wrong, and a mislabelled row is permanently wrong, so later trees carve tiny regions around those rows. Training loss collapses because the noise is being memorised; validation loss turns up because those rounds add nothing real.
solid answer
~50 sBoosting is sequential error-chasing: every round fits the current residuals, and a row with a wrong label produces a large residual that no amount of genuine structure will explain. So the later trees spend their capacity carving narrow regions around those 5% of rows. Training loss reaches zero because a long enough ensemble of shallow trees can isolate individual points; validation loss turns up at round 300 because from there on the marginal tree is fitting noise rather than signal. The training curve is therefore useless as a stopping signal — only the held-out curve tells you anything. Practical responses: stop at the validation minimum, lower the learning rate so that minimum is better resolved, subsample rows so no mislabelled row appears in every round's fit, and for regression prefer a loss that grows linearly in large residuals, such as Huber, over squared error.
go deeper
Recognise the shape: training loss falling to zero while validation loss climbs means the extra rounds are memorising the training set rather than learning anything transferable.
Explain the mechanism — each round fits the current residuals, and a mislabelled row keeps a large residual forever, so later trees isolate it — and name stopping at the validation minimum as the fix.
Show a plan, not just a diagnosis: stop at the minimum, lower the rate to resolve it, subsample rows, pick a robust regression loss, and mine the persistent large residuals for the suspect rows.
Argue about where the effort should go. Once validation loss sits at the noise floor, further modelling is waste, and the call to stop investing in the model is yours to make and defend.
## Why boosting is specifically vulnerable here Bagging-style averaging fits many models to the same target independently; a mislabelled row contributes one wrong vote and gets outvoted. Boosting does the opposite by design. Each round computes what the current ensemble still gets wrong — the pseudo-residuals — and fits the next tree to exactly that. A row whose label is wrong is a row the ensemble can never explain by learning real structure, so its residual stays large round after round and it keeps presenting itself as the most valuable thing left to fix. The algorithm has no way to distinguish 'hard but real' from 'impossible because the label is wrong'. Both look like a large residual. That is the whole mechanism. Boosting's greatest strength — relentlessly attacking remaining error, which is why it reduces bias where averaging reduces variance — becomes its failure mode when a slice of the remaining error is not error at all but corruption. ## Reading the two curves The pattern in the question is the classic signature: - **Training loss to zero.** Given enough rounds, an additive ensemble of shallow trees can isolate essentially any individual training point: successive trees keep splitting until a mislabelled row sits in a region of its own with the prediction it was told to produce. Zero training loss means the model has memorised its way to the corrupt labels, and it says nothing about generalisation. - **Validation loss turning up after round 300.** Round 300 is roughly where the marginal tree stops finding structure that transfers and starts fitting the corruption. Before it, each round improves both curves; after it, they diverge. That inflection point is the model. The corollary matters more than the diagnosis: **the training curve carries no stopping information**. In a setting with 5% noise it will descend to zero whatever you do. Anyone tuning against it will train forever and ship the worst available model. ## The floor you cannot cross With 5% of labels wrong, there is a lower bound on validation loss that no model can beat: even a perfect predictor of the true relationship disagrees with the corrupted labels on those rows. A team chasing validation loss below that floor will keep adding capacity and keep making the model worse. Recognising that the curve has bottomed out at the noise floor — rather than at some fixable modelling shortfall — is the senior judgment in this scenario, and it is the point at which the effort should move off the model. ## What actually helps 1. **Stop at the validation minimum.** This is the direct fix and it is nearly free: the round count already tells you where the noise-fitting begins. Ship 300-ish rounds, not 3,000. 2. **Lower the learning rate.** At a coarse rate the ensemble crosses the minimum in a few big jumps and you cannot tell where it was; at a small rate the curve is smooth and the turning point is well resolved, so the stopping decision is more accurate. The rate does not make the model immune to noise, but it makes the optimum findable. 3. **Subsample the rows.** If each round fits on 70-80% of the rows, no mislabelled row is present in every round's target, so the ensemble's pursuit of it is intermittent rather than relentless. This measurably softens the turn-up. 4. **Choose a less aggressive loss for regression.** Squared error penalises a residual by its square, so one grossly wrong target dominates the round's fit. Huber loss is quadratic near zero and linear beyond a threshold, and absolute-error loss is linear throughout; both cap how much a single outrageous residual can pull the tree. For classification, log loss is the standard choice and is already gentler than losses that grow exponentially in the margin. 5. **Look at what the last trees are fitting.** The rows with the largest persistent residuals at the end of training are, disproportionately, the corrupt ones. That list is often the most useful artefact of the whole exercise. ## What does not help - **More rounds.** The noise does not average out; boosting is not averaging. Each additional round after the turn makes it worse. - **A larger learning rate 'to get there faster'.** Fewer, coarser steps make the minimum harder to locate, not easier to avoid. - **Judging progress on the training curve.** It will be zero, and it will be zero for a bad model. ## Answering the question in an interview Name the mechanism first — sequential residual-chasing means a permanently-wrong row attracts unbounded attention — then the two curves and what they mean, then the fact that the training curve is not a signal, then the mitigations in order of directness. Finishing with the noise-floor observation is what separates a diagnosis from a plan.
- Why does bagging-style averaging tolerate the same 5% noise better?Because averaging fits each model to the target independently, so a corrupt row influences a minority of the models and gets diluted in the average. Boosting instead directs each new tree at exactly what is still unexplained, which is precisely where the corrupt rows sit, so their influence compounds across rounds instead of washing out.
- How would you use the trained model to find the suspect rows?Look at which training rows still carry the largest residuals late in training, or which rows the ensemble only fits after the validation minimum has passed. Those are disproportionately the mislabelled ones. It is a useful by-product: the model's own difficulty ranking is a cheap candidate list for review.
- Would raising the learning rate limit the damage, since fewer rounds are needed?No. It reduces how many rounds you run, but each one is coarser, so the ensemble crosses the validation minimum in a few big jumps and you can no longer see where the turn happened. You end up with a worse-resolved stopping point on a higher-variance model. A small rate plus stopping at the minimum is strictly better.
A student given unlimited time with an answer key that has typos in it. The longer they study, the more confidently they reproduce the typos.
saying these in an interview costs you the question
- Says zero training loss means the model is good
- Adds more rounds so the noise will average out
- Claims boosting is inherently robust to mislabelled rows
- Tunes the stopping point against the training curve
- Expects validation loss to keep falling below the noise floor