skip to content

Why does training error fall monotonically as capacity grows while test error is U-shaped?

level: middleimportance: must knowfreq 74%

answer

  1. measured on the data it optimised
  2. the bigger class contains the smaller
  3. two forces pulling opposite ways
  4. no minimum to read on one curve

basics

~20 s

Enlarging a nested hypothesis space keeps every earlier fit available, so training error can only fall. Test error instead trades falling systematic error against rising sensitivity to the sample, so it bottoms out and climbs.

solid answer

~50 s

Training error is measured on exactly the data the fit was chosen to minimise. If the larger model class contains the smaller one - degree-15 polynomials contain every straight line, a depth-30 tree contains every depth-3 tree - then whatever the smaller class achieved is still reachable, so the minimum training error is monotonically non-increasing in capacity. Test error is measured on data the fit never saw, and it moves under two opposing forces: as capacity grows the systematic error of the class shrinks, while sensitivity to the particular sample grows. The first dominates early, the second dominates late, so the curve is U-shaped with a minimum somewhere in between. The practical consequence is the whole point: training error can never select capacity, because it always votes for the largest class on offer. Only held-out data locates the minimum.

go deeper

for a junior

Be able to sketch the two curves from memory: one falling all the way, one dipping and rising, with the gap between them widening on the right. Say which of the two is measured on data the model never saw.

for a middle

Explain the containment argument for the training curve in your own words, and name the two opposing effects behind the turning point of the held-out curve without hand-waving about complexity.

for a senior

Demonstrate that you check whether the held-out split is genuinely independent before trusting a small gap, and that you treat the location of the minimum as specific to this dataset rather than a transferable setting.

for a principal

Be ready to argue for the evaluation discipline this implies across a team - that a capacity dial tuned against training numbers is untuned - and to decide how much data to reserve for the estimate that actually locates the minimum.

## Two curves, drawn against the same x-axis Put capacity on the horizontal axis - polynomial degree 1, 2, 3, ..., 15 fitted to the same 60 apartment-rent listings, or tree depth 1, 2, 3, ..., 30 on the same 5,000 listings. Plot two things: the error the fitted model achieves on the listings it was fitted to, and the error it achieves on listings held back. They behave completely differently, and understanding why is the core of this topic. ## Why training error only goes down The argument is a containment argument, not a statistical one. Suppose the classes are **nested**: every function expressible at degree 5 is also expressible at degree 6 (set the extra coefficient to zero); every partition a depth-3 tree can make is also available to a depth-4 tree. Now the fitting procedure minimises training loss over the class. Moving to the larger class hands the optimiser a superset of the options it already had, so the minimum it can reach is at least as good. Therefore minimum training error is non-increasing as capacity grows, and in practice strictly decreasing because the extra freedom is almost always worth a little on the sample at hand. Push this to the extreme and training error reaches zero: with 60 listings, a polynomial of degree 59 can in general pass exactly through all of them, and a tree grown until every leaf holds one listing reproduces the training set perfectly. Perfect training error is not an achievement; it is what unlimited capacity does by construction. Two caveats a good answer includes. First, the monotonicity needs nesting - comparing two *different* families (a shallow tree against a linear model) gives you no such guarantee. Second, it needs the optimiser to actually reach the minimum; with a procedure that stops early or gets stuck, a larger class can come out slightly worse on the training data. ## Why test error turns around Held-out error is not measured on the data that was optimised, so the containment argument does not apply to it. Instead two effects run in opposite directions as capacity grows: - **Systematic error falls.** A larger class contains functions closer to the true relationship, so the best-case approximation improves. On the rent data, moving from a straight line to a gentle curve genuinely captures how price per square metre changes with size. - **Sensitivity to the sample rises.** A larger class can fit the particular 60 listings you drew - the odd landlord, the mistyped floor area. Draw a different 60 listings and you get a visibly different curve. That instability is error on new data even when the class is, on average, well-centred. Early on, the first effect is large and the second is small, so held-out error drops. Later, the class is already rich enough to express the pattern, so there is little systematic error left to remove, while every extra degree of freedom adds instability. The curve bottoms out and rises. The minimum is the capacity you want, and its location is a property of *this* dataset, not a universal constant. ## The consequence that matters in practice Because the training curve falls monotonically, it has no minimum to read off. Choosing capacity by training error always selects the largest class you offered it. This is precisely why a held-out estimate exists at all: it is the only one of the two curves that has an interior minimum to find. Any workflow that tunes a capacity dial while looking at training numbers is not tuning anything - it is just walking to the right-hand edge of the plot. ## Reading the gap The vertical distance between the two curves is informative on its own. - **Small gap, both errors high** - you are left of the minimum; the class is too restrictive. - **Large gap, training error near zero** - you are right of the minimum; the fit is tracking sample-specific structure. - **Small gap, both errors low** - you are near the bottom. A warning about that last case: a small gap is not proof of a good model if the held-out split is not really independent - duplicated listings, or the same building appearing on both sides of the split, will shrink the gap for reasons that have nothing to do with capacity. ## What the classical picture assumes The U-shape as described assumes you fit each capacity setting to convergence on a fixed sample and evaluate on data from the same distribution. It is a statement about the sample size you have: with more data the curve's minimum sits further right, and the whole held-out curve sits lower. It is not a claim that some absolute degree or depth is universally correct.

  • Is falling training error guaranteed for any way of increasing capacity?
    Only under two conditions: the larger class must contain the smaller one, and the fit must actually reach the training-loss minimum. Comparing unrelated families gives no guarantee, and an optimiser that stops early or lands in a poor solution can make a richer class score worse on the training data than a simpler one.
  • Why can the training curve never tell you where to stop increasing capacity?
    Because it has no interior minimum - it falls all the way to the right-hand edge, so reading a stopping point off it always means picking the largest class you tried. The turning point exists only in the held-out curve, which is why an independent estimate is not optional.
  • What does a near-zero training error tell you by itself?
    Almost nothing about quality. It tells you the class had enough capacity to reproduce the sample, which any sufficiently rich class can do by construction - a tree grown to one listing per leaf achieves it trivially. Whether that fit generalises is a separate question that only unseen data answers.

saying these in an interview costs you the question

  • Says training error is U-shaped too
  • Picks a capacity setting from training error
  • Treats zero training error as a success
  • Claims a bigger model always generalises worse
  • Ignores that monotonicity needs nested classes

context