skip to content

On a learning curve over training-set size, what do converged curves versus a persistent gap tell you?

level: middleimportance: must knowfreq 72%

answer

  1. read the shape, not one score pair
  2. does the gap close as rows grow?
  3. level measures error, gap measures variance
  4. met and flat means the model is the limit
  5. wide at every size means the sample is the limit

basics

~20 s

Curves that meet at a high error mean high bias: the model is too simple, and more rows will not help. A gap that stays wide at every training size means high variance, where more rows still help.

solid answer

~50 s

A learning curve puts the number of training rows on the x-axis and plots two errors: the error on the rows the model was trained on, and the error on a fixed held-out set. As rows are added, training error rises (the model can no longer nearly memorise) and validation error falls, so the two curves move toward each other. If they meet and flatten together at, say, 8% error, that is the high-bias shape: the model plus its features cannot do better, and buying rows will not move it. If validation error sits about 12 points above training error at every size, that is the high-variance shape: the model is fitting sample-specific noise, and the still-closing gap says more rows are genuinely buying accuracy. Read the trend across sizes, never a single train/validation pair.

go deeper

for a junior

Be ready to say what is on each axis and which curve is which: rows on the x-axis, training error and held-out error as two lines. Recall that the two lines moving apart and the two lines meeting mean different things.

for a middle

Explain the mechanics: why training error rises and validation error falls with more rows, and why the gap and the level are two separate readings. An interviewer will hand you a sketched pair of curves and expect a named verdict.

for a senior

Show you have read these curves on real projects: catching the flat-and-terrible case that was actually a broken join, stating the verdict at the data scale you operate at, and refusing to diagnose from a single training run.

for a principal

Own the framing question — the curve exists to decide whether data or model capacity is the binding constraint, and that decision commits budget. Be able to say when the curve is not worth drawing because the answer is already known.

## What a learning curve is A **learning curve** plots error against the **number of training rows**. You choose a set of training-set sizes — 1,000, 2,000, 5,000, 10,000, 20,000 is a typical log-spaced ladder — and at each size you train the same model on a random subsample of that many rows. You then record two numbers: the error on the rows the model just trained on (the *training* curve) and the error on a held-out evaluation set that is identical for every point (the *validation* curve). Plotting both against size gives two lines. **The shapes carry the diagnosis; the individual points do not.** This is different from any curve swept over a model setting: here the model and its settings are held constant and only the amount of data changes. The question a learning curve answers is therefore very specific — *would more rows help?* ## Why the two curves move toward each other With a handful of rows, a reasonably flexible model can come very close to reproducing them exactly, so training error starts near zero. It has, however, learned the peculiarities of those particular rows, so it generalises badly and validation error starts high. Add rows and the model can no longer satisfy them all; it is forced to compromise, and **training error rises**. At the same time the estimate it forms is less driven by the accidents of one small sample, so **validation error falls**. Both curves are converging on the same quantity — the error the model family achieves on the population. ## The two canonical shapes **1. Converged and flat.** Both curves meet — say at 8% error — and then run flat as size increases. The gap between them is essentially gone, which means sampling noise is no longer the limitation. What remains is what this model, on these features, is capable of, plus whatever error is irreducible. Extra rows have nothing left to buy: you are already estimating the best function this family can express, and it is not good enough. This is the **high-bias / underfitting** verdict. **2. A persistent, wide gap.** Training error is low, validation error is far above it — 12 points apart, and still 12 points apart at your largest size, or narrowing only slowly. The model can fit the sample beautifully and cannot carry that to new data, which is the definition of fitting noise. This is the **high-variance / overfitting** verdict, and here the curve is telling you that data is the lever: as long as validation error is still falling, each new block of rows is still paying. ## Reading the trend, not the endpoint The single most common error is to look at one training run — "train 2%, validation 14%" — and stop. That pair is consistent with both stories: a variance problem that more rows would fix, and a model that has already saturated and happens to be over-parameterised for the sample. Only the curve across sizes separates them, because only the curve shows whether the gap is closing and whether the validation line is still descending. Equally, do not read a small gap as good news. A model that is uniformly terrible on both curves has a small gap; it is simply consistently wrong. **The gap measures variance; the level measures total error.** You need both readings. ## Practical cautions - **Flat and terrible at every size** can also mean the pipeline is broken — a target leaked into nothing, mis-joined labels, a feature column that is all missing. Before declaring high bias, sanity-check that a deliberately over-flexible model can drive training error down at all; if it cannot, the problem is not bias, it is a bug. - **The two curves are not measured on the same number of rows.** Training error at n=1,000 is an average over 1,000 rows; validation error is over the fixed held-out set. That is fine — you are comparing in-sample to out-of-sample error, not two samples of equal size — but it does mean the training curve is noisier at small n. - **Curves can cross the story mid-way.** A model may look high-variance up to 5,000 rows and converged by 50,000. The verdict is always stated *at the data scale you are operating at*, not in the abstract. - **New rows must come from the same population** you will be scored on. Twenty thousand more rows drawn from a source you never serve do not move the curve you care about; they change the problem. ## What the shapes commit you to A converged plateau is a statement that the current model and feature representation are exhausted; a persistent gap is a statement that the sample, not the model family, is the binding constraint. Naming the shape correctly is the whole point of drawing the curve: everything you do next is downstream of getting this call right.

  • Why does error on the training rows typically rise as the training set grows?
    With few rows a flexible model can nearly memorise them, so its in-sample error is optimistically low. As rows are added it can no longer satisfy every one of them and must compromise, so training error climbs toward the error the model family actually achieves on the population. A rising training curve is the expected, healthy pattern, not a warning sign.
  • Both curves plateau together at a high error — could more rows ever still help?
    Not from the same population with the same model and features: the plateau is what that combination can do. The exception is rows that broaden coverage rather than deepen it — new sites, new devices, rare classes barely present in the current sample. Those change the problem the curve was drawn for, so you redraw it rather than extrapolate the old one.
  • Both curves are flat and terrible at every size. How do you tell high bias from a broken pipeline?
    Try to overfit on purpose. Take a few hundred rows and fit a deliberately over-flexible model; a healthy pipeline will drive training error near zero on that tiny sample. If it cannot, the signal is not reaching the model at all — misaligned labels, a leaked-out target, an all-missing feature block — and no amount of capacity or data will help until that is fixed.

Two runners on a track: if they finish side by side but slowly, the pace is the limit and a longer race changes nothing. If one is always far ahead of the other, the trailing one is still catching up as the distance grows.

saying these in an interview costs you the question

  • Reads one train/validation score pair instead of the trend
  • Says more data always improves any model
  • Calls a converged high-error curve overfitting
  • Treats a small train-validation gap as proof the model is good
  • Expects training error to fall as the training set grows

context