skip to content

How do you construct a learning curve over training-set size so its shape is trustworthy?

level: seniorimportance: should knowfreq 40%

answer

  1. one independent variable only
  2. same evaluation rows at every point
  3. one draw per size is noise
  4. log-spaced sizes, replicated and averaged
  5. refit scaling and encodings inside the subsample

basics

~10 s

Vary only the training-set size. Score every point against one fixed held-out set, draw several random subsamples per size and average them, space sizes logarithmically, and refit preprocessing inside each subsample so nothing leaks.

solid answer

~50 s

The curve is a controlled experiment, so change one thing: the number of training rows. Keep **one held-out evaluation set fixed** across every point, or a change in the curve could be a change in the test sample rather than in the model. At each size draw **several independent random subsamples**, refit, and plot the mean with its spread — a single draw at 1,000 rows is noisy enough to invent a bump that is not there. Space sizes logarithmically (1k, 2k, 5k, 10k, 20k), since gains arrive per doubling. Stratify the subsamples when classes are imbalanced, and sample by group when rows share an entity. Fit every preprocessing step — scaling statistics, encodings, imputation values — **inside** each subsample; fitting them on the full pool leaks information into the small-n points and flattens the curve. Finally, state whether hyperparameters were fixed or re-tuned at each size, because the two conventions give different curves.

go deeper

for a junior

Recall the basic recipe: train on growing random subsamples, score each on the same held-out rows, and plot both errors against size. Knowing that the evaluation set must not change is the key detail here.

for a middle

Explain why each control matters — what breaks if the test sample moves, what a single draw at small size does to the shape, and why preprocessing must be refit per subsample rather than once.

for a senior

Demonstrate that you have produced these curves for real: replicate counts, spread bands, group-aware subsampling, a stated tuning convention, and the discipline to say 'cannot tell at this resolution' when bands overlap.

for a principal

Own the standard: a curve that will justify spending on data must be reproducible and carry its uncertainty, because someone will act on its slope. Decide how much compute the organisation spends on getting that slope right.

## The curve is an experiment, and it has one independent variable Everything below follows from a single principle: a learning curve is only interpretable if **training-set size is the only thing that changes** between points. Every classic mistake in drawing one is some other variable moving at the same time. ## Hold the evaluation set fixed Score every trained model on the **same** held-out rows. A tempting alternative — re-split the data at each size, so that a smaller training set leaves a larger test set — sounds efficient and quietly ruins the plot: the validation numbers are now computed on different samples of different sizes, and a step in the curve could equally be a step in the difficulty of that particular test sample. Fix the evaluation set once, before the ladder starts, and never touch it during the sweep. ## Average repeated draws at each size At small sizes, which model you get depends heavily on **which** rows you happened to draw. Two independent 1,000-row subsamples can differ by several points of held-out error. Plotting one draw per size gives a jagged line whose bumps are pure sampling noise and which people then over-interpret. Instead, draw several subsamples per size — five is a common working number, more at the small end where variability is worst — refit each, and plot the mean along with a band showing the spread. The band is itself informative: a wide band at small n is the variance story showing up before the gap does. A useful refinement is **nested subsamples**: make each larger sample a superset of the smaller one within a replicate, so the curve traces the effect of *adding* rows rather than the effect of drawing an unrelated sample. ## Space the sizes logarithmically Going 1k, 2k, 5k, 10k, 20k tells you far more than 4k, 8k, 12k, 16k, 20k, because gains arrive per doubling, not per row. Evenly spaced sizes waste most of the compute on the flat right-hand end of the curve and leave the steep, informative region under-sampled. Log spacing also matches how the curve will be read and extrapolated later. ## Fit preprocessing inside each subsample Every quantity learned from data — a scaling mean and spread, a category-to-number mapping, an imputation fill value, a target-based encoding, a feature-selection step — must be recomputed **from the subsample being trained on**. If you compute them once on the full training pool and reuse them at every size, then the 1,000-row model quietly benefits from statistics over all 20,000 rows. The small-n points come out too good, the curve looks flatter than it is, and you conclude data does not help when it does. This is the subtlest and most common defect in a hand-rolled curve. ## Decide, and state, the hyperparameter convention There are two defensible conventions and they give different pictures: - **Fixed settings.** Use one configuration at every size. Simple and cheap, but a configuration chosen for 20,000 rows is usually too flexible for 1,000, so the small-n end looks worse than it needs to and variance is exaggerated. - **Re-tuned at each size.** Repeat the selection procedure within each subsample. This shows what a competent practitioner would actually achieve at each data scale, which is the honest input to a data-acquisition decision — at the cost of much more compute, and with the requirement that all tuning happen inside the subsample. Neither is wrong. Reporting which one you used is mandatory, because a reader comparing your curve with another's will otherwise compare two different experiments. ## Keep the subsamples representative Draw **stratified** subsamples when the target is imbalanced, so a 500-row point is not accidentally almost free of positives. When rows are grouped — several records per patient, per document, per customer — subsample by **group**, not by row, or the same entity appears in both training and evaluation and the small-n points are flattered again. If the data is a time series, respect ordering rather than sampling rows at random. ## Read the noise band before reading the trend When the plot is finished, look at the spread bands first. If successive points overlap heavily, the apparent slope between them is not evidence of anything, and the honest conclusion is "cannot tell at this resolution" — either add replicates or accept a coarser reading. A curve reported without any notion of its own uncertainty invites confident conclusions the experiment does not support. ## The checklist One fixed evaluation set; several replicates per size with a spread band; log-spaced sizes; preprocessing refit inside each subsample; a stated hyperparameter convention; stratified or group-aware sampling. Miss any one and the shape you read may be an artefact of the procedure rather than a property of the problem.

  • Should hyperparameters be re-tuned at each training-set size, or held fixed?
    Either, provided you say which. Fixed settings are cheap but a configuration chosen for the full data is usually too flexible for the smallest sizes, exaggerating the left-hand error. Re-tuning inside each subsample shows what is actually achievable at each data scale, which is the more honest basis for a data-acquisition decision, at a much higher compute cost.
  • Why refit scaling statistics and category encodings inside each subsample?
    Because those quantities are learned from data. Computing them once over the full pool lets a 1,000-row model borrow information from 20,000 rows, so the small-size points come out too good and the curve looks flatter than reality. You then conclude that extra rows buy nothing when in fact your procedure gave the small models a preview of them.
  • Your rows are grouped — many records per patient. How does that change the subsampling?
    Subsample by group rather than by row, and split the evaluation set by group too. Sampling rows independently puts records from the same patient on both sides, so the model recognises the entity rather than generalising, and every point on the curve is optimistic — most severely at small sizes, which distorts the whole shape.

saying these in an interview costs you the question

  • Re-splits the data at every training size
  • Plots a single random subsample per size
  • Fits scaling and encodings once on the full data
  • Uses evenly spaced sizes over a wide range
  • Reports the curve with no notion of its own noise

context