skip to content

Diagnosing Fit

Reading learning and validation curves to tell a high-bias model from a high-variance one, then acting on it. Interviewers hand you a train/validation score pair and ask what you would do next.

on this pageshow

explore

questions

10

What does a validation curve plot, and how do you read it to pick a hyperparameter value?

level: juniorimportance: must knowfreq 62%

answer

  1. one dial, two lines
  2. x-axis is capacity, not data size
  3. look where validation turns over
  4. the distance between the lines means variance

basics

~20 s

A validation curve plots training score and held-out validation score against one hyperparameter that controls model capacity. Pick the value where validation peaks; a widening gap between the two lines past that point is the overfitting signal.

solid answer

~50 s

A validation curve sweeps a single capacity dial — decision-tree depth, the number of neighbours, a polynomial degree — and plots both the training score and the held-out validation score at every value on that dial. Reading it left to right: at low capacity both lines are low and close together, which is underfitting; as capacity rises both improve; past some point training keeps climbing toward perfect while validation flattens and then falls, and the widening gap between the lines is the variance signal. I choose the value where validation peaks, not where the gap is smallest. On a 1,400-row employee-attrition model swept from depth 1 to 30, validation F1 peaked at depth 5 with a 1-point gap, while depth 20 reached 0.99 train against 0.79 validation — a 20-point gap and a worse model. Score each point with cross-validation, not one split.

go deeper

for a junior

Be ready to say out loud what each axis and each line is, and to point at the region of a sketched curve where the model overfits. Knowing that you pick the peak of the validation line, not the training line, is the whole screening bar here.

for a middle

Explain the mechanics: one refit per point, why the training line keeps rising as capacity grows, and why the two lines diverge. Expect to be handed three train/validation pairs and asked which value ships and why.

for a senior

Show the discipline around the plot — cross-validating each point, choosing the metric that matches the business cost, never scoring on the test set, and refusing to over-read differences smaller than the fold-to-fold jitter.

for a principal

Own the framing that this curve is a one-dimensional slice of a multi-dimensional response surface, conditional on every other setting. Argue when a team should stop turning dials because the curve is flat and the real ceiling is data or features.

## What a validation curve is A validation curve is a plot with **one hyperparameter on the x-axis** and a **performance score on the y-axis**, carrying two lines: the score measured on the data the model was fitted on (the *training score*) and the score measured on data held out from fitting (the *validation score*). A hyperparameter is a setting you choose before fitting — it is not learned from the data the way weights or split points are. One point on the plot means: refit the model from scratch with the hyperparameter set to that value, then measure both scores. Sweeping thirty values means fitting thirty models. Nothing is reused between points. The dial on the x-axis is deliberately a **capacity** dial — a setting that controls how flexible the fitted function is allowed to be. Typical choices: the maximum depth of a decision tree, the degree of a polynomial expansion, the number of neighbours a nearest-neighbour classifier votes over, the minimum number of samples required in a leaf. A validation curve is *not* a plot over training-set size (that is a different diagnostic, drawn with the hyperparameters held fixed), and it is not a plot over a decision threshold. ## The three regions of the curve **Underfitting (left of the peak).** Capacity is too low for the signal in the data. Both lines sit low and almost on top of each other — the model cannot even fit the training data well, so there is nothing for it to overfit. This is the high-bias regime. **The sweet spot.** Validation reaches its maximum (or, if you plot error, its minimum). This is the value you ship. **Overfitting (right of the peak).** Training score keeps improving — often all the way to near-perfect — while validation flattens and then declines. The model is now fitting noise specific to the training rows. This is the high-variance regime, and the **train-minus-validation gap** is its fingerprint. ## The gap is a signal, not the criterion On the attrition sweep above, depth 1 gives train 0.55 / validation 0.54 — a 1-point gap, and a bad model. Depth 5 gives 0.86 / 0.84 — a 2-point gap and the best validation score on the grid. Depth 20 gives 0.99 / 0.79 — a 20-point gap and a clearly worse model. The common beginner error is to optimise the gap. You do not: you optimise **validation score**, and read the gap as the explanation for why validation is falling. A model with a 5-point gap and 0.88 validation beats a model with a 1-point gap and 0.70 validation every time. ## Drawing one honestly - **Cross-validate each point** rather than scoring against a single split. A single split produces a jagged curve whose apparent peak moves when you change the random seed, and you will tune to that noise. - **Choose grid spacing to match the dial.** Integer dials like depth take every value over a small range; dials spanning orders of magnitude are swept multiplicatively. - **Plot the metric you actually care about.** A depth that maximises accuracy on an imbalanced attrition problem is often not the depth that maximises recall on the leavers. - **Never draw the curve on the test set.** The curve is a tuning instrument; the test set is spent once, at the end. ## What the curve does not tell you Every point on it is computed with **all the other hyperparameters held fixed** at whatever you chose. The best depth given a minimum leaf size of 1 need not be the best depth given a minimum leaf size of 50 — the dials interact. So a validation curve is a *diagnostic*: it shows you the shape of the model's response to one dial and tells you which regime you are in. It does not, on its own, prove you have found the best configuration. It also cannot tell you whether a different model family, better features, or more rows would beat everything on the curve. A curve that is flat and low across the entire sweep is telling you the dial is not the binding constraint. Finally, neighbouring points near the peak are often indistinguishable given how much the score jitters from fold to fold. When depth 5, 6 and 7 all land within that jitter, take the one that is cheapest to train and to serve — you are not throwing away accuracy you actually have evidence for.

  • Why score each point on the curve with cross-validation instead of a single validation split?
    A single split gives one noisy estimate per point, so the curve comes out jagged and the apparent peak moves when you change the split. Averaging over folds smooths that jitter and makes the peak stable enough to act on. It also tells you how much fold-to-fold variation there is, which is what you compare small differences against before believing them.
  • The curve holds every other hyperparameter fixed. Why does that matter when you act on it?
    Because capacity dials interact. The best tree depth when leaves may hold a single row is not the best depth when leaves must hold fifty, so the curve you drew is conditional on those other settings. Treat it as a diagnosis of the regime you are in and the shape of the response, then confirm the value jointly rather than shipping it as a proven optimum.
  • Both lines are low and sit on top of each other across the whole sweep. What are you looking at?
    Underfitting everywhere on that dial. The model cannot fit even the training data, so there is no variance problem to trade against — turning the capacity dial further will not help. The curve is telling you the constraint lies elsewhere: the feature representation, the model family, or the signal available in the data.

It is like tuning a radio dial while watching two needles: one for how loud the station sounds in the studio, one for how it sounds in the car. Past a point the studio needle keeps rising while the car needle drops.

saying these in an interview costs you the question

  • Picks the hyperparameter value with the best training score
  • Optimises for the smallest train-validation gap instead of validation score
  • Confuses the x-axis with the number of training rows
  • Draws the curve against the test set
  • Trusts every wiggle of a curve scored on one split

context

open as a page

On a learning curve over training-set size, what do converged curves versus a persistent gap tell you?

level: middleimportance: must knowfreq 72%

basics

~20 s

Curves that meet at a high error mean high bias: the model is too simple, and more rows will not help. A gap that stays wide at every training size means high variance, where more rows still help.

open as a page

A model hits 2% training error but 14% validation error - what do you try first?

level: middleimportance: must knowfreq 78%

basics

~20 s

A 12-point train-validation gap is high variance, not high bias. Rank the fixes by cost: more training rows, then a stronger penalty or fewer features, then a lower-capacity model. Adding capacity would make the gap worse.

open as a page

Why does a model's held-out error flatten above zero no matter how many training rows you add?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Part of the error is irreducible: the label is not fully determined by the features, so near-identical inputs carry different labels. Data cannot remove that. The rest of the floor is the model's own bias, which rows also cannot fix.

open as a page

A k-NN validation curve over k = 1 to 200 shows zero training error at k = 1. Which end is high capacity?

level: middleimportance: should knowfreq 45%

basics

~20 s

The small-k end. At k = 1 every training point is its own nearest neighbour, so training error is zero by construction. Capacity falls as k grows, so this curve's complexity axis runs right to left.

open as a page

How do you construct a learning curve over training-set size so its shape is trustworthy?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Vary only the training-set size. Score every point against one fixed held-out set, draw several random subsamples per size and average them, space sizes logarithmically, and refit preprocessing inside each subsample so nothing leaks.

open as a page

A speech model shows 9% training and 10% validation word error rate, humans 6% - what next?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Most of the error is avoidable bias: three points sit between the model's 9% training error and the 6% human benchmark, against a one-point train-validation gap. Attack bias first - more capacity, richer features, weaker regularization - not more data.

open as a page

Your validation curve is still improving at the largest hyperparameter value you swept — what do you do?

level: seniorimportance: nice to knowfreq 32%

basics

~10 s

An edge value is a property of your grid, not of the model: the sweep ended before the optimum. Extend the range and refit until validation clearly flattens or turns over, then choose.

open as a page

Given a learning curve over training-set size, how do you decide whether 20,000 more labels are worth buying?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Extrapolate the curve one or two doublings at most, price the projected error drop in hours and money, and ask whether it changes a decision the product makes. Buy a small tranche first to test the projection.

open as a page

Expert annotators agree only 92% of the time - how much headroom does your model really have?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Pairwise agreement of 92% is not an 8% error floor. If annotators err independently, each is wrong on roughly 4% of items. Adjudicate a sample to split irreducible ambiguity from fixable guideline drift before funding more modelling.

open as a page