What does a validation curve plot, and how do you read it to pick a hyperparameter value?
answer
- one dial, two lines
- x-axis is capacity, not data size
- look where validation turns over
- the distance between the lines means variance
basics
~20 sA validation curve plots training score and held-out validation score against one hyperparameter that controls model capacity. Pick the value where validation peaks; a widening gap between the two lines past that point is the overfitting signal.
solid answer
~50 sA validation curve sweeps a single capacity dial — decision-tree depth, the number of neighbours, a polynomial degree — and plots both the training score and the held-out validation score at every value on that dial. Reading it left to right: at low capacity both lines are low and close together, which is underfitting; as capacity rises both improve; past some point training keeps climbing toward perfect while validation flattens and then falls, and the widening gap between the lines is the variance signal. I choose the value where validation peaks, not where the gap is smallest. On a 1,400-row employee-attrition model swept from depth 1 to 30, validation F1 peaked at depth 5 with a 1-point gap, while depth 20 reached 0.99 train against 0.79 validation — a 20-point gap and a worse model. Score each point with cross-validation, not one split.
go deeper
Be ready to say out loud what each axis and each line is, and to point at the region of a sketched curve where the model overfits. Knowing that you pick the peak of the validation line, not the training line, is the whole screening bar here.
Explain the mechanics: one refit per point, why the training line keeps rising as capacity grows, and why the two lines diverge. Expect to be handed three train/validation pairs and asked which value ships and why.
Show the discipline around the plot — cross-validating each point, choosing the metric that matches the business cost, never scoring on the test set, and refusing to over-read differences smaller than the fold-to-fold jitter.
Own the framing that this curve is a one-dimensional slice of a multi-dimensional response surface, conditional on every other setting. Argue when a team should stop turning dials because the curve is flat and the real ceiling is data or features.
## What a validation curve is A validation curve is a plot with **one hyperparameter on the x-axis** and a **performance score on the y-axis**, carrying two lines: the score measured on the data the model was fitted on (the *training score*) and the score measured on data held out from fitting (the *validation score*). A hyperparameter is a setting you choose before fitting — it is not learned from the data the way weights or split points are. One point on the plot means: refit the model from scratch with the hyperparameter set to that value, then measure both scores. Sweeping thirty values means fitting thirty models. Nothing is reused between points. The dial on the x-axis is deliberately a **capacity** dial — a setting that controls how flexible the fitted function is allowed to be. Typical choices: the maximum depth of a decision tree, the degree of a polynomial expansion, the number of neighbours a nearest-neighbour classifier votes over, the minimum number of samples required in a leaf. A validation curve is *not* a plot over training-set size (that is a different diagnostic, drawn with the hyperparameters held fixed), and it is not a plot over a decision threshold. ## The three regions of the curve **Underfitting (left of the peak).** Capacity is too low for the signal in the data. Both lines sit low and almost on top of each other — the model cannot even fit the training data well, so there is nothing for it to overfit. This is the high-bias regime. **The sweet spot.** Validation reaches its maximum (or, if you plot error, its minimum). This is the value you ship. **Overfitting (right of the peak).** Training score keeps improving — often all the way to near-perfect — while validation flattens and then declines. The model is now fitting noise specific to the training rows. This is the high-variance regime, and the **train-minus-validation gap** is its fingerprint. ## The gap is a signal, not the criterion On the attrition sweep above, depth 1 gives train 0.55 / validation 0.54 — a 1-point gap, and a bad model. Depth 5 gives 0.86 / 0.84 — a 2-point gap and the best validation score on the grid. Depth 20 gives 0.99 / 0.79 — a 20-point gap and a clearly worse model. The common beginner error is to optimise the gap. You do not: you optimise **validation score**, and read the gap as the explanation for why validation is falling. A model with a 5-point gap and 0.88 validation beats a model with a 1-point gap and 0.70 validation every time. ## Drawing one honestly - **Cross-validate each point** rather than scoring against a single split. A single split produces a jagged curve whose apparent peak moves when you change the random seed, and you will tune to that noise. - **Choose grid spacing to match the dial.** Integer dials like depth take every value over a small range; dials spanning orders of magnitude are swept multiplicatively. - **Plot the metric you actually care about.** A depth that maximises accuracy on an imbalanced attrition problem is often not the depth that maximises recall on the leavers. - **Never draw the curve on the test set.** The curve is a tuning instrument; the test set is spent once, at the end. ## What the curve does not tell you Every point on it is computed with **all the other hyperparameters held fixed** at whatever you chose. The best depth given a minimum leaf size of 1 need not be the best depth given a minimum leaf size of 50 — the dials interact. So a validation curve is a *diagnostic*: it shows you the shape of the model's response to one dial and tells you which regime you are in. It does not, on its own, prove you have found the best configuration. It also cannot tell you whether a different model family, better features, or more rows would beat everything on the curve. A curve that is flat and low across the entire sweep is telling you the dial is not the binding constraint. Finally, neighbouring points near the peak are often indistinguishable given how much the score jitters from fold to fold. When depth 5, 6 and 7 all land within that jitter, take the one that is cheapest to train and to serve — you are not throwing away accuracy you actually have evidence for.
- Why score each point on the curve with cross-validation instead of a single validation split?A single split gives one noisy estimate per point, so the curve comes out jagged and the apparent peak moves when you change the split. Averaging over folds smooths that jitter and makes the peak stable enough to act on. It also tells you how much fold-to-fold variation there is, which is what you compare small differences against before believing them.
- The curve holds every other hyperparameter fixed. Why does that matter when you act on it?Because capacity dials interact. The best tree depth when leaves may hold a single row is not the best depth when leaves must hold fifty, so the curve you drew is conditional on those other settings. Treat it as a diagnosis of the regime you are in and the shape of the response, then confirm the value jointly rather than shipping it as a proven optimum.
- Both lines are low and sit on top of each other across the whole sweep. What are you looking at?Underfitting everywhere on that dial. The model cannot fit even the training data, so there is no variance problem to trade against — turning the capacity dial further will not help. The curve is telling you the constraint lies elsewhere: the feature representation, the model family, or the signal available in the data.
It is like tuning a radio dial while watching two needles: one for how loud the station sounds in the studio, one for how it sounds in the car. Past a point the studio needle keeps rising while the car needle drops.
saying these in an interview costs you the question
- Picks the hyperparameter value with the best training score
- Optimises for the smallest train-validation gap instead of validation score
- Confuses the x-axis with the number of training rows
- Draws the curve against the test set
- Trusts every wiggle of a curve scored on one split