Why does tuning a model require a validation set separate from the test set?
answer
- every peek costs you honesty
- selection is itself a fit
- maximum of noisy scores is optimistic
- validation absorbs the tuning optimism
- one partition, one final number
basics
~20 sEvery choice made by looking at a score - hyperparameters, features, thresholds - bends the model toward those rows. Validation absorbs that optimism; the untouched test set then gives an honest estimate of performance on new data.
solid answer
~50 sThe training set fits the parameters, the validation set is where you make decisions - which hyperparameters, which feature set, which decision threshold, when to stop - and the test set exists to produce one honest number after everything is frozen. Selection is itself a kind of fitting: if you compare many candidates on the same rows and keep the winner, you are taking the maximum of several noisy estimates, and the maximum of noise is optimistic. That optimism attaches to whichever partition drove the choices, which is exactly why the reported estimate has to come from rows that drove none of them. So a two-way split plus tuning gives you a number that flatters the model by an amount you cannot measure from the inside. Practically: split first, decide on validation, freeze, then score the test set once and report it with its uncertainty.
go deeper
Be ready to name the three partitions and say what each one is for: fitting parameters, comparing candidates, and producing one final estimate. Knowing that the test set is used once is most of the answer here.
Explain why making a choice from a score is a form of fitting, and describe how the optimism grows with the number of candidates compared. Interviewers expect the mechanism, not just the rule.
Show how you enforce the separation on a real project: when the hold-out is unsealed, who may run against it, and what you do when you discover somebody already peeked before the model was frozen.
Own the tradeoff between spending rows on an honest verdict and spending them on a better model, and set the team's standing policy for how often a hold-out is refreshed and how its use is reported.
## Three partitions, three jobs **Training set** - the rows the learning algorithm fits parameters on: regression coefficients, tree split points and leaf values, cluster centres, support vectors. This is the only partition the optimiser is allowed to touch. **Validation set** - rows used to *compare* things. No parameters are fitted here, but decisions are made here: how deep the trees go, how strong the regularisation is, which of two feature sets wins, where the classification threshold sits, when to stop adding boosting rounds. The output of the validation set is a *choice*. **Test set** - rows used exactly once, after every choice is frozen, to produce the number you report to other people. Its output is an *estimate*. The mistake the three-way split exists to prevent is using one partition for both of the last two jobs. ## Why selecting is a kind of fitting A score computed on a finite sample is an estimate with sampling error. Suppose you evaluate `k` candidate configurations on the same validation rows and each estimate has standard error `sigma`. Picking the best one means taking the maximum of `k` noisy numbers, and the maximum of noise is positive on average. If the candidates were genuinely equally good, the expected score of the winner sits roughly `sigma * sqrt(2 * ln k)` above the truth. With 100 candidates and a standard error of 1.5 accuracy points, that is about 4.5 points of pure illusion - `sqrt(2 * ln 100)` is about 3.0 - even though no candidate is actually better than any other. Nothing about this requires the optimiser to have touched the rows. Information flows out of a dataset every time a human or a search loop looks at a number computed on it and acts on that number. That is why "we never trained on the test set" is not a defence. ## Relative questions versus absolute questions The validation set answers a *relative* question: which of these candidates is better? A relative comparison tolerates a bias that is shared across candidates, because the bias largely cancels when you rank them. The test set answers an *absolute* question: how well will this specific frozen model do on data like this? An absolute claim has no cancellation to lean on, so it needs rows that carry no selection history at all. This is the cleanest way to explain in an interview why one hold-out cannot do both jobs. ## The order of operations 1. Split the rows into train, validation and test **before** making any decision informed by looking at data. 2. Fit candidates on train, score them on validation, pick a winner. 3. Freeze the configuration - hyperparameters, features, threshold, preprocessing choices. 4. Optionally refit the frozen configuration on train plus validation, since those validation rows are now just labelled data. 5. Score the test set **once** and report the number together with its uncertainty. 6. If the test number disappoints and you go back to step 2, say so out loud. That test set has been spent; the next honest number needs fresh rows. ## Common muddles **"The test set is unbiased because we never trained on it."** Fitting is not the only channel by which data influences a model. Any choice conditioned on those rows - including the choice to keep iterating until the number looks good - leaks. **"We used cross-validation, so we do not need a test set."** Cross-validation replaces the *validation* block; it is a better way to spend scarce rows on the selection question. It does not produce an untouched partition, because the selection was driven by exactly those folds. **"We will report the validation score."** The validation score is a selection signal, not an estimate of generalisation, and the gap between them widens with the number of candidates compared. **"Feature selection happened before the split, so it does not count."** Choosing which columns to keep by looking at the whole table is a decision made on the test rows, and it spends them just as surely as tuning does. ## When the three-way split is not the right shape With a few hundred rows, an honest three-way split leaves partitions so small that both the selection signal and the final estimate are dominated by noise; the usual answer is to hold out a test set and run cross-validation over everything else, which recovers the selection signal without a dedicated validation block. With millions of rows the opposite pressure applies: the partitions can be small percentages and still be enormous, and a single hold-out is both precise and cheap. What never changes is that the number you publish must come from rows that drove no decision.
- If you tune on a single hold-out and report its best score, how wrong is that number?Optimistically biased, and the bias grows with how many configurations you tried and how noisy the estimate is. With a few hundred hold-out rows and dozens of candidates, the winner's score can sit several points above its true value simply because you kept the maximum of a set of noisy numbers.
- Can cross-validation replace the validation set entirely?For the selection job, yes - scoring candidates by cross-validation over the non-test rows gives a lower-variance signal without reserving a dedicated block, which is what you do when data is scarce. It cannot replace the test set, because the winner was still chosen using every one of those rows.
- Where should the decision threshold of a classifier be chosen?On validation. A threshold is a fitted quantity like any hyperparameter: it is picked by looking at scores, so picking it on the test set converts the test set into a validation set. Choose it on validation, freeze it, then apply that exact threshold when you score the test partition.
Validation is the rehearsal where you keep changing the script; the test set is opening night. Rehearse on opening night and you no longer know how the show plays to a fresh audience.
saying these in an interview costs you the question
- Treats the test set as a second validation set
- Says a score is unbiased because no model trained on those rows
- Reports the best validation score as final performance
- Thinks a random split by itself removes selection optimism
- Iterates against the test set until the number improves