skip to content

Your team has scored the same hold-out test set 40 times this quarter - what now?

level: principalimportance: should knowfreq 40%

answer

  1. each look changed the next decision
  2. selection, not fitting, did the damage
  3. it became a validation set
  4. bias grows with the comparison count
  5. budget, gatekeeper, log, refresh

basics

~20 s

That hold-out is no longer a test set: forty scored comparisons made it a validation set, and its number is optimistic by an unmeasurable amount. Get fresh untouched rows for the real verdict, and budget future use.

solid answer

~50 s

The first thing to say is diagnostic, not procedural: every one of those forty evaluations fed a decision, so the surviving model was selected partly on those rows and the score is biased upward. If the candidates had been genuinely equally good, the expected winner sits around `sqrt(2 * ln 40)` - roughly 2.7 - standard errors above the truth, so on a hold-out with a one-point standard error that is a couple of points of pure illusion. You cannot measure the bias from inside the same rows, so I stop calling it a test score and relabel it as validation. Then I source a clean estimate: rows deliberately held back and never queried, or a fresh batch nobody has touched. Going forward I would put the hold-out behind a budget - a fixed number of evaluations, one owner who runs them, a log of every run and what it decided - and refresh it on a schedule rather than when someone notices.

go deeper

for a junior

Be ready to state the rule this violates: the test set is scored once, after the model is frozen. Knowing that repeated scoring makes the number optimistic is enough at this level.

for a middle

Explain the mechanism - each evaluation informed a decision, so the surviving model was selected on those rows - and describe how the optimism scales with the number of comparisons and the noise of each one.

for a senior

Show what you actually do next: relabel the number, source uncontaminated rows, present the exposure honestly to whoever is making the launch call, and instrument the pipeline so it cannot recur silently.

for a principal

Own the policy and its costs: evaluation budgets, a gatekeeper, an audit log, and a refresh cadence - weighed against slower iteration, labelling spend and the loss of comparability with historical numbers.

## What actually happened A hold-out set does not degrade because it was *looked at*. It degrades because each look changed what the team did next: a feature was dropped, a model family was abandoned, a threshold moved, an experiment was continued rather than killed. Forty scoring events over a quarter is forty decisions conditioned on those rows. The model that survived is, in part, the model that fitted the hold-out's noise - not by gradient descent, but by natural selection over a series of human choices. This is usually called adaptive overfitting to the hold-out, and it is the same mechanism that makes a tuned validation score optimistic. The only difference is that here it happened to the partition whose whole purpose was to be uncontaminated. ## How much bias? The honest answer is that you cannot measure it from inside the same rows - if you could, you would just subtract it. But you can bound the scale. If `k` candidates were genuinely equally good and each hold-out estimate had standard error `sigma`, the expected score of the best-looking candidate sits about `sigma * sqrt(2 * ln k)` above the truth. For `k = 40` that factor is about 2.7. On a hold-out where accuracy has a standard error of one point, the winner is expected to look nearly three points better than it is - and that is the *equally good* case; real selection sequences are adaptive, which can be worse, while genuinely large quality gaps between candidates make it better. The useful takeaway for an interview is the shape: the optimism grows with how many comparisons were made and with how noisy each one was, so a small hold-out queried often is the worst combination. ## What to do now **1. Relabel, do not relitigate.** Stop publishing that number as a test result. It is a validation score. Say so in the report, with the count of evaluations behind it. The reputational cost of correcting it now is far smaller than the cost of a launch decision made on a number that was three points high. **2. Get rows nobody has queried.** In order of preference: a partition that was reserved from the beginning and genuinely never opened; a fresh batch of labelled data collected since; or, if neither exists, a re-partition of everything followed by a clean re-run of the whole selection procedure. The last one is expensive and is the reason the discipline matters. **3. Quantify the exposure while you wait.** If the decision cannot wait for clean rows, present the number with an explicit caveat and, where possible, a range: the cross-validated estimate over the training rows and the contaminated hold-out score bracket the plausible truth, and if they disagree materially that gap is itself the signal to report. **4. Do not average the forty scores.** Averaging estimates of forty *different* models tells you about the search, not about the survivor. ## Preventing the next one The organisational fix is a policy, and a lead is expected to own it: - **A budget.** Decide up front how many times the hold-out may be scored in a project - often once, occasionally a handful of times at pre-declared milestones - and treat the budget as spent when it is spent. - **A gatekeeper.** One person, or one automated job, runs test evaluations. Individual engineers get the training rows and a resampling scheme; they do not get the hold-out. - **A log.** Every evaluation records the date, the model, the score and the decision it informed. This is what lets you say "forty" instead of "a lot", and it is what lets a reader discount the number appropriately. - **A pre-declared decision rule.** Write down what score would mean ship and what would mean stop *before* running. Most hold-out abuse is the search for a run that clears a bar nobody wrote down. - **A refresh cadence.** Plan for the hold-out to be replaced periodically with newly collected data, and budget the labelling for it. A test set is a consumable, not a fixture. ## The tradeoffs a lead has to weigh Strict gatekeeping slows iteration and frustrates engineers who want a quick read on whether an idea works; the counter is to make the resampling-based estimate on training rows fast, well-tooled and trusted, so nobody *needs* the hold-out to iterate. Refreshing the test set costs labelling budget and breaks comparability with historical numbers, so you decide deliberately whether continuity or honesty matters more for a given metric. And when the data is genuinely scarce, you may have to accept a contaminated estimate and manage the risk explicitly rather than pretend a clean one exists. What is not defensible is continuing to publish a number as a test score once you know what has been done to it.

  • Can you correct the reported score for the number of times the hold-out was used?
    Not reliably from those same rows. You can bound the scale - optimism grows roughly with the square root of the log of the number of comparisons, multiplied by the estimate's standard error - and you can report that as a caveat. A trustworthy corrected number needs data that was not part of the selection.
  • Is it acceptable to score the hold-out more than once on a long project?
    Yes, if it is planned. A small number of pre-declared checkpoints, each logged with the decision it informed, keeps the optimism small and bounded. What corrodes the set is unbudgeted, ad-hoc scoring where the number of looks is unknown and the decision rule was written after the score appeared.
  • Someone proposes reshuffling the split before each evaluation to keep the test set fresh. Why is that wrong?
    Reshuffling recycles rows that have already influenced decisions back into the test partition, so the contamination spreads rather than clears. It also destroys comparability between runs. Freshness comes from rows that have never been scored, not from rearranging the ones that have.

A test set is a sealed envelope with the exam answers. Once forty people have opened it to check their work, resealing it does not make the next score honest.

saying these in an interview costs you the question

  • Says the score is still valid because no model trained on those rows
  • Averages the forty past scores into one estimate
  • Reshuffles the split to make the hold-out feel fresh
  • Treats a test set as a permanent fixture, never refreshed
  • Keeps reporting the number without disclosing how often it was used

context