skip to content

Is a last-year holdout a valid backtest if you tuned the model on the full history?

level: seniorimportance: should knowfreq 50%

answer

  1. leakage through the choice, not the fit
  2. every comparison is a use of data
  3. the winner is a minimum over noise
  4. select before, score after, score once
  5. count how many configurations were tried

basics

~20 s

No. If the last year influenced which model or hyperparameters you picked, its error is an optimistic in-sample number for that choice, not an independent estimate. Selection must happen on origins entirely before the holdout begins.

solid answer

~50 s

Look-ahead does not only enter through the fit; it also enters through the *choice*. If you compared model orders, window lengths or hyperparameters using the whole history — including the final year — then the winner was chosen partly because it happens to do well on that year, and reporting its error there is reporting a number you optimised. The fix is to structure selection and evaluation in time: run an inner backtest over origins that lie entirely before the holdout starts, pick the configuration on those results, freeze it, then score the holdout exactly once. Also report how many configurations were compared. The best of fifty noisy scores is biased low by the maximum-of-noise effect, so a small gap to the runner-up may be selection noise rather than a real difference. If you then refit on all data for deployment, that model has no unused window left to evaluate on.

go deeper

for a junior

Know that choosing a model by comparing scores on a window counts as using that window, so the same window cannot then serve as an independent test of the winner.

for a middle

Explain the structure that fixes it: run selection on origins that end before the reserved window, freeze the configuration, then score the reserved window a single time.

for a senior

Show you handle the messy version — short histories, the maximum-of-noise bias when many candidates are compared, and the discipline of not re-consulting the holdout after seeing a disappointing result.

for a principal

Own the norm that what gets evaluated is the whole procedure, selection included, and that reports must state how many configurations were compared and how often the reserved window was consulted.

## Two channels for future information A backtest can be contaminated in two distinct ways. The first is the obvious one: training data that includes or reaches into the scored window. The second is subtler and at least as common: **selection leakage**, where the scored window influenced which model you are scoring. Every comparison you run is a use of data. Choosing a model order, a training-window length, a regularisation strength, a feature set, a transformation, even the decision to "try a different family because the first one looked bad on the last year" — all of these consume information from whatever window you looked at. If the final holdout is part of that window, its error is no longer an out-of-sample estimate of the chosen model; it is the minimum of many scores computed on that same window. ## Why the reported number is biased low Suppose you evaluate fifty configurations on the same last-year window. Each configuration's measured error is its true error plus noise from that particular year — one unusual promotion, one weather event, one supply disruption. Selecting the minimum measured error preferentially picks configurations whose noise happened to be favourable. The winner's measured error is therefore an underestimate of its true error, and the gap grows with the number of configurations compared and with the noisiness of the window. This is why a small margin between the winner and the runner-up should not be treated as a real difference. With fifty candidates and a single scored year, differences of a few percent are routinely noise. ## The structure that fixes it Separate the two jobs in time: 1. **Reserve the final window.** Decide up front which stretch of the series is the evaluation window and do not look at it — not for tuning, not for plotting model errors, not for deciding which family to try. 2. **Select on an inner backtest.** Within the pre-holdout history, run a rolling-origin backtest: several origins, each fitting on data before it and scoring the horizon after it, all of it ending before the holdout begins. Compare configurations on the error across those inner origins. 3. **Freeze.** Fix the configuration — model form, hyperparameters, window policy, refit cadence, feature set. 4. **Score once.** Run the frozen configuration through the holdout and report that number. Once. The last point is the hard one in practice. If you look at the holdout, dislike the result, change something and look again, the holdout has become a tuning set and its independence is gone. A number obtained on the third look is not an unbiased estimate of anything. ## Practical realities **Short histories.** Reserving a year for a single-use holdout is expensive when you only have three. The alternative is to make the *entire* evaluation a rolling-origin backtest with selection nested inside each origin: at each origin, choose the configuration using only data before that origin, then score it on the horizon after it. This is more expensive to run and more fiddly to implement, but it uses the history far better and it measures the whole procedure — selection included — rather than one frozen model. **Selection is part of the system.** That last point deserves emphasis. What you deploy is not a model but a procedure that periodically picks and fits a model. Evaluating the procedure, selection and all, is the more faithful simulation. A backtest that scores a hand-picked configuration answers a narrower question than the one you actually have. **Refitting for deployment.** After evaluation it is normal to refit the frozen configuration on all data, including the holdout, before deploying. That is legitimate — but understand that the deployed model has no unused window left, so the holdout number describes the configuration, not that particular fit, and it will not be reproducible against fresh data until fresh data exists. **Honest reporting.** Record how many configurations were compared, on which origins, and how many times the holdout was consulted. Those three facts let a reader judge how much to discount the headline number, and their absence is a reasonable ground for scepticism. ## Interview framing Answer no, and name the mechanism: selection is a use of data, so a window used for selection cannot also serve as an independent estimate. Then give the structure — inner backtest before the holdout, freeze, score once — and add the maximum-of-noise point, which is the part that separates a candidate who has merely heard the rule from one who understands why it exists.

  • The best of fifty configurations beats the runner-up by a small margin. How much should you trust that?
    Not much on its own. The winner's score is the minimum of fifty noisy measurements, so it is biased low, and a small gap is well within selection noise. Check whether the ordering holds across individual origins rather than only in the average, and prefer the simpler or more stable configuration when the margin is inside the origin-to-origin spread.
  • Your history is too short to spare a final year. What do you do instead?
    Nest the selection inside the rolling backtest: at each origin, choose the configuration using only data before that origin, fit it, and score the horizon that follows. You then measure the entire procedure — selection included — rather than one hand-picked model. It costs more compute and more care to implement, but it wastes no history.
  • Is it legitimate to refit the chosen configuration on all data, including the holdout, before deploying?
    Yes, and it is usually the right call, since the deployed model should use the freshest data. Just be clear about what the reported number describes: it estimates the configuration's performance, not that specific final fit, and no unused window remains to validate it. Plan to re-measure once genuinely new data has accumulated.

saying these in an interview costs you the question

  • The holdout is out-of-sample because the model never fit on it
  • Tuning does not count as using the data
  • The best of many candidates is an unbiased estimate of its error
  • Checking the holdout a few times and adjusting is harmless
  • Only the final fit needs to respect the time cutoff

context