skip to content

In a 14-day-ahead backtest, why put a 14-day gap between training and test data?

level: seniorimportance: should knowfreq 44%

answer

  1. the label has duration, not just the features
  2. what was knowable at the origin
  3. the last rows reach into the test window
  4. purging and embargo
  5. widen the gap by the reporting lag

basics

~20 s

Because a row dated within 14 days of the forecast origin has a target that lands at or after the origin, so it was not yet observed there. Dropping those rows keeps training to outcomes genuinely known at prediction time.

solid answer

~50 s

The gap exists because the *target* has a time extent, not just the features. If the row at time t carries the value at t+14, or the total over t+1 to t+14, then knowing that label requires data through t+14. Standing at origin T, you can only know labels for rows with t no later than T-14, so the last 14 days of would-be training rows must be dropped — that is the gap, also called a purge or embargo. Train on them and you have trained on outcomes overlapping the very window you are about to score. A gap also separates the boundary rows, which are strongly correlated with the first test points and would otherwise flatter the score. If data arrives with a reporting lag, widen the gap by that lag. No gap is needed when every training label is genuinely observed at the origin, as in a plain one-step-ahead fit.

go deeper

for a junior

Know that a forward-dated target is not observed at the row's own timestamp, so the newest training rows may have to be dropped before a backtest is trustworthy.

for a middle

Be able to derive the gap rather than recite it: work out the first date each training label was knowable, compare it to the forecast origin, and show that a horizon-h target removes the last h rows.

for a senior

Demonstrate that you audit real availability — reporting lags, revised series, late-settling records — and reconstruct what was known at the origin rather than what is known today. Expect to explain when a gap is unnecessary.

for a principal

Make the availability contract explicit across teams: what data exists at prediction time, with what lag and what revision behaviour. Backtests should be built against that contract so reported accuracy and deployed accuracy do not diverge.

## The problem: labels have duration Most discussions of time-ordered evaluation stop at "train before the cutoff, test after it." That is necessary but not always sufficient, because the label attached to a training row is often *not* observed at that row's timestamp. Consider a supervised setup for a 14-day-ahead forecast. Row t holds features computed from information available at time t, and its label is the outcome 14 days later — either the value at t+14 for a direct h-step model, or the total over t+1 to t+14 for a cumulative target. Either way, the label is only knowable once t+14 has happened. Now stand at forecast origin T, the moment you would actually make the prediction in production. Which rows may be in the training set? Only those whose labels are already observed at T, that is rows with t + 14 <= T, or t <= T - 14. Rows dated between T-13 and T have labels that fall at or after T — outcomes that have not yet occurred at prediction time and which overlap the test window you are about to score. Including them is look-ahead, and it is created purely by the split, independently of how the features were built. So the training set ends 14 days before the origin, the test window begins after the origin, and between them sits a gap of the horizon length. The gap is also called **purging** — removing training rows whose label windows overlap the test period — and **embargo** when an additional buffer is placed after the test window before training resumes at the next origin. ## The second reason: boundary correlation Even when labels are technically observed, the rows immediately adjacent to the cutoff are the most similar to the first test points: nearly the same level, the same weekday context, often the same promotion or weather. Fitting on them lets the model approach the test window from arm's length rather than from genuine distance. A gap removes the most-correlated rows and makes the reported error closer to what you will see in production, where the newest usable data is already stale by the time it is processed. ## When the gap is unnecessary The gap is not a ritual. It follows from the target definition and the data-availability schedule: - **Plain one-step-ahead fitting on the series itself.** If you estimate a model on observations up to T and forecast T+1, every value used in the fit is already realised at T. There is no unobserved label and no gap is required. - **Recursive multi-step forecasting.** The model is still trained one step ahead and then iterated forward, so training labels are all observed at T. The horizon is produced by iteration, not by a forward-dated target, so again no gap. - **Direct multi-step or window-aggregate targets.** Here the gap of the horizon length is required, because the target explicitly reaches forward. The test to apply is mechanical: for each training row, ask what date you would first have known its label. If that date is on or after the origin, the row must go. ## Data-availability lag Production rarely has data up to the instant of prediction. Sales may settle for three days, returns may be reconciled weekly, a third-party feed may arrive with a delay. If the freshest reliable observation at origin T is from T-3, then your backtest must also cut features at T-3, and the label constraint compounds: training rows need t + 14 <= T - 3. Backtests that quietly assume instant data are one of the most common sources of an optimistic number that evaporates on deployment. Where possible, reconstruct what was *known* at the origin, including the vintage of any revised series, rather than what is known today. ## What the gap costs You lose horizon-many rows of training data at every origin, and the reported error goes up. Both are features, not bugs: the lost rows are the ones you would not have had, and the higher error is the honest one. If losing them materially hurts the fit, the real message is that the history is short relative to the horizon, which is itself worth reporting. ## Interview framing State the rule as a consequence rather than a recipe: training may use only labels that were observed at the forecast origin; when the target reaches h steps forward, that removes the last h rows, which is the gap. Add the availability lag, add the boundary-correlation argument as a secondary benefit, and note the case where no gap is needed so it is clear you are reasoning rather than reciting.

  • Your sales data settles three days late. How does that change the backtest?
    Both cuts move. Features at origin T may use only data through T-3, and the training-label constraint becomes t plus the horizon no later than T-3, so the gap widens by the lag. Otherwise the backtest silently assumes data you would not have had, and production error will exceed the reported figure by an amount nobody predicted.
  • When is no gap needed at all?
    When every training label is already realised at the origin and no input reaches forward — the plain one-step-ahead case, including recursive multi-step forecasting, where the model is fit one step ahead and then iterated. The gap is required by forward-dated or window-aggregate targets, so check the target definition rather than applying a gap by default.
  • Does adding a gap make the backtest pessimistic?
    It makes it lower than the no-gap number, but the no-gap number was inflated by information the model could not have had. Relative to production, a correctly sized gap is not pessimistic — it is calibrated. The only genuine cost is the horizon-many training rows dropped at each origin, which matters mainly on short histories.

It is like grading a weather forecaster on last month's predictions: you may only use outcomes that had already been recorded when the forecast was issued, not ones that settled afterwards.

saying these in an interview costs you the question

  • The test set comes after training, so leakage is impossible
  • A gap only wastes data and lowers the score
  • Look-ahead can only come from how features were built
  • The same gap applies regardless of how the target is defined
  • Assuming data is available the instant it is generated

context