skip to content

Why does a 5-day-ahead return label force a gap between train and validation folds?

level: middleimportance: should knowfreq 40%

answer

  1. a row occupies a span, not a date
  2. when is the label actually known?
  3. the last training rows reach forward
  4. purge, then embargo
  5. gap at least the label horizon

basics

~20 s

A 5-day-ahead label dated day t is computed from outcomes through day t+5, so the last training rows overlap the validation window and share its outcomes. Dropping those final five days of training removes the overlap.

solid answer

~40 s

A labelled row is not a point in time, it is an interval. With a 5-trading-day-ahead return label, a row dated 31 March carries a label resolved from prices through 5 April. If validation starts on 1 April, that training row's label was computed from the validation period, so training and validation share outcome data and their labels are correlated by construction. The split looks correct - everything in training precedes everything in validation by feature date - and it still leaks, in the optimistic direction. The fix is to purge the final rows of the training block whose label windows reach into validation: with a 5-day label, the last five trading days. Where training resumes after a validation block, an embargo drops a further short stretch of rows immediately following it.

go deeper

for a junior

Recall that a label can be defined over future time, so a row is only fully known days after its features are. Be able to say why that makes a plain date cutoff insufficient.

for a middle

Explain the mechanism: which training rows overlap the validation window, why their labels are correlated with validation labels, and that the gap must be at least as wide as the label span.

for a senior

Show you would assert the boundary in code - maximum training label-end date before minimum validation feature date - and that you expect the score to fall once the gap is in, without treating that drop as a regression.

for a principal

Own the tradeoff between how much history purging costs across many folds and how much optimism it removes, and set the label-span and gap definition as a reviewed convention rather than a per-experiment choice.

## A labelled row spans an interval, not a point Most cross-validation intuition treats a row as living at a single moment: the timestamp of its features. With a forward-looking label that intuition is wrong, and it is the root of this whole failure. Take a row dated day `t` whose label is "the return over the next five trading days". Its features are known at `t`. Its label is not known until `t+5`. The row therefore occupies the interval `[t, t+5]` on the timeline. Any split that files it under `t` alone has misplaced part of it. ## How the overlap becomes leakage Suppose forward-chaining folds: train on everything up to 31 March, validate on April. The last training row is dated 31 March; its label is the return through roughly 5 April, computed from prices that sit *inside* the validation window. So the training set encodes what happened in the first days of April. Worse, the first few validation rows have labels built from those same days. Two rows on opposite sides of the boundary now have labels driven by one shared stretch of price movement - they are correlated by construction, not by any pattern the model discovered. A model that fits the training labels closely near the boundary will appear to anticipate the start of the validation window. The damage is concentrated at the fold boundary and always points the same way: the estimate is optimistic. It also survives every other precaution. You can order the data perfectly, refit every transform inside the fold, and hold out a clean future block - and this leak remains, because it originates in the *labels*, not the features. ## Purge and embargo Two repairs, from the financial machine-learning literature, are usually named together: **Purging** removes from the training set every row whose label window overlaps the validation window. With a 5-trading-day label and validation starting 1 April, that is the last five trading days of March. Those rows are discarded for that fold. That is the cost, and it is real: with many folds and a long label span you can lose a noticeable slice of training data. **Embargo** applies when training data continues *after* the validation block, which happens whenever the validation window sits in the middle of the timeline. Rows immediately following the validation block are still entangled with it - features built on lookback windows summarise days inside it, and errors are serially correlated - so a short additional buffer is dropped after the block before training resumes. When every fold trains strictly on the past, purging is what does the work, and practitioners often just call the resulting gap "the embargo". ## How wide should the gap be? Start from the label: the gap must be at least the span the label looks forward, five trading days here. Then widen it for anything else that couples the two sides - a target defined over an event window that can run long, a feature whose value is only finalised after a settlement delay, or strong serial correlation in the series that makes adjacent days near-copies of each other. Business-day versus calendar-day arithmetic matters: five trading days is a week or more of wall-clock time, and a gap measured in calendar days can silently be too short. ## Verifying it The check is mechanical, and worth writing as an assertion rather than trusting by eye: the maximum label-end date in the training fold must be strictly earlier than the minimum feature date in the validation fold. If your data has no explicit label-end column, derive one when the labels are built - a row that only records `t` has already thrown away the information the check needs. ## Reading the result When the gap goes in, the validation score drops. That is the honest direction, and the most common mistake at this point is to conclude the gap "broke the model" and quietly remove it. Nothing about the model changed; you removed a subsidy. The purged number is the one that has a chance of matching production, because in production you genuinely will not know 5 April's prices when you score 31 March. A second, subtler consequence: purging makes the estimate slightly pessimistic relative to a model refit on all available history at deployment time, since each fold trains on a little less data. That bias is small, it shrinks as history grows, and it errs in the safe direction - unlike the overlap it replaces.

  • How wide should the gap be, and what widens it beyond the label span?
    At minimum the label span itself. Widen it when the label is defined over an event window that can run longer than nominal, when a field is only finalised after a reporting or settlement delay, or when the series is strongly autocorrelated so adjacent rows are near-copies. Count in the same units the label uses - five trading days is more than five calendar days.
  • Does purging bias the performance estimate, and in which direction?
    Slightly pessimistic. Each fold trains on a little less data than the deployed model will, so measured performance is a touch below what a full refit achieves. That bias shrinks as history grows and points the safe way, unlike the optimistic bias from overlapping labels that it replaces.
  • Why does an embargo drop rows after the validation block, not just before it?
    Because information flows both ways across the boundary once training resumes on later data. Rows just after the block have features summarising days inside it and errors correlated with it, so training on them lets the model learn the validation period from the other side. The embargo is that buffer.

Grading an exam where the last question's answer is printed on the first page of the next exam. The papers are in the right order; the content still overlaps.

saying these in an interview costs you the question

  • Believes a time-ordered split alone removes every overlap
  • Assumes the label is known at the feature timestamp
  • Chooses a gap width unrelated to the label horizon
  • Removes the gap because the score dropped
  • Measures a trading-day label span in calendar days

context