skip to content

Rolling-Origin Backtesting

Refitting on an expanding or sliding history and forecasting the next block, repeatedly, so the score reflects real deployment. Look-ahead features are the classic way this quietly breaks.

on this pageshow

questions

5

Why is a shuffled random train/test split invalid for evaluating a daily demand forecast?

level: juniorimportance: must knowfreq 78%

answer

  1. the arrow of time
  2. neighbouring days are near-duplicates
  3. training rows come from after the test day
  4. drift hidden on both sides of the split
  5. cutoff date, not a coin flip

basics

~20 s

A shuffled split trains on future days and tests on past ones, so the model interpolates between neighbouring dates instead of forecasting. Because adjacent days are highly correlated, the score looks excellent and says nothing about future performance.

solid answer

~40 s

Forecasting is extrapolation forward in time, and a shuffled split does not simulate that. Randomly assigning rows puts days from after the test day into the training set, so the model is filling a gap between two known neighbours rather than predicting an unknown future. Daily demand is strongly autocorrelated, so those neighbours are near-duplicates of the held-out day, and the error collapses to something you can never reproduce in production. Shuffling also hides drift: trend, price changes and new products all appear on both sides of the split, so the test distribution matches training by construction. The honest alternative is a time-ordered split — train on everything before a cutoff, score on what follows — and better still a rolling-origin backtest that repeats that cutoff at many origins.

go deeper

for a junior

Be ready to say in one sentence that a shuffled split trains on days that come after the test days, and that the honest alternative is a cutoff date with training before it and testing after it.

for a middle

Explain the mechanism, not just the rule: autocorrelation makes held-out days easy to interpolate from their neighbours, and shuffling puts the same trend and regime on both sides so drift never shows up in the score.

for a senior

Show you audit the whole pipeline for the same defect. Scaling constants, target encodings, imputation values and model selection all have to respect the cutoff, and you should be reporting error across many origins rather than one lucky window.

for a principal

Own the standard: define once, for the whole organisation, what an acceptable forecast evaluation looks like, and make the reported number the one that matches how the model is deployed. Numbers produced by shuffled splits should not be allowed into planning decisions.

## What a split is actually simulating An evaluation split is a simulation of a decision you will make later. For a demand forecast the decision is: standing at some date, with only the data that exists on that date, predict what happens next. Every property of the split should mirror that situation. If the split lets information from after the prediction date into the fit, the number it produces answers a question nobody will ever ask. A shuffled random split breaks the mirror in the most direct way possible. Rows are assigned to train or test by coin flip, so roughly 80 percent of the days *after* any held-out day end up in training. The model is not forecasting; it is interpolating between two observed neighbours. ## Why the number is not just wrong but flattering Daily demand series are strongly autocorrelated: today looks a lot like yesterday and a lot like the same weekday last week. Under a shuffled split, a held-out Tuesday almost certainly has the Monday before it and the Wednesday after it in the training set. Predicting the middle of a sandwich is close to trivial, and the model can score well while having learned nothing about the future. The effect compounds with three other properties of real series: - **Trend and level shifts.** A series that grows over three years has a very different level in year three than in year one. Shuffling spreads all three years across both sides, so the training data always covers the test data's level. A time-ordered split makes the model extrapolate a level it has never seen — which is exactly what production does. - **Regime changes.** A pricing change, a promotion policy, a new competitor or a supply disruption changes the data-generating process partway through. Shuffled evaluation trains on the post-change regime while pretending to predict it. - **Structural repeats.** Holidays, promotions and stock-outs often affect a run of consecutive days. Shuffling splits such a run across train and test, letting the model see part of an event it is supposedly predicting. ## What a valid split looks like The minimum fix is a **time-ordered holdout**: pick a cutoff date, fit on everything strictly before it, and score on the window after it. Nothing from on or after the cutoff may influence the fit — not the model coefficients, not the chosen model, not scaling constants estimated from the data. The better fix is a **rolling-origin backtest**. Instead of one cutoff you use many: at each origin, fit on the history available at that origin, forecast the horizon you actually care about, score those forecasts against what really happened, then move the origin forward and repeat. This gives you many independent-ish estimates of forecast error rather than one, so you can see whether the model was merely lucky in one quarter, and you can look at how error behaves over time instead of collapsing everything into a single average. ## Common objections, answered *"The model never sees the test labels, so it cannot cheat."* Not seeing a label is not the same as not using future information. Training on days that occur after the test day is future information regardless of which labels are visible. *"I fixed the random seed, so the result is reproducible."* Reproducibility is not validity. A seed makes a meaningless number stable. *"I will just make the test set bigger."* Size does not repair the ordering. A larger shuffled test set gives a more precise estimate of the wrong quantity. *"My series has no autocorrelation, so shuffling is harmless."* If a series really has no autocorrelation, no trend and no seasonality, there is very little to forecast beyond its mean, and the question of split design is moot. In practice the claim is almost never true of demand data — and even then, a time-ordered split costs nothing, so there is no reason to take the risk. ## How to talk about it in an interview Say what the split is simulating, say that shuffling breaks the simulation by training on the future, then name the concrete mechanism — autocorrelation turns forecasting into interpolation, and drift is hidden because both sides of the split cover the same period. Finish by naming the fix: a time-ordered cutoff at minimum, a rolling-origin backtest when you want an error estimate you can trust.

  • Is shuffling acceptable if the series shows no visible trend or seasonality?
    No. Absence of visible structure is not absence of dependence, and a flat-looking series can still shift regime later. More importantly, a time-ordered split costs you nothing when the data really is exchangeable, so there is no upside to shuffling. If a series truly had no time structure at all, there would be almost nothing to forecast beyond its mean.
  • How long should the held-out window be for a daily series?
    Long enough to cover at least one full seasonal cycle and several repetitions of the horizon you care about, so weekly and annual effects are represented rather than one unusual stretch. If the history is too short for that, use many rolling origins instead of one long holdout, and report how error varies across origins rather than a single number.
  • What about scaling or normalisation constants — do they need the same discipline?
    Yes. Any constant estimated from data is part of the fit. Computing a mean or a scale over the whole series, including the test window, leaks future information into training even when the split itself is time-ordered. Estimate such constants from the training portion available at the origin and apply them unchanged to the test window.

Shuffling is like grading a student on fill-in-the-blank questions taken from a book they have already read cover to cover, then claiming the score predicts how they will handle next year's news.

saying these in an interview costs you the question

  • Shuffling is fine as long as the random seed is fixed
  • The model cannot cheat because test labels are hidden
  • Autocorrelation affects fitting but not evaluation
  • A bigger test set fixes a shuffled split
  • Time order only matters for models that use lagged inputs

context

open as a page

In rolling-origin backtesting, how do expanding and sliding training windows differ?

level: middleimportance: must knowfreq 66%

basics

~20 s

An expanding window keeps its start fixed so the training set grows at every origin; a sliding window keeps a fixed length and drops the oldest data. Expanding uses more history, sliding adapts faster after a regime change.

open as a page

In a 14-day-ahead backtest, why put a 14-day gap between training and test data?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because a row dated within 14 days of the forecast origin has a target that lands at or after the origin, so it was not yet observed there. Dropping those rows keeps training to outcomes genuinely known at prediction time.

open as a page

Is a last-year holdout a valid backtest if you tuned the model on the full history?

level: seniorimportance: should knowfreq 50%

basics

~20 s

No. If the last year influenced which model or hyperparameters you picked, its error is an optimistic in-sample number for that choice, not an independent estimate. Selection must happen on origins entirely before the holdout begins.

open as a page

How should a rolling-origin backtest reflect how often the model is refit in production?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The backtest should refit on the same schedule the deployed system uses. Refitting at every origin while production retrains quarterly reports the accuracy of a model far fresher than the one that will actually be serving forecasts.

open as a page