Why is a shuffled random train/test split invalid for evaluating a daily demand forecast?
answer
- the arrow of time
- neighbouring days are near-duplicates
- training rows come from after the test day
- drift hidden on both sides of the split
- cutoff date, not a coin flip
basics
~20 sA shuffled split trains on future days and tests on past ones, so the model interpolates between neighbouring dates instead of forecasting. Because adjacent days are highly correlated, the score looks excellent and says nothing about future performance.
solid answer
~40 sForecasting is extrapolation forward in time, and a shuffled split does not simulate that. Randomly assigning rows puts days from after the test day into the training set, so the model is filling a gap between two known neighbours rather than predicting an unknown future. Daily demand is strongly autocorrelated, so those neighbours are near-duplicates of the held-out day, and the error collapses to something you can never reproduce in production. Shuffling also hides drift: trend, price changes and new products all appear on both sides of the split, so the test distribution matches training by construction. The honest alternative is a time-ordered split — train on everything before a cutoff, score on what follows — and better still a rolling-origin backtest that repeats that cutoff at many origins.
go deeper
Be ready to say in one sentence that a shuffled split trains on days that come after the test days, and that the honest alternative is a cutoff date with training before it and testing after it.
Explain the mechanism, not just the rule: autocorrelation makes held-out days easy to interpolate from their neighbours, and shuffling puts the same trend and regime on both sides so drift never shows up in the score.
Show you audit the whole pipeline for the same defect. Scaling constants, target encodings, imputation values and model selection all have to respect the cutoff, and you should be reporting error across many origins rather than one lucky window.
Own the standard: define once, for the whole organisation, what an acceptable forecast evaluation looks like, and make the reported number the one that matches how the model is deployed. Numbers produced by shuffled splits should not be allowed into planning decisions.
## What a split is actually simulating An evaluation split is a simulation of a decision you will make later. For a demand forecast the decision is: standing at some date, with only the data that exists on that date, predict what happens next. Every property of the split should mirror that situation. If the split lets information from after the prediction date into the fit, the number it produces answers a question nobody will ever ask. A shuffled random split breaks the mirror in the most direct way possible. Rows are assigned to train or test by coin flip, so roughly 80 percent of the days *after* any held-out day end up in training. The model is not forecasting; it is interpolating between two observed neighbours. ## Why the number is not just wrong but flattering Daily demand series are strongly autocorrelated: today looks a lot like yesterday and a lot like the same weekday last week. Under a shuffled split, a held-out Tuesday almost certainly has the Monday before it and the Wednesday after it in the training set. Predicting the middle of a sandwich is close to trivial, and the model can score well while having learned nothing about the future. The effect compounds with three other properties of real series: - **Trend and level shifts.** A series that grows over three years has a very different level in year three than in year one. Shuffling spreads all three years across both sides, so the training data always covers the test data's level. A time-ordered split makes the model extrapolate a level it has never seen — which is exactly what production does. - **Regime changes.** A pricing change, a promotion policy, a new competitor or a supply disruption changes the data-generating process partway through. Shuffled evaluation trains on the post-change regime while pretending to predict it. - **Structural repeats.** Holidays, promotions and stock-outs often affect a run of consecutive days. Shuffling splits such a run across train and test, letting the model see part of an event it is supposedly predicting. ## What a valid split looks like The minimum fix is a **time-ordered holdout**: pick a cutoff date, fit on everything strictly before it, and score on the window after it. Nothing from on or after the cutoff may influence the fit — not the model coefficients, not the chosen model, not scaling constants estimated from the data. The better fix is a **rolling-origin backtest**. Instead of one cutoff you use many: at each origin, fit on the history available at that origin, forecast the horizon you actually care about, score those forecasts against what really happened, then move the origin forward and repeat. This gives you many independent-ish estimates of forecast error rather than one, so you can see whether the model was merely lucky in one quarter, and you can look at how error behaves over time instead of collapsing everything into a single average. ## Common objections, answered *"The model never sees the test labels, so it cannot cheat."* Not seeing a label is not the same as not using future information. Training on days that occur after the test day is future information regardless of which labels are visible. *"I fixed the random seed, so the result is reproducible."* Reproducibility is not validity. A seed makes a meaningless number stable. *"I will just make the test set bigger."* Size does not repair the ordering. A larger shuffled test set gives a more precise estimate of the wrong quantity. *"My series has no autocorrelation, so shuffling is harmless."* If a series really has no autocorrelation, no trend and no seasonality, there is very little to forecast beyond its mean, and the question of split design is moot. In practice the claim is almost never true of demand data — and even then, a time-ordered split costs nothing, so there is no reason to take the risk. ## How to talk about it in an interview Say what the split is simulating, say that shuffling breaks the simulation by training on the future, then name the concrete mechanism — autocorrelation turns forecasting into interpolation, and drift is hidden because both sides of the split cover the same period. Finish by naming the fix: a time-ordered cutoff at minimum, a rolling-origin backtest when you want an error estimate you can trust.
- Is shuffling acceptable if the series shows no visible trend or seasonality?No. Absence of visible structure is not absence of dependence, and a flat-looking series can still shift regime later. More importantly, a time-ordered split costs you nothing when the data really is exchangeable, so there is no upside to shuffling. If a series truly had no time structure at all, there would be almost nothing to forecast beyond its mean.
- How long should the held-out window be for a daily series?Long enough to cover at least one full seasonal cycle and several repetitions of the horizon you care about, so weekly and annual effects are represented rather than one unusual stretch. If the history is too short for that, use many rolling origins instead of one long holdout, and report how error varies across origins rather than a single number.
- What about scaling or normalisation constants — do they need the same discipline?Yes. Any constant estimated from data is part of the fit. Computing a mean or a scale over the whole series, including the test window, leaks future information into training even when the split itself is time-ordered. Estimate such constants from the training portion available at the origin and apply them unchanged to the test window.
Shuffling is like grading a student on fill-in-the-blank questions taken from a book they have already read cover to cover, then claiming the score predicts how they will handle next year's news.
saying these in an interview costs you the question
- Shuffling is fine as long as the random seed is fixed
- The model cannot cheat because test labels are hidden
- Autocorrelation affects fitting but not evaluation
- A bigger test set fixes a shuffled split
- Time order only matters for models that use lagged inputs