In rolling-origin backtesting, how do expanding and sliding training windows differ?
answer
- one keeps the start fixed
- the other keeps the length fixed
- bias-variance, framed by stationarity
- old regimes never leave an expanding fit
- window length is itself a hyperparameter
basics
~20 sAn expanding window keeps its start fixed so the training set grows at every origin; a sliding window keeps a fixed length and drops the oldest data. Expanding uses more history, sliding adapts faster after a regime change.
solid answer
~50 sA rolling-origin backtest fixes a sequence of forecast origins. At each origin you fit on the history available then, forecast the horizon you care about, score against what actually happened, and move the origin forward. The two window policies differ in what "history available" means. An **expanding window** starts at the beginning of the series every time, so the training set grows — lower estimation variance, but old regimes keep pulling on the fit and training size is not comparable across origins. A **sliding window** keeps a fixed length, say the last two years of weekly data, so every origin is fit on an equal amount of recent data and pre-break history falls out — more adaptive, but noisier and unable to learn long seasonal patterns. A known structural break argues for sliding; a stable series with limited history argues for expanding.
go deeper
Be able to state the mechanical difference: expanding keeps the start of training fixed and grows, sliding keeps a fixed length and drops the oldest data, and both refit as the origin moves forward.
Explain the tradeoff and the loop around it — horizon, step between origins, how many origins you end up with — and why a structural break in the series pushes you toward a fixed-length window.
Show judgment on real series: pick the window from evidence rather than habit, inspect error over time instead of one average, and notice that overlapping test windows make origin errors correlated when you talk about uncertainty.
Frame it as a policy question. The window policy should match how quickly the business believes its demand process changes, and that belief should be tested, documented and revisited rather than baked silently into a pipeline.
## The backtest loop Rolling-origin backtesting replaces a single train/test cutoff with a sequence of them. Choose a set of forecast origins t1 < t2 < ... < tk. At each origin ti you: fit the model on the training data that exists at ti, produce forecasts for the horizon of interest (say the next h periods), score those forecasts against the values that actually occurred, then advance to the next origin. The output is a collection of errors — one per origin, or one per origin-and-lead-time — rather than a single number. Two knobs define the loop besides the origins themselves: the **horizon** h, which should match the decision the forecast supports, and the **step** between origins. Stepping by h gives non-overlapping test windows and roughly independent errors, at the cost of fewer origins. Stepping by one gives many more origins, but their test windows overlap heavily, so the errors are correlated and the effective sample size is far smaller than the count of origins suggests. ## Expanding window An expanding (or growing, or anchored) window fixes the *start* of training at the beginning of the series and lets the end move with the origin. Training size grows monotonically across origins. Advantages: - **Uses all available data.** With short histories this matters a great deal; a model with several parameters and seasonal structure may simply not be estimable from two years of weekly data. - **Lower estimation variance at later origins.** More data means more stable parameter estimates and typically better forecasts, all else equal. - **Can learn long-period structure.** Annual seasonality needs several years in the window before it can be estimated at all. Disadvantages: - **Old regimes never leave.** If the process changed three years ago, every fit still contains the pre-change data, biasing the model toward a world that no longer exists. - **Training size is not comparable across origins.** The first origin is fit on far less data than the last, so error differences across origins mix genuine difficulty with sample-size effects — a confound to keep in mind when reading the error-over-time plot. ## Sliding window A sliding (or rolling) window keeps a fixed length w and moves both endpoints, so old observations drop out as new ones arrive: five years of weekly data with a two-year window gives roughly three years' worth of origins, each fit on about 104 weeks. Advantages: - **Adapts to regime change.** After a permanent shift in level, price elasticity or mix, the pre-change data leaves the window within w periods and stops distorting the fit. - **Comparable fits across origins.** Every origin sees the same training size, so error differences across time are about the data, not about how much of it there was. - **Bounded cost.** Fit time does not grow with the length of the series, which matters when you refit at many origins. Disadvantages: - **Discards information.** On a stable process this is pure loss: more variance for no bias reduction. - **Cannot see long cycles.** A window shorter than the seasonal period cannot estimate that seasonality, and a window of only two or three cycles estimates it badly. ## Choosing between them The choice is a bias-variance decision framed by stationarity. Ask: is the process stable over the span of the history? If yes, keep everything — expanding. If there is a known or suspected structural break, or if the relationship between drivers and demand drifts, prefer a sliding window long enough to hold several seasonal cycles but short enough that the old regime exits. You do not have to guess. The window length is a hyperparameter, and it can be selected the same way any other is: run the backtest at several candidate lengths and compare error across origins. Two cautions. First, that selection must be confined to data before whatever final holdout you intend to report, or the holdout stops being independent. Second, compare the error *profile over time*, not just the mean: a sliding window often loses on average while winning decisively in the periods right after a break, and if breaks are what hurt the business, the mean hides the thing you care about. A middle position is common in practice: an expanding window with observation weights that decay with age, so old data contributes but contributes less. That gives some of the adaptivity of sliding without throwing history away. ## Reporting Whichever policy you choose, report the distribution of errors across origins, not only the average — the number of origins, the spread, and how error moves through time. A single averaged figure hides both the lucky quarter and the quarter where the model broke.
- How do you choose the step size between consecutive origins?Step by the horizon when you want roughly independent test windows and honest uncertainty about the mean error. Step by one when history is short and you need many origins, but remember the overlapping test windows make errors highly correlated, so the effective sample size is far below the number of origins and confidence in the average is weaker than the count suggests.
- How many origins are enough for the backtest to be informative?Enough to cover several full seasonal cycles and a variety of conditions — a quiet stretch, a peak season, at least one disruption if the history contains one. Ten origins that all fall in the same calm quarter tell you less than five spread across two years. Judge by the conditions covered, not by the count alone.
- How would you pick the sliding-window length without contaminating your final estimate?Treat it as a hyperparameter and select it on an inner set of origins that lie entirely before the window you intend to report. Compare candidate lengths on error across those origins, fix the winner, then run the reported backtest once. Selecting the length on the same origins you report makes the reported number optimistic.
An expanding window is a diary you never close; a sliding window is a two-year rolling logbook where each new page pushes the oldest one out.
saying these in an interview costs you the question
- Expanding is always better because more data is always better
- Sliding and expanding differ only in compute cost
- Ten overlapping origins give ten independent error estimates
- The horizon can be chosen for convenience rather than from the decision
- One averaged error across origins is the whole result