How should a rolling-origin backtest reflect how often the model is refit in production?
answer
- the backtest simulates a policy
- the served model is older than the tested one
- data rolls forward, parameters do not
- price the staleness, do not guess it
- where the gap concentrates matters most
basics
~20 sThe backtest should refit on the same schedule the deployed system uses. Refitting at every origin while production retrains quarterly reports the accuracy of a model far fresher than the one that will actually be serving forecasts.
solid answer
~50 sA backtest is a simulation of a deployed policy, and refit cadence is part of that policy. If the backtest refits at every weekly origin but production retrains once a quarter, the reported error belongs to a model that is always fresh, while the served model spends most of its life stale — the gap shows up as unexplained degradation after launch. Simulate the real schedule: refit only at the cadence you will actually use, and between refits keep parameters frozen while rolling the data forward, which is what a deployed system does. To choose the cadence, run the same origins twice — refit-every-origin and refit-quarterly — and read the difference as the price of staleness, then weigh it against compute, review and deployment-risk costs. Look at where the gap concentrates: if it appears only around regime shifts, the answer may be event-triggered retraining rather than a faster calendar.
go deeper
Know that a model does not update itself: between scheduled refits its parameters are fixed, even though it keeps consuming new data.
Be able to explain how a backtest imitates a slower refit schedule — roll the data forward at every origin but re-estimate parameters only on the deployed cadence — and why doing otherwise inflates the score.
Show you price staleness empirically by running the same origins under two cadences, and that you read the error profile over time and the worst origin rather than only the average.
Own the cadence as a policy decision balancing accuracy, compute, review overhead and the disruption of a forecast that changes under planners, and require every published backtest to state the cadence it assumed.
## The backtest simulates a policy, not a model It is easy to think of a backtest as measuring a model. It is more accurate to say it measures a *policy*: a rule for how a model is built, refreshed and served over time. Refit cadence is a first-class part of that policy, alongside the training window and the horizon, and if the backtest and production disagree about it, the backtest is measuring a system nobody is going to run. The usual mismatch runs one way. Backtests refit at every origin because that is the simplest loop to write. Production retrains far less often, because retraining involves compute, a validation step, sometimes human review or an approval gate, and a deployment with its own risk. The consequence is a systematically optimistic report: the backtest never serves a model older than one step, while production routinely serves one that is weeks or months old. ## What "stale" actually means Between refits, two things can move forward independently: - **The data the model consumes.** A deployed forecaster keeps ingesting new observations, so its inputs stay current even when its parameters do not. - **The parameters themselves.** These change only when you refit. Some model families make the distinction sharp: state can be updated with each new observation while the estimated parameters stay fixed until the next scheduled refit. A faithful backtest imitates exactly that — roll the data forward at every origin, refit the parameters only on the cadence you intend to deploy. The cost of staleness is the accuracy you give up by serving parameters estimated from older data. It depends on how fast the process drifts. On a stable series it can be near zero for months; through a demand shift it can be large and concentrated in a short period. ## Measuring the cost rather than guessing it The decision does not have to be intuition. Run the identical set of origins under two cadences — refit at every origin, and refit on the candidate schedule — and compare origin by origin. That gives: - **The average gap**, which is the headline price of the slower cadence. - **The shape of the gap over time**, which is usually more informative. A gap concentrated in the weeks following a structural break tells a different story from a gap spread evenly, and points to event-triggered retraining rather than a faster calendar. - **The worst-origin gap**, which matters when the business cost of a bad forecast is not linear — a stock-out or an overcommitment during peak season is not compensated by an accurate quiet week. Against that, weigh the costs of frequent refitting: compute and orchestration, the human time in any review gate, the deployment risk of a model changing under people who plan against it, and the loss of a stable reference point when forecasts change for reasons users cannot see. ## Making the call A reasonable decision procedure: 1. Measure the staleness cost at two or three candidate cadences on the same origins. 2. If the slowest acceptable cadence loses little on average and little at the worst origin, take it and spend the saved effort elsewhere. 3. If the loss concentrates around regime shifts, keep the slow calendar cadence but add a trigger: retrain when monitored error or input distribution moves beyond a threshold. 4. Whatever you choose, make the backtest use that exact schedule from then on, so the reported accuracy is the accuracy of the thing being deployed. ## Second-order interactions Cadence interacts with the training window. A sliding window plus infrequent refitting compounds staleness — the model is both forgetting old data and not learning new data. An expanding window plus infrequent refitting is more inert but drifts more slowly. It also interacts with organisational rhythm. If forecasts feed a quarterly planning cycle, a model that changes weekly may create more confusion than the accuracy is worth, because plans get revised for reasons stakeholders cannot attribute. Conversely, a short-horizon operational forecast that drives daily replenishment can tolerate and benefit from frequent refreshes. ## Reporting When you publish backtest accuracy, state the refit cadence alongside the horizon, the window policy and the number of origins. Those four facts define what was actually measured, and omitting the cadence is how an optimistic number reaches a planning meeting without anyone noticing the assumption behind it. ## Interview framing The strong answer treats cadence as a policy choice with a measurable price and a set of costs on the other side, insists that the backtest simulate whatever cadence is chosen, and mentions where the gap concentrates rather than only its average. A weak answer says "refit as often as possible" and never connects the backtest loop to how the system will actually run.
- The quarterly-refit backtest is only slightly worse on average. Is that enough to adopt it?Only after checking where the difference sits. Look at the worst origins and at the periods following any regime shift; an average that hides a large loss during peak season or after a demand break is misleading when bad forecasts are expensive. If the loss really is spread thin and small, take the cheaper cadence and reinvest the effort.
- How would you decide between a fixed calendar cadence and event-triggered retraining?Use the shape of the staleness cost. Evenly spread degradation argues for a calendar schedule matched to the drift rate; degradation concentrated after shifts argues for triggers on monitored error or input drift, usually with a slow calendar refit as a floor. Triggers need monitoring you trust and a guard against retraining on a transient anomaly.
- How does refit cadence interact with the training-window policy?They compound. A short sliding window with infrequent refits is the worst combination: the model has already forgotten older data and is not yet learning newer data, so it is stale in both directions. If you must refit rarely, a longer or expanding window is the safer partner, since its parameters change more slowly anyway.
Judging a quarterly-updated map by how well a continuously-updated one navigates the city tells you nothing about the detours drivers will actually take.
saying these in an interview costs you the question
- Refit as often as possible; more is always better
- Refit cadence is an engineering detail, not part of evaluation
- The average gap between cadences is the whole story
- Backtest freshness has no effect on reported accuracy
- Publishing backtest error without stating horizon or cadence