For a 28-day-ahead forecast, how do recursive and direct multi-step strategies differ?
answer
- one model looped, or many models trained
- predictions become inputs
- errors feed the next input
- long horizons cannot use short lags
- compounding traded against models to maintain
basics
~20 sRecursive forecasting trains one one-step model and feeds its own predictions back as lags, so errors compound over 28 steps. Direct forecasting trains a model per horizon using only lags of at least that horizon, so nothing is fed back.
solid answer
~50 sRecursive uses a single model that maps recent lags to the next value. To reach day 28 you predict day 1, append that prediction as if it were an observation, rebuild the lags, predict day 2, and iterate. It is sample-efficient — every row trains the one model — but from step two onward the model runs on inputs it generated itself, so bias accumulates and the input distribution drifts from training. Direct trains one model per horizon (or one multi-output model), each mapping features available at the origin to `y[t+h]`, using only lags of at least `h`. Nothing compounds and each horizon can pick its own features, but you fit and monitor many models, each sees fewer effective examples, long horizons lose the recent lags entirely, and the 28 forecasts need not form a coherent path. A middle option pools horizons in one model with the horizon itself as a feature.
go deeper
Know the two shapes: one model applied repeatedly with its own output fed back, versus one model trained per horizon that predicts that horizon in a single step.
Explain the mechanics — which lags each strategy may legally use, and why iterating means the model runs on inputs it produced rather than observed.
Show that you have felt the tradeoff in production: compounding and flattening on one side, model count, thinner data and jagged paths on the other, and a plan to evaluate both across the horizons that matter.
Own the operational consequence. Decide how many forecasting artifacts the organisation maintains, monitors and retrains, and whether a few accuracy points at the far horizon justify that ongoing cost.
## The problem A supervised forecasting model maps a feature row to one target. Asking for 28 days of forecasts from one origin means producing 28 targets, and there are two structurally different ways to get them. ## Recursive (iterated) forecasting Train one model for a horizon of one step: features are the lags and rolling aggregates available at time `t`, and the target is `y[t+1]`. Every historical row is a training example, so the model sees the maximum amount of data. To forecast further ahead, run it in a loop. Predict `yhat[T+1]`. Treat that prediction as if it were the observation for `T+1`, rebuild the lag row for `T+2` — where lag-1 is now the prediction rather than a measurement — and predict again. Repeat 28 times. The attraction is simplicity and data efficiency: one model, one training run, one artifact, one set of hyperparameters, and every horizon benefits from all the data. The cost is that the model is used outside the conditions it was trained under. It learned the mapping from *observed* lags to the next value; from step two onward at least some of its inputs are its own estimates. Two things follow. First, errors compound: an error at step one enters the input for step two, and the process is autoregressive in error as well as in signal. Second, predictions are systematically smoother and less variable than real observations — a conditional mean has lower variance than the thing it estimates — so the feature distribution drifts towards the middle of the training range, and with it the forecasts drift towards the series mean. Long recursive paths are usually flatter than reality, and that flattening is a bias, not just added noise. ## Direct forecasting Train a separate model for each horizon. The model for horizon `h` maps features available at the origin to `y[t+h]`. Its feature set may only include lags of `h` or more: for `h = 28`, lag-1 is unavailable, so the shortest usable autoregressive column is lag-28, and the row leans much more on calendar structure and long rolling aggregates. Nothing is ever fed back, so nothing compounds; each model is trained and evaluated on exactly the task it will perform. Each horizon can also specialise — short horizons weight recent lags, long horizons weight seasonal and calendar terms — which frequently beats forcing one mapping to serve both. The costs are real. You now have as many models as horizons to fit, tune, deploy, version and monitor. Each is trained on fewer *usable* rows, because long lags require long history and the early part of every series drops out. The 28 forecasts come from 28 independently fit functions, so the path can be jagged or internally inconsistent — day 14 above day 13 and day 15 for no structural reason — which matters when a human reads the curve or when the numbers feed an optimiser that assumes a smooth trajectory. ## The middle ground Several practical hybrids exist and mentioning one shows range. **Pooled direct with a horizon feature.** Stack the training data for all horizons into one table, add `h` as a column, and fit a single model. It behaves like direct — no feedback — but shares statistical strength across horizons and yields one artifact. The features must still be restricted to what is available at the origin for that `h`, which usually means building the row from origin-relative lags. **Multi-output models.** One model emitting a vector of 28 values at once, sharing all internal structure and producing a coherent path. **DirRec.** A hybrid that extends the feature set at each step while training a separate model per horizon; it appears in the literature and is worth naming, though it is rarely the pragmatic choice. **Blending.** Recursive for the near horizons where autoregression dominates, direct for the far ones where calendar structure dominates, stitched with a smooth handover. ## Choosing between them Favour recursive when the horizon is short relative to the strength of the autocorrelation, when data per series is scarce, when the one-step model is well calibrated, and when operational simplicity matters more than a few points of accuracy at the tail. Favour direct when the horizon is long, when the drivers at long range are deterministic (calendar, holidays, planned promotions) rather than autoregressive, when compounding has been observed to hurt, or when different horizons genuinely deserve different features. Long-horizon retail and demand planning problems usually land here, which is why competition solutions on that shape of data commonly train per-horizon or pooled-direct models. The empirically honest position is that neither strategy dominates in general; the reliable way to decide is to evaluate both over the horizons that matter and inspect where the error curve separates. An interviewer is usually satisfied by a candidate who names the compounding mechanism, names the availability restriction on lags at long horizons, and treats the choice as a measurable one rather than a doctrinal one.
- Why do long recursive forecast paths tend to flatten towards the series mean?Because each fed-back input is a conditional mean, which is less variable than a real observation. Iterating pushes the feature row towards the centre of the training distribution, and a model evaluated there returns central predictions. The result is a systematic flattening — a bias in the shape of the path, not merely extra noise around it.
- What is the cheapest way to get most of the benefit of direct forecasting without maintaining 28 models?Pool the horizons into one model with the horizon as a feature. Stack training rows for every horizon, add `h` as a column, and restrict each row's lags to those available at its own origin. You keep the no-feedback property and one deployable artifact, and horizons share statistical strength instead of each being fit in isolation.
- Which feature columns matter most at a 28-day horizon compared with a one-day horizon?Short lags disappear, since lag-1 through lag-27 are unobserved at the origin, so the row shifts weight onto lag-28 and longer, long trailing aggregates shifted back by the horizon, and deterministic calendar structure such as day of week, holidays and annual harmonics. A 28-day model is largely a calendar model anchored to a level estimate.
Recursive forecasting is navigating by dead reckoning, each step measured from your last estimated position; direct forecasting is taking a fresh bearing from the known starting point for every distance you care about.
saying these in an interview costs you the question
- Claims direct forecasting is always more accurate
- Feeds recursive predictions back without noting compounding
- Uses lag-1 in a 28-day-ahead direct model
- Ignores that each direct model needs separate monitoring
- Assumes independently fit horizons produce a smooth path
- Treats the choice as doctrine rather than something to measure