When would you use direct multi-step forecasting instead of feeding predictions back in recursively?
answer
- does the prediction become an input?
- one model versus one per horizon
- errors accumulate along the horizon
- direct trains on real observed history
- far-horizon direct models look flat
basics
~20 sRecursive forecasting iterates one one-step model, feeding its predictions back as inputs, so errors compound. Direct forecasting fits a separate model per horizon on real observed history: no compounding, but many models and thinner data.
solid answer
~50 sTake ICU bed occupancy forecast one to fourteen days ahead. The recursive approach trains a single one-day-ahead model and applies it repeatedly, feeding each prediction back in as the newest lag; by day fourteen almost every input is model output rather than observation, so bias and error accumulate and the uncertainty is understated. The direct approach trains fourteen models, one per horizon, each with the label at `t+h` and features computed only from genuinely observed history — no compounding, and each model can lean on whatever features actually help at its own distance. The costs are fourteen artifacts to train and operate, a trajectory that can be non-monotonic because the models never see each other, and far-horizon models with weaker signal. I would go recursive for short horizons with strong autocorrelation and a single well-behaved series, and direct once the horizon is long enough that compounding dominates.
go deeper
Know the two names and the one-line difference: recursive reuses its own predictions as inputs, direct trains one model per horizon on real history. Being able to draw the three-step recursion is enough at this level.
Explain why compounding happens mechanically — the model was trained on observed inputs and is served predicted ones — and be able to state the data and model-count costs of the direct alternative for a fourteen-step horizon.
Show that you would decide empirically: plot error against horizon for both schemes on a holdout, weigh the operational cost of many models, and have an answer for exogenous drivers whose future values are unknown.
Own the tradeoff at portfolio scale — how many models the organisation can actually operate, whether a pooled direct model with horizon as a feature is the better compromise, and what the forecast consumers need from trajectory coherence.
## The problem multi-step forecasting creates A model trained to predict one step ahead answers exactly one question. Real requirements are rarely one step: an ICU capacity planner wants bed occupancy for each of the next fourteen days, a scheduler wants each of the next 48 hours. There are three standard ways to get a whole horizon out of a supervised learner, and interviewers ask about the first two by name. ## Recursive (iterated) forecasting Train one model `f` for horizon 1: label `y[t+1]`, features from `y[1..t]`. To forecast further, apply it repeatedly and append each prediction to the history as if it were an observation. ``` y_hat[t+1] = f(y[t], y[t-1], ...) y_hat[t+2] = f(y_hat[t+1], y[t], ...) y_hat[t+3] = f(y_hat[t+2], y_hat[t+1], ...) ``` By step 14 the newest thirteen lags are all model output. Three consequences follow: - **Error compounds.** Each step's error becomes an input error at the next step, and the model was trained assuming its inputs were true observations. A small systematic bias — say the model slightly under-predicts weekends — is re-fed and amplified rather than corrected. - **The training distribution and the serving distribution diverge.** The model only ever saw real, noisy observations as inputs; at inference it receives its own smoothed predictions, which are less variable than real data. - **Uncertainty is understated** unless you simulate: a single recursive path treats predicted values as certain. What you get in exchange is real: one model to train, tune and monitor; every training row usable; a coherent trajectory; and a horizon you can extend arbitrarily without retraining. ## Direct (horizon-specific) forecasting Train a separate model `f_h` for every horizon `h` in 1..14, each on the table whose label is `y[t+h]` and whose features stop at `t`. At inference each model fires once, in parallel, and nothing feeds back. - **No compounding.** Every model is trained on the exact task it performs, with real observations as inputs both in training and at serving. - **Horizon-appropriate behaviour.** The day-1 model naturally leans on the most recent observation; the day-14 model learns that recent noise tells it little and leans on seasonal and calendar structure, so its forecasts look flatter. That flattening is honest, not a defect. - **Costs.** Fourteen models to train, store, monitor and explain. Each model gets slightly fewer usable rows (you lose `h` rows at the end of the history). And because the models are fitted independently, the resulting trajectory can wobble — day 8 above day 7 and day 9 below both — in a way that looks wrong to a human reader even when each point is individually well estimated. ## Multi-output A third option: one model that emits the whole vector `[y[t+1], ..., y[t+14]]` at once. It avoids compounding like the direct approach and shares structure across horizons like the recursive one, but forces a single feature set and a single loss across all horizons, and not every learner supports vector targets natively. ## Choosing The deciding questions are: 1. **How long is the horizon relative to the series' memory?** If the autocorrelation decays quickly, a recursive model is guessing from its own guesses very fast. Long horizons favour direct. 2. **How much does compounding actually cost here?** Measure it: run both on a holdout and plot error against horizon. Recursive error typically grows faster than direct as `h` increases; if the curves are close through the horizon you care about, take the cheaper option. 3. **What is the operational budget?** Fourteen models per series is fine for one series and a serious burden for many. This is the argument that most often decides the question in practice. 4. **Do you need future values of exogenous drivers?** Recursion needs the target's own future lags, which it manufactures. Both approaches need future values of any external driver they use, which is why drivers known in advance — calendar, scheduled events, published prices — are preferred over ones you would have to forecast yourself. ## A pragmatic middle ground A common compromise is a **pooled direct model with the horizon as a feature**: one model trained on rows for all horizons, with `h` supplied as an input column. It keeps direct's freedom from compounding and shares parameters across horizons, at the price of forcing one functional form to serve both the near and far ends. ## What a strong answer sounds like Name both schemes precisely, state that the difference is whether predictions become inputs, identify error compounding as the core cost of recursion and model count plus data thinning as the cost of direct, and then say how you would decide empirically rather than by rule.
- Why do direct models at long horizons often produce nearly flat forecasts?Because at fourteen days out the most recent observations carry little information, so the fitted model leans on seasonal and calendar structure and shrinks the rest toward the conditional mean. That flattening is the correct response to weak signal, not a bug — the check is whether it still beats a naive rule at that horizon.
- Where does a single model that predicts all fourteen horizons at once fit in?That is the multi-output approach: one model, one shared representation, a vector label. It avoids compounding like the direct scheme and shares parameters like the recursive one, but forces a single feature set and a single loss across every horizon, and not all learners accept vector targets.
- What breaks if a recursive model depends on an external driver you do not know in advance?You must forecast that driver too, so its error enters on top of the target's own compounding. Direct forecasting removes only the need to manufacture the target's future lags — it still needs the driver's future values. That is why drivers known ahead of time, such as calendars and scheduled events, are strongly preferred.
saying these in an interview costs you the question
- Thinks recursive and direct give identical accuracy
- Misses that recursion feeds predictions in as inputs
- Claims direct forecasting needs the target's future values
- Uses recursion at long horizons without measuring error growth
- Reads a wobbly direct trajectory as proof the models are broken