skip to content

Why does a recurrent forecaster that feeds its own predictions back in drift over a 14-day horizon?

level: middleimportance: should knowfreq 58%

answer

  1. after step one the inputs are guesses
  2. bias does not cancel, it accumulates
  3. one-step loss, fourteen-step deployment
  4. emit all horizons from one state
  5. no feedback path means no compounding

basics

~20 s

Because every step after the first is conditioned on a predicted value rather than an observed one. Small one-step errors re-enter as inputs and accumulate across the horizon, and a one-step training loss never penalised the 14-step trajectory at all.

solid answer

~50 s

Recursive decoding predicts day 1, appends that prediction to the input, predicts day 2 from it, and repeats. From step two onward the model reads its own output, so any one-step error is baked into the next input: systematic bias accumulates in one direction and variance grows with the horizon. The objective makes it worse — if training minimised one-step error, nothing pushed the parameters to make step 14 good, and the model learns to lean on the most recent observed value, exactly the input that turns synthetic in a rollout. The alternative is a direct multi-horizon head: one decoder emitting all 14 values at once from the encoder's final state, trained on the mean loss over all 14 steps. No output is ever an input, so nothing compounds. The cost is a horizon fixed at training time and a trajectory that can come out jagged.

go deeper

for a junior

Know that a recursive forecast reuses its own predictions as inputs, and that this is why a long rollout is less trustworthy than the first step of it.

for a middle

Explain the compounding path step by step, why systematic bias accumulates instead of cancelling, and what a head that emits the whole horizon at once changes about the gradient.

for a senior

Diagnose drift from a per-horizon error curve and a flattening mean forecast, and justify a decoding choice against the horizon the business actually consumes.

for a principal

Frame it as an objective-design decision: which horizon the loss should represent, whether trajectory coherence is a product requirement, and what fixing the horizon costs in model proliferation.

## What recursive decoding actually does A recurrent forecaster trained to map a window of history to the next single value can be extended to any horizon by iteration: predict `y_hat[t+1]`, shift it into the input, predict `y_hat[t+2]` from the extended input, and repeat 14 times. It is attractive because one model with one set of weights answers any horizon, and because the trajectory is internally consistent — each step is conditioned on the path the model has already committed to. Drift is the price. ## The compounding mechanism At step 1 the model reads only observed data, so its error is the honest one-step error. At step 2 one of its inputs is `y_hat[t+1]`, which carries that error. At step 3, two inputs are model output. By step 14 the entire recent context is synthetic. Two distinct things degrade: - **Variance accumulates.** Each step adds its own noise on top of an input that already contains earlier noise, so the spread of the 14-step forecast is much wider than the one-step spread. Even a perfectly unbiased model has this. - **Bias compounds in one direction.** This is the dangerous one. If the model systematically under-predicts peaks by a little, the under-predicted value becomes the context for the next step, which is then predicted from an artificially low level, and the forecast slides steadily away from reality. Errors do not cancel because they are correlated with the state that produced them. The classic visible symptom is a rollout that flattens into a near-constant line: the model's small pull toward the mean is applied 14 times in a row. ## The objective mismatch The deeper issue is that a one-step objective and a 14-step deployment optimise different things. Under one-step loss, the single most informative input is the most recent observation, and a strong model learns to rely on it heavily — nearly a persistence forecast with corrections. That is a great one-step strategy and a terrible 14-step one, because after step 1 the most recent input is the model's own guess. This is why an excellent one-step validation number can sit next to a poor 14-day backtest, and why one-step error must never be used to select a model that will be rolled out. ## The direct multi-horizon head The structural alternative keeps the recurrent encoder over the input window but replaces iterative decoding with a head that produces all `H` values in one shot from the encoder's final state — a sequence-to-sequence arrangement whose output side emits the whole horizon simultaneously. The loss is the mean error over all `H` positions, so the gradient directly rewards being right 14 steps out. What you gain: - No feedback path at all, so no compounding. Step 14's error is whatever the model's step-14 estimate is worth, not the accumulation of thirteen previous mistakes. - The long-horizon error appears in the objective, so the model can trade one-step sharpness for horizon-wide accuracy. - Inference is a single pass instead of `H` sequential passes, which matters when you serve thousands of series. What you give up: - **The horizon is fixed.** `H` output units are trained for exactly 14 steps. Serving 7 or 21 needs another head or another model, whereas recursion answers any horizon with the same parameters. - **Capacity scales with the horizon.** A long horizon means a wide output layer, and each position gets its own share of a fixed training set. - **No trajectory coherence.** Each output unit is free to move independently, so the 14-day path can be jagged or its aggregate inconsistent. If downstream consumers need a coherent path or a correct total, add a smoothness or aggregation term to the loss or post-process — the architecture alone guarantees nothing. ## Choosing between them Short horizons relative to the sampling rate favour recursion: there is little room for compounding, and the flexibility is free. Long horizons, or horizons where the business cares about the far end (a 14-day replenishment decision hinges on days 10 to 14, not day 1), favour the direct head. A hybrid many teams land on is a direct head for the fixed operational horizon plus a recursive model kept for ad-hoc queries. Whichever you choose, evaluate at the horizon you serve, and evaluate **per step**. A single averaged number hides exactly the pattern that diagnoses drift: a per-horizon error curve that rises steeply and a mean forecast that flattens toward the series average tells you the rollout is decaying, not that the model is uniformly weak.

  • Your one-step validation error is excellent but the 14-day backtest is poor. What does that tell you?
    That the objective and the deployment horizon disagree. A model tuned on one-step error learns to lean on the most recent observation, which is precisely the input that becomes synthetic after step one in a rollout. Select and compare models on the horizon you actually serve, scored per step, before you conclude anything about architecture.
  • Does emitting all 14 values from one head guarantee a smooth, coherent trajectory?
    No. Each output unit moves independently, so a direct head can produce a jagged path across horizons and totals that do not reconcile with any sensible trajectory. If consumers need coherence, add a smoothness or aggregation term to the loss, or post-process the output. Removing the feedback loop removes compounding, not inconsistency.
  • When is recursive decoding still the better choice?
    When the horizon varies at serve time, or is short enough that compounding has little room to grow. One recursive model answers any horizon with a fixed parameter count, while a direct head fixes the horizon at training time and spends output capacity on every extra step. For a one- or two-step-ahead service, recursion is simpler and usually just as accurate.

saying these in an interview costs you the question

  • Claims a larger hidden state fixes long-horizon drift
  • Uses one-step validation error to select a model that is rolled out
  • Assumes rollout errors cancel out on average
  • Treats drift as an implementation bug rather than an objective mismatch
  • Thinks a multi-output head removes the need to fix a horizon

context