skip to content

Recurrent Forecasters

Feeding a numeric series to a recurrent net means choosing a window, a horizon and a per-series scaling, then earning the win over ARIMA. Interviewers ask when the neural model is worth it.

on this pageshow

questions

4

How do you window a half-hourly electricity load series into training examples for a recurrent forecaster?

level: juniorimportance: must knowfreq 72%

answer

  1. supervised pairs, not a feature table
  2. two numbers: lookback and horizon
  3. window must cover the dominant cycle
  4. last valid start is n - L - H
  5. cut the timeline before cutting windows

basics

~20 s

Slide a fixed-length input window over the series, pairing each window with the next H values as its target. Size the window to cover the dominant seasonal cycle, split chronologically first, and drop any window whose target crosses the split.

solid answer

~50 s

You pick two numbers: an input window length `L` and a forecast horizon `H`. Every training example is the pair (`series[t : t+L]`, `series[t+L : t+L+H]`), and you slide `t` forward by a stride to generate many examples from one series. For 48-hour-ahead load forecasting on half-hourly data, `H` is 96 steps and a natural `L` is 336 steps, one full week, so the window contains the daily cycle and the weekday/weekend contrast the model needs. The last valid start is `n - L - H`. Crucially you cut the timeline into train and validation periods **before** cutting windows, and discard any training example whose target extends past the cutoff — otherwise the model is fitted on the periods you are about to score it on. Unlike a tabular setup, the model reads the raw window directly, so nothing is hand-aggregated into columns.

code

python · 12 lines
python
series = [10, 12, 9, 14, 11, 13, 15, 12, 16, 14]
L, H, stride = 3, 2, 1          # input window, forecast horizon, step

examples = []
for t in range(0, len(series) - L - H + 1, stride):
    x = series[t:t + L]         # what the model reads
    y = series[t + L:t + L + H] # what it must predict
    examples.append((x, y))

print(len(examples))            # 6 examples from 10 observations
for x, y in examples[:3]:
    print(x, "->", y)

go deeper

for a junior

Be ready to state the two numbers that define an example, the input window length and the forecast horizon, and to show how sliding that window turns one series into many training rows.

for a middle

Explain how the window length follows from the seasonal period, what a larger stride buys and costs, and why the last usable start index is n minus lookback minus horizon.

for a senior

Show that you cut the timeline before cutting windows, that no training target crosses the validation cutoff, and that gaps, short series and serve-time forecast origins are handled deliberately.

for a principal

Own the tradeoff between a long window that captures yearly structure and the example count, training cost and refresh cadence it consumes across an entire forecasting platform.

## The shape of the problem A recurrent forecaster does not consume a time series; it consumes a batch of fixed-shape examples. Turning one long numeric series into those examples is called windowing, and it is where most forecasting bugs are born. Two numbers define the transformation: - **Input window length `L`** (also called the lookback or context length): how many past steps the model reads before it predicts anything. - **Horizon `H`**: how many future steps it must produce. Example `t` is then the pair `x = series[t : t+L]`, `y = series[t+L : t+L+H]`. Slide `t` by a **stride** `s` and you get `floor((n - L - H)/s) + 1` examples from `n` observations. At stride 1 the last valid start index is `n - L - H`; starting any later leaves the target incomplete. Each example is shaped `(L, channels)` on the input side — one channel if only the target series is fed in, more if you also feed known-future or static inputs — and `(H,)` on the target side. ## Choosing L `L` is a real hyperparameter, not a formality. The rule of thumb is that the window must contain at least one, preferably two or three, repetitions of the dominant cycle you expect the model to exploit. Half-hourly electricity load has a strong daily cycle (48 steps) sitting inside a weekly one (336 steps). A 96-step window gives the model two days and no way to distinguish a Sunday from a Wednesday; a 336-step window gives it the whole weekly pattern. Going much longer buys yearly structure only if you have several years of data, and it costs on three fronts: more compute per example, a longer sequential path for the recurrence to walk during training, and fewer complete examples at the start of the series. ## Stride and effective sample size Stride 1 maximises the row count, and that number is seductive: three years of half-hourly data is over 50,000 windows. But neighbouring windows share `L-1` of their `L` observations. The examples are almost perfectly correlated, and the **effective** sample size is far closer to the number of distinct seasonal cycles — about 150 weeks — than to the row count. This is why a deep forecaster on a single series overfits so readily despite an apparently huge training set, and why a larger stride often costs little accuracy while cutting epoch time proportionally. ## Splitting without leaking The single most common error is generating all windows and then splitting them randomly. Two things go wrong. First, overlapping windows put nearly the same time steps in both the training and validation sets. Second, a randomly chosen training window can have a target that lies *after* a validation window's target, so the model is literally fitted on the future it is being scored on. The correct order is: choose a cutoff time, assign periods to train and validation, and only then window each period. Any training example whose target range crosses the cutoff must be dropped — with horizon `H`, that means discarding the last `H` training starts. Note the asymmetry: a *validation* example whose input window reaches back into the training period is perfectly legitimate, because at real forecast time those past values genuinely are observed. Inputs may look back; targets may not look forward. ## Practical details that show experience - **Gaps and missing readings.** Decide explicitly whether to drop a window containing a gap or to impute and flag it with a was-missing channel. Silent interpolation teaches a smoothness the data does not have, and whatever you do in training must be reproducible at serve time. - **Many series.** With thousands of meters or SKUs, window each series independently and pool the examples. Series shorter than `L + H` produce no examples at all and need either a shorter window or a padded, masked variant. - **Batching and state.** Windows are trained as independent examples with a fresh initial hidden state each time, so batches may be shuffled freely once the chronological split has been made. Shuffling is not the leak; splitting at random is. - **Where the window starts matters for reporting.** Line the validation origins up with how the model will actually be called in production — one forecast per day at 06:00, say — rather than at every possible index, or your offline sample will not match the deployed one. ## The trade you are really making Windowing converts an ordering constraint into a data-shaping decision. A longer window means richer context and fewer, slower, more correlated examples; a shorter window means more examples that individually know less. Getting `L`, `H`, the stride and the split boundary right costs an afternoon and determines whether every number you measure afterwards means anything.

  • At stride 1 you have almost as many examples as time steps. Why is that count misleading?
    Neighbouring windows share all but one observation, so the examples are massively correlated. The effective sample size is closer to the number of independent seasonal cycles than to the row count, which is why a model can overfit badly on tens of thousands of windows drawn from three years of data. A larger stride usually costs little and cuts epoch time proportionally.
  • The recurrence forces training to walk 336 steps in order. What non-recurrent model reads the same window in parallel?
    A stack of dilated causal convolutions. Each layer convolves the entire window at once with a kernel that only looks backwards, and stacking layers with growing dilation reaches far back in time with no sequential dependency during training. You trade an in-principle unbounded memory for a fixed lookback, but every position in the window is computed in parallel, so wall-clock training time on long windows drops sharply.
  • Some meters have gaps in their readings. What do you do with a window that contains one?
    Either drop it or impute and flag it. Silently interpolating teaches the model a smoothness the data does not have; a common compromise is short-gap interpolation plus a binary was-missing channel next to the value channel, and dropping windows whose gap is long relative to the window. Whatever you choose has to be reproducible at forecast time.

Windowing is like cutting a filmstrip into clips: each training example is a fixed run of frames plus the frames that come next. Slide by one frame and you get many clips that mostly show the same scene.

saying these in an interview costs you the question

  • Generates all windows, then splits train and validation randomly
  • Picks the window length arbitrarily, ignoring the seasonal period
  • Counts heavily overlapping windows as independent examples
  • Lets a training window's target extend past the validation cutoff
  • Never defines a horizon, just feeds one long sequence

context

open as a page

Why does a recurrent forecaster that feeds its own predictions back in drift over a 14-day horizon?

level: middleimportance: should knowfreq 58%

basics

~20 s

Because every step after the first is conditioned on a predicted value rather than an observed one. Small one-step errors re-enter as inputs and accumulate across the horizon, and a one-step training loss never penalised the 14-step trajectory at all.

open as a page

How should you scale 3,000 SKU series whose volumes span four orders of magnitude for one shared forecaster?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Scale each series by its own statistics, not the pooled distribution. A global scaler squashes a two-unit-a-day SKU toward a constant while a 20,000-unit SKU dominates the loss. Fit each scale on training data only, and invert it on forecasts.

open as a page

With 120 daily observations from a single ATM, would you ship a recurrent forecaster or a classical model?

level: principalimportance: should knowfreq 44%

basics

~20 s

Ship the classical model. One short series yields roughly a hundred overlapping windows and about seventeen weekly cycles, far too little to fit thousands of recurrent weights. Deep forecasters earn their keep by pooling many related series, not on one short one.

open as a page