skip to content

How do you turn a raw time series into a supervised table of lag and rolling-window features?

level: juniorimportance: must knowfreq 78%

answer

  1. reframe the series as tabular regression
  2. one row per forecast origin
  3. features come from the past only
  4. lags plus trailing-window statistics
  5. no window may end after t

basics

~20 s

Build one row per time step whose features use only past values — lags such as 1, 7 and 14 steps back, plus rolling means or spreads over trailing windows — and whose label is the value at the forecast horizon.

solid answer

~50 s

I fix the horizon first, say one day ahead, because that decides what the label is. Each row is then a forecast origin `t`: the label is the value at `t+h`, and every feature must be computable from data at or before `t`. The usual features are lags (the value at `t`, `t-1`, `t-7`), rolling statistics over trailing windows (mean and standard deviation of the last 7 and 28 observations), differences between the current value and a lag to capture momentum, and calendar attributes of the target timestamp, which are known in advance. The discipline that matters is that no window may reach past `t`: a window centred on `t`, or a mean computed once over the whole series, leaks the future and produces offline scores that collapse in production. The earliest rows have undefined long lags and are normally dropped.

code

python · 9 lines
python
series = [12, 15, 14, 18, 21, 19, 25, 27]
rows = []
for t in range(3, len(series) - 1):          # need 3 lags, and a next-step label
    lag1, lag2, lag3 = series[t], series[t - 1], series[t - 2]
    roll3 = sum(series[t - 2:t + 1]) / 3      # trailing mean, ends at t
    rows.append(((lag1, lag2, lag3, round(roll3, 2)), series[t + 1]))

for features, label in rows:
    print(features, "->", label)

go deeper

for a junior

Be ready to state the construction out loud: one row per time point, features from the past, label from the future, horizon fixed first. Knowing what a lag and a trailing rolling mean are is the whole ask here.

for a middle

Explain the mechanics precisely — which lags exist for a horizon of h, why a centred window leaks, and why the earliest rows are dropped. Justify your chosen lag set from the series' periodicity rather than reciting 1, 7, 28.

for a senior

Show the discipline of auditing every feature against its origin, and talk about what the table looks like in a real pipeline: series boundaries, exogenous drivers known versus unknown in advance, and how you would prove no feature reaches past t.

for a principal

Own the framing as a contract the whole team codes against. Decide where feature construction lives so training and serving cannot drift apart, and what review rule catches a leaking window before it reaches a model.

## The reframing A time series is an ordered sequence of observations `y[1], y[2], ... , y[T]` — daily units sold, hourly load, weekly bed occupancy. A general-purpose regression learner does not understand order; it consumes a flat table of rows, each an independent feature vector with a label. Making forecasting tractable with ordinary regression models therefore means manufacturing that table yourself, and every design decision in it is about **what a forecaster standing at one moment in time could legitimately have known**. ## Origin, horizon, label Two quantities define the task before any feature exists. - The **forecast origin** `t` is the moment you stand at. Everything up to and including `t` is observed history. - The **horizon** `h` is how far ahead you must predict. A one-day-ahead daily forecast has `h = 1`; a 48-hour-ahead hourly forecast has `h = 48`. One training row is then: *label* = `y[t+h]`, *features* = functions of `y[1..t]` (plus anything else known at `t`). Sliding `t` forward one step at a time over the history produces the table. This is often called the sliding-window or rolling-origin construction, and it is why one series of length 1,000 can yield close to 1,000 training rows. ## Lag features A lag feature is simply an earlier observation copied onto the current row: `lag_1 = y[t]`, `lag_2 = y[t-1]`, `lag_7 = y[t-6]` and so on. Lags are the primary carriers of signal, and which ones you include should follow the structure of the series rather than habit: - **Short lags** (1–3) carry persistence and momentum: tomorrow usually resembles today. - **Seasonal lags** carry the repeating cycle: lag 7 for a weekly pattern on daily data, lag 24 for a daily pattern on hourly data, lag 364 or 365 for an annual pattern on daily data. - **Derived contrasts** — `y[t] - y[t-7]`, or the ratio between the two — encode "is this week above or below last week", which is often more learnable than the raw levels, especially for tree-based models, which cannot extrapolate a rising trend beyond the range of values they saw in training. Note the indexing convention: with a horizon of `h`, the *closest* usable lag is the value at `t`, which is `h` steps away from the label. There is no such thing as a feature between `t` and `t+h`. ## Rolling-window features Rolling (trailing) statistics summarise a stretch of recent history: the mean, standard deviation, minimum, maximum, or a count of zeros over the last 7, 28 or 90 observations. They smooth noise that a single lag would inject, and the spread measures give the model a handle on how volatile the series has recently been. Two rules keep them honest: 1. **The window must end at or before `t`.** A window "centred" on the target timestamp, which some smoothing utilities produce by default, includes future values and is a leak. 2. **The window must be recomputed per row**, not once over the whole series. A single global mean or standard deviation computed across the entire history is a statistic of the future as well as the past. ## Everything else that is allowed Calendar attributes of the **target** timestamp are safe even though they lie in the future, because they are known in advance: day of week, month, whether it is a public holiday, whether it is the last day of the month. So is any exogenous driver whose future value is genuinely known — a scheduled promotion, a published price, a planned closure. An exogenous driver whose future value is *not* known (tomorrow's weather, tomorrow's traffic) can only be used as a lag or as a forecast of its own, and in the latter case its error enters yours. ## Boundary rows and multiple series The first rows of the table have undefined long lags — there is no 365-day lag on day 30. Dropping them is usually cleaner than imputing, because filling with zeros or the series mean invents history the model then learns from. If the table stacks many series on top of one another, every lag and window must be computed **within** a series; a naive shift down a stacked table pulls the tail of one series into the head of the next and quietly corrupts thousands of rows. ## Why this framing is worth getting right Almost every catastrophic offline-to-production gap in forecasting traces back to one row of this table containing something its origin could not have known. The table is the contract: if you can point at any feature and say exactly which timestamps it was computed from, and all of them are at or before `t`, the framing is sound.

  • Why are the earliest rows of a lag table usually dropped rather than filled?
    Their longest lags and widest windows have no history to compute from. Filling them with zeros or the series mean invents observations the model then learns a pattern from, and those rows carry no real signal anyway. Dropping is cleaner unless the series is so short that losing the head costs you most of the data, in which case shorten the longest lag instead.
  • How do you decide which lags to include?
    Start from the known periodicity of the series — 24 for hourly data with a daily cycle, 7 for daily data with a weekly cycle — and add short lags for persistence. Autocorrelation at each lag is a useful screen. Then add contrasts such as the current value minus the same lag, because tree-based models cannot extrapolate a trend from raw levels.
  • What changes when the table stacks many series together?
    Every lag and rolling window must be computed within a series, respecting its own ordering. A shift applied to the stacked table pulls the end of one series into the beginning of the next, producing features that mix unrelated histories. The rows at each series' head then need dropping individually, not once globally.

Each row is a flashcard showing only what a forecaster could have seen that morning, with the eventual outcome written on the back.

saying these in an interview costs you the question

  • Centres a rolling window on the target timestamp
  • Computes rolling statistics once over the entire series
  • Uses a feature only knowable after the forecast origin
  • Fills undefined early lags with zeros without thinking
  • Shifts lags across a stacked table ignoring series boundaries

context