skip to content

Lag Features and Horizons

Turning a series into a supervised table of lag, rolling and calendar or Fourier terms, then forecasting several steps out recursively or with one model per horizon. Leakage hides in this step.

on this pageshow

questions

6

In a daily demand forecasting model, what is a lag-7 feature and why include it?

level: juniorimportance: must knowfreq 74%

answer

  1. the target used as its own input
  2. why Saturdays resemble Saturdays
  3. the target column shifted in time
  4. same weekday, one week back
  5. unusable once horizon exceeds seven

basics

~20 s

A lag-7 feature is the target value from seven days earlier, placed on today's row as an input column. It hands the model last week's same-weekday demand, capturing weekly seasonality that a lag-1 feature alone misses.

solid answer

~50 s

For a row dated `t` whose target is demand on day `t`, the lag-7 feature holds the demand observed on day `t-7`. Because a week has seven days, that value is the same weekday, so it carries the weekly rhythm directly: Saturdays are compared against Saturdays rather than against Fridays. In practice it ships alongside `lag-1`, which carries the most recent level, and `lag-28`, which is four weeks back on the same weekday and is less sensitive to one odd week. The critical constraint is availability: the feature must be known when the forecast is made. Forecasting one day ahead, `y[t-7]` is safely in the past. Forecasting ten days ahead from the same origin, it is not — you would need a value that has not happened yet, so long-horizon models use lags at least as long as the horizon.

go deeper

for a junior

Be ready to state plainly that lag-7 is the target from seven days earlier and that it captures the weekly cycle. Mention that it must already be observed when the forecast is made.

for a middle

Explain why lag-1, lag-7 and lag-28 are usually shipped together, and show the mechanics: sort by time, shift by calendar date within each series, drop the incomplete first rows.

for a senior

Demonstrate that you check every lag against the forecast horizon and origin before it enters the table, and that you have handled missing dates and closed days rather than assuming a dense daily index.

for a principal

Own the decision of which lag set becomes the shared feature contract across many models and horizons, and the cost of maintaining one table whose columns are only valid for some of its consumers.

## What a lag feature is A time series is a target measured repeatedly over time: daily units sold, hourly requests, weekly revenue. Tabular forecasting turns that sequence into a supervised learning table, one row per timestamp, with the value to predict in a target column and everything known before that timestamp in feature columns. A **lag feature** is the simplest such column: the target's own earlier value, copied forward onto a later row. Writing the series as `y[1], y[2], ..., y[T]`, the lag-k feature on the row for time `t` is `y[t-k]`. Lag-1 is yesterday, lag-7 is the same day last week, lag-28 is four weeks back. The lag is a shift of the target column downward in time; nothing is aggregated or transformed. The first `k` rows have no value for lag-k and are usually dropped or masked. ## Why lag-7 specifically Daily human-driven series almost always have a weekly cycle. Retail demand peaks on weekends, B2B traffic peaks midweek, payroll-driven spending clusters on particular days. That cycle repeats with period 7, so `y[t-7]` sits at the same phase of the cycle as `y[t]`. That matters because lag-1 is a poor guide across a weekday boundary. If Sunday sells three times what Monday sells, a model leaning on lag-1 has to learn a large correction that depends on which day it is. Lag-7 removes most of that: it is already a Sunday when the target is a Sunday. A typical daily feature row therefore carries `lag-1` for the current level, `lag-7` for the same weekday last week, and `lag-28` for the same weekday four weeks back, which averages away a single unusual week without a long window. Lag features also carry information a calendar flag cannot. A day-of-week indicator says only *which* weekday it is, so the model can learn an average weekday effect. Lag-7 carries the actual number, which moves with trend, price changes, promotions and store-level shocks. The two are complementary, not substitutes. ## The horizon constraint Every feature must be computable at the moment the forecast is produced. Call that moment the forecast origin, and call the number of steps ahead the horizon `h`. A row predicting `y[t]` from an origin `h` steps earlier may only use values up to `y[t-h]`. That rule immediately restricts which lags are legal. For a next-day model (`h = 1`), lags 1, 7 and 28 are all fine. For a ten-day-ahead model (`h = 10`), lag-7 would require `y[t-7]`, which is three days *after* the origin — it does not exist yet. The usable lags start at 10: lag-14, lag-21, lag-28 keep the same-weekday property while remaining observable. Building a feature table once and reusing it for every horizon is a common way to smuggle unavailable values into training. There are two standard escapes. A **direct** model per horizon simply drops the lags shorter than `h` and trains on what remains. A **recursive** model predicts one step, feeds its own prediction back in as the lag for the next step, and iterates to the required horizon — cheaper in models but its errors accumulate as the horizon grows. ## Building lags correctly Three mechanical points cause most bugs. First, sort by timestamp within each series before shifting; an unsorted table produces lags that mix unrelated rows. Second, shift in the right direction — a sign error yields `y[t+7]`, a future value, which is a leak that makes offline metrics look brilliant and production collapse. Third, respect series boundaries: with many series in one table, lags must be computed within each series, otherwise the first rows of one product inherit the tail of another. Missing dates deserve care too. If a store is closed on Sundays and those rows are simply absent, a naive shift of seven rows lands six calendar days back, not seven. Reindex to a complete calendar and lag by date, not by row position. ## What an interviewer is checking The question looks definitional but it separates candidates who have built a forecasting table from those who have only read about one. Strong answers name the shift, tie the choice of 7 to the weekly period of the data, and volunteer the availability constraint without prompting. Weak answers describe lag-7 as "a rolling average of the last week", which is a different feature entirely, or claim any lag can be used at any horizon.

  • If you must forecast ten days ahead, is a lag-7 feature still usable?
    Not directly. At the forecast origin you would need the value from three days after the origin, which has not been observed. Either restrict the feature set to lags of at least ten (lag-14, lag-21, lag-28 keep the same-weekday property), or generate the intermediate values recursively by feeding the model's own predictions back in and accept the compounding error that comes with it.
  • How does a lag-7 feature differ from a day-of-week indicator column?
    The indicator says only which weekday the row falls on, so the model learns an average effect per weekday, constant across the whole history. Lag-7 carries the actual observed level from that weekday last week, so it moves with trend, price changes and local shocks. Most daily models carry both: the indicator for the stable pattern, the lag for the current level.
  • What can go wrong when the series has missing dates?
    Shifting by seven rows is not the same as shifting by seven days once dates are missing. If closed days are absent from the table, a row-based shift silently reaches back six or five calendar days and destroys the same-weekday property. Reindex onto a complete daily calendar first, lag by date, and keep an explicit flag for the days that were genuinely unobserved.

It is the shopkeeper's habit of checking last Saturday's till before ordering for this Saturday, rather than checking yesterday, which was a quiet Tuesday.

saying these in an interview costs you the question

  • Describes lag-7 as an average of the last seven days
  • Assumes any lag is usable at any forecast horizon
  • Computes lags without sorting by timestamp first
  • Shifts in the wrong direction and pulls in future values
  • Lags by row position when calendar dates are missing
  • Computes lags across series boundaries in a pooled table

context

open as a page

How does a centred rolling mean feature leak future values into a training row?

level: seniorimportance: must knowfreq 62%

basics

~20 s

A centred window straddles the row: a 7-day centred mean at day t averages t-3 through t+3, so three future values enter that row. The model then trains on inputs that will not exist at prediction time.

open as a page

Why add a 28-day rolling mean and rolling standard deviation alongside raw lag features?

level: middleimportance: should knowfreq 60%

basics

~20 s

Rolling aggregates compress recent history into stable columns. A 28-day trailing mean gives the current level with daily noise averaged out; the rolling standard deviation gives recent volatility. A single lag carries one noisy day and can express neither.

open as a page

For a 28-day-ahead forecast, how do recursive and direct multi-step strategies differ?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Recursive forecasting trains one one-step model and feeds its own predictions back as lags, so errors compound over 28 steps. Direct forecasting trains a model per horizon using only lags of at least that horizon, so nothing is fed back.

open as a page

Would you train one global model across thousands of store-item series, or one model per series?

level: principalimportance: should knowfreq 38%

basics

~20 s

Default to one global model trained on the pooled rows of all series, with a series identifier and static attributes as columns. It shares structure across series, covers brand-new ones, and leaves one artifact to operate. Reserve per-series models for high-value, atypical series.

open as a page

How do Fourier terms with period m=365 encode yearly seasonality as model columns?

level: middleimportance: nice to knowfreq 34%

basics

~10 s

Fourier terms are sine and cosine columns built from the date index: sin(2pikt/365) and cos(2pikt/365) for k = 1..K. Those 2K columns approximate a smooth yearly cycle using far fewer parameters than day-of-year indicators.

open as a page