Why add a 28-day rolling mean and rolling standard deviation alongside raw lag features?
answer
- one column instead of many noisy days
- level and volatility, separately
- trees cannot extrapolate a trend
- short reacts fast, long stays stable
- for horizon h the window ends at t-h
basics
~20 sRolling aggregates compress recent history into stable columns. A 28-day trailing mean gives the current level with daily noise averaged out; the rolling standard deviation gives recent volatility. A single lag carries one noisy day and can express neither.
solid answer
~50 sA raw lag is one observation, so it inherits that day's noise, promotions and outages. A trailing 28-day mean over `y[t-28] ... y[t-1]` estimates where the series currently sits; a trailing standard deviation over the same window says how erratic it has been, which lets a model treat a spike differently in a calm series than in a volatile one. Window length is the real decision: a short window tracks level shifts quickly but is noisy, a long window is stable but reacts late, so a common feature row carries both a 7-day and a 28-day version. The window must end strictly before the value being predicted — for horizon `h` it ends at `t-h`. Also record how many observations the window actually contained, because early rows and gap-ridden series produce means over far fewer points than the nominal width.
go deeper
Know that a rolling feature summarises a stretch of recent history in one number, and that the window must stop before the day you are predicting.
Explain the mean as a level estimate and the standard deviation as a volatility estimate, and articulate the short-versus-long window tradeoff rather than naming one width.
Show the operational care: windows shifted by the horizon, minimum observation counts, widths aligned to the seasonal period, and counts emitted alongside sparse aggregates.
Own how many window variants the feature set carries. Every extra width costs compute, storage and reviewer attention, and buys correlated columns; be able to justify the set you standardise on.
## Two kinds of history A forecasting feature row can look backwards in two ways. A **lag** picks out one earlier observation: `y[t-1]`, `y[t-7]`, `y[t-28]`. A **rolling aggregate** summarises a whole stretch of earlier observations into one number: the mean, standard deviation, minimum, maximum, median or count over a trailing window. Both answer "what has this series been doing", but they trade off differently between precision and stability. A single lag is precise about one day and silent about every other. It is also fragile: if `y[t-1]` happened to be a stockout, a promotion or an instrumentation failure, that shock lands directly in the feature. A window mean averages the shock down by roughly the window width, so a 28-day mean moves only slightly when one day misbehaves. ## What the mean contributes The rolling mean is the model's estimate of the current level. It matters most for models that cannot extrapolate. Gradient-boosted trees and random forests predict by partitioning feature space and averaging targets inside each leaf, so they can never produce a value outside the range seen in training. If a series has grown steadily, a tree fed only calendar features will systematically under-predict the future. Feeding it a trailing mean re-anchors the prediction: the tree learns a relationship between the recent level and the next value, and the level column carries the growth in from outside. A useful refinement is to feed the target relative to the trailing mean, or to feed the ratio of a short window mean to a long window mean. The ratio is a momentum column: above one means the series is running hot relative to its own recent baseline, below one means it is cooling. That framing generalises across series of very different magnitudes far better than raw levels. ## What the standard deviation contributes The rolling standard deviation is a volatility column. Two series can share the same mean and behave completely differently: one steady at 100 units a day, the other alternating between 20 and 180. A model that knows only the mean treats a forecast of 100 as equally reliable in both. The rolling standard deviation lets it condition on the difference — useful for the point forecast, and essential when the model produces intervals or the downstream consumer sizes safety stock. Volatility also acts as a regime signal. A sustained rise in the rolling standard deviation often marks a shift in behaviour: a promotion campaign started, a data feed became unreliable, demand fragmented across channels. As a feature this lets the model discount the recent level exactly when the recent level has become less trustworthy. ## Choosing the window Window length is a bias-variance decision expressed in time. A short window (7 days) reacts within a week to a genuine level shift, but its estimate wobbles because it averages few points. A long window (28, 56, 91 days) is a much steadier estimate, but after a step change it takes most of the window length before it reflects reality, so the model keeps forecasting the old level. Because no single answer is right, the standard practice is to include several windows and let the learner decide which to weight. Aligning at least one window to the seasonal period helps: for daily data with a weekly cycle, a 7-day or 28-day window contains whole numbers of weeks, so the mean is not biased by containing four Saturdays and only three Sundays. A 30-day window does not have that property and quietly mixes weekday composition from row to row. ## Availability and the horizon The same rule that governs lags governs windows: every input must exist at the forecast origin. For a next-day model the window covers `y[t-28]` through `y[t-1]` and stops there. For a horizon of `h` steps, the window must end at `t-h`, which means the aggregate is effectively a lagged rolling feature — compute the trailing statistic, then shift it back by `h`. Reusing a next-day feature table for a two-week-ahead model is one of the most common ways a project ends up training on values it will never have in production. ## Degenerate windows Two edge cases spoil rolling features quietly. At the start of a series the window is not full: the first row has one observation behind it, not 28. Averaging whatever is available produces a high-variance number that looks like a normal feature, so either require a minimum count before emitting a value, or carry the observation count as its own column so the model can learn to distrust thin windows. The same applies to intermittent series and gaps: a window nominally 28 days wide that contains three non-missing days is a very different quantity from one containing 28, and the standard deviation over two points is close to meaningless. Finally, keep the window trailing. A window centred on the current row reaches into the future by half its width, and everything downstream of that mistake — validation scores, feature importances, the go-live decision — becomes uninformative.
- How would you choose between a 7-day and a 91-day rolling mean?Not by choosing — ship both and let the model weight them. Conceptually the short window tracks level shifts within a week but wobbles, and the long window is a stable baseline that lags a step change by most of its width. If the series is prone to abrupt shifts, weight towards short windows; if it is noisy but stable, towards long. Prefer widths that are whole multiples of the seasonal period.
- Why does a rolling mean help a gradient-boosted tree more than it helps a linear model with a trend term?Trees predict by averaging targets within leaves, so they cannot output a value above anything seen in training and will under-forecast a growing series. A trailing mean smuggles the current level in as an input, so the tree learns a relationship to the level rather than to absolute time. A linear model with an explicit time term can already extrapolate the trend directly.
- What extra column should accompany a rolling mean on a sparse series?The count of non-missing observations actually inside the window. A 28-day mean built from three observations is far less reliable than one built from 28, but both arrive as a single number the model cannot distinguish. Emitting the count, or suppressing the feature below a minimum count, lets the model learn how much to trust it.
saying these in an interview costs you the question
- Lets the window include the day being predicted
- Reuses a next-day window table for longer horizons
- Treats a rolling mean as a substitute for raw lags
- Ignores that early rows have partly empty windows
- Picks a 30-day window over 28 on daily weekly-cycle data
- Reports a standard deviation computed over two points