How does a centred rolling mean feature leak future values into a training row?
answer
- training must simulate prediction time
- the window straddles the row
- half the width lies ahead
- the target sits inside its own feature
- trailing only, shifted back by the horizon
basics
~20 sA centred window straddles the row: a 7-day centred mean at day t averages t-3 through t+3, so three future values enter that row. The model then trains on inputs that will not exist at prediction time.
solid answer
~50 sA trailing window ends at the last observed point; a centred window straddles the row, so half its width lies after it. A 7-day centred mean at time `t` averages `y[t-3] ... y[t+3]`, and since `y[t]` itself is inside, the feature is partly the target. Training error drops to something implausible, that column dominates feature importance, and the moment the model runs live the future half of the window does not exist. The tell is usually a validation score far better than any plausible baseline, plus a large gap between validation and the first live week. The fix is mechanical: every window is trailing, and for a horizon of `h` steps the window ends at `t-h`, not at `t`. The same audit catches its cousins — interpolating a missing target across a gap, and per-series statistics computed over the entire history rather than only the part preceding the row.
go deeper
Know that features may only use information available before the moment of prediction, and that a window centred on the current row breaks that rule.
Trace the arithmetic: name which timestamps a centred window of a given width covers, and state the trailing-window fix including the shift by the forecast horizon.
Show the diagnostic instinct. Describe the signature of a leak, the column-by-column availability audit you run, and the related traps in imputation and whole-history statistics.
Own prevention as process rather than vigilance: an as-of-timestamp convention for every source, a feature definition reviewed once, and a rule that an unexplained jump in accuracy blocks release until explained.
## The rule being violated Every feature on a training row must be computable from information that exists at the forecast origin — the moment the prediction would actually be made. Training is meant to simulate production. If a column uses information that will not be available when the model runs, the model learns a relationship it cannot exploit, and the evaluation measures a task nobody asked for. That is **look-ahead leakage** in feature construction. It is distinct from leakage introduced by how data is split for evaluation; here the column itself is unbuildable in production, so no evaluation scheme can rescue it. ## Why centring is the classic case Smoothing is a natural instinct: a noisy series is easier to reason about once averaged. Centred windows are the default in a lot of visualisation and smoothing work because they do not shift the curve sideways — a trailing mean lags the series by roughly half its width, while a centred mean sits on top of it. That property is exactly what makes it a good picture and a fatal feature. A centred window of width `w` at time `t` covers `t - (w-1)/2` through `t + (w-1)/2`. At `w = 7` that is `y[t-3]` through `y[t+3]`. Three observations are future relative to `t`, and `y[t]` itself is in the average, so the feature contains a fraction `1/7` of the exact target. A model with any capacity finds that relationship immediately: multiply the smoothed column by seven, subtract the neighbours it can partly infer, recover the target. The symptoms are consistent. Training and validation error land far below any sensible benchmark — better than a domain expert believes possible. The smoothed column takes an overwhelming share of feature importance. Removing it destroys performance in a way no legitimate feature does. Then the first production week reports errors several times the validation figure, because the column is now built from whatever the pipeline does when half the window is missing: a shorter window, a forward fill, or nulls. ## The fix Use trailing windows only, and offset them by the horizon. For a forecast made `h` steps before the target, the window ends at `t-h`; compute the trailing statistic and then shift it back `h` steps. A next-day model may use a window ending yesterday; a fourteen-day-ahead model may not, and needs its own feature table. The general audit is a single question asked of every column: *what is the latest timestamp any input to this number carries, and is that timestamp at or before the forecast origin?* If the answer is later than the origin, or if you cannot answer at all, the column is not admissible. Doing this column by column is tedious and catches nearly everything. ## The neighbours of this bug **Interpolating a missing target across a gap.** Linear interpolation between the last observation before a gap and the first one after it computes the filled values from data on both sides. Every interpolated row therefore contains information from after itself, and if those rows are used as targets or as inputs to lags, future information propagates through the table. Filling forward from the past is different in kind — it uses only earlier data — but carries its own hazard: if a missing target is filled with the previous day's value, that row's target becomes exactly its own lag-1 feature, and the model learns to copy. Metrics improve on rows that were never real observations. The honest handling is to mark gaps explicitly, exclude imputed rows from the target, and carry an observed flag. **Statistics computed over the whole history.** A per-series summary — the mean level of an item across the entire dataset, the item's peak, a ranking of items by total volume — computed once over all rows and joined back gives every training row a summary of its own future. The correct form is expanding: for the row at time `t`, the statistic is computed only from data strictly before `t` (or before `t-h`), so it evolves along the series exactly as it would in production. **Aligning an external series at the wrong timestamp.** Weather, prices and promotion calendars are often stored by the date they describe rather than the date they became known. A promotional discount recorded against its effective date is a legitimate feature if the plan was locked in beforehand, and leakage if it is a value only known after the fact. Store the timestamp at which each fact became knowable, not only the timestamp it refers to. ## Interview framing This question is a filter for people who have shipped a forecast rather than only fit one. Weak answers say leakage means "including the target in the features" and stop. Strong answers describe the mechanism precisely, name the diagnostic signature — implausibly good validation error, one dominant column, a large live-versus-offline gap — and give an operational rule for prevention rather than a promise to be careful. The best answers add that suspiciously good results deserve the same investigation as suspiciously bad ones, which is a habit, not a technique.
- What would make you suspect leakage before the model ever reaches production?An error far below any credible baseline, a single column carrying most of the feature importance, and a sharp collapse when that column is dropped. Then audit the column directly: ask what the latest timestamp feeding it is and whether that is at or before the forecast origin. Suspiciously good results deserve the same scrutiny as suspiciously bad ones.
- Is forward-filling a missing target across a gap the same kind of leak?Not quite. Forward fill copies an earlier value forward, so it uses no future information, but it makes the filled row's target identical to its own lag-1 feature and the model learns to copy while metrics improve on rows that were never observed. Interpolating across the gap is a genuine future leak, since each filled value is computed partly from data after it.
- How should a per-series average level be computed if it is used as a feature?As an expanding statistic: on the row at time t, use only observations strictly before the forecast origin for that row, so the value evolves down the series exactly as it would in production. Computing one average over the entire history and joining it to every row hands each training row a summary that includes its own future.
saying these in an interview costs you the question
- Centres rolling windows because the curve looks smoother
- Defines leakage only as putting the target in the features
- Celebrates an implausibly low validation error
- Interpolates missing target values across a gap
- Joins a whole-history per-series average onto every row
- Reuses one feature table across every forecast horizon
- Aligns external data by the date it describes, not when it was known