skip to content

Why does a per-city average demand computed over the full three-year history leak, even with forward-chaining folds?

level: seniorimportance: should knowfreq 44%

answer

  1. the folds are fine, the column is not
  2. when was this value computable?
  3. one number built from every year
  4. the future joined onto the past
  5. compute it as of each row's date

basics

~20 s

The aggregate summarises the entire timeline, so a row from year one already carries year three's demand. Every training row sees the future and every validation row sees its own outcomes, no matter how correctly the folds are ordered.

solid answer

~50 s

The split is fine; the feature is not. A per-city average computed once over three years and joined onto every row carries information backwards: a January-2023 row's feature value depends on demand recorded in 2025. Forward-chaining folds cannot repair that, because the leak happened upstream, when the column was built. Two things go wrong at once. Validation rows carry a summary that partly contains their own targets, so the score is optimistic. And the value the model was trained to rely on is one that will never exist at scoring time, when only the past is available. The fix is to make the feature as-of: for each row, compute the city average using only data dated before that row, so the column can be reproduced live. Recomputing it inside each fold's training data is necessary but not sufficient on its own.

go deeper

for a junior

Recall that a feature summarising the whole dataset carries information from dates after the row it is attached to, and that a correctly ordered split does not undo that.

for a middle

Explain both failures: validation rows whose feature partly contains their own target, and training rows fed a value that could not have been computed at their date. Describe the as-of alternative.

for a senior

Demonstrate the audit habit - for every column, when could this have been computed - plus the train-serve reasoning about whether the live pipeline can reproduce the value, and a fallback rule for thin history.

for a principal

Own the structural fix over the vigilance fix: a feature path that only accepts data older than the row being served, so whole-timeline aggregates are not expressible and no reviewer has to catch them by eye.

## The leak is in the pipeline, not the split This one is dangerous because the code that reviewers stare at - the fold construction - is correct. Folds are ordered, nothing from the future appears on the training side by date, the boundary assertions pass. The contamination entered earlier, when a feature was built by aggregating across the whole dataset. A per-city average demand computed once over three years of store history and joined onto every row is a single number per city, derived from every day in the file. Attach it to a row dated January 2023 and that row now carries a compressed summary of 2024 and 2025. The row's feature vector contains its own future. ## The two separate failures **Backwards flow of the future into training.** The model learns relationships conditioned on a number that already encodes how the series turned out. If a city's demand climbed steeply in year three, its whole-history average is elevated, and year-one rows for that city carry an elevated feature. The model can pick up "cities whose average is high relative to their current level will grow", which is not a discovered pattern - it is the answer, written into the input. **Validation rows summarising their own outcomes.** The aggregate was computed over the validation period too. A validation row's feature is partly a function of that row's own target and of its neighbours' targets. Predicting a value from a summary that contains it is close to trivial, and the fold score reflects that rather than any forecasting skill. ## Why refitting the feature per fold is necessary but not enough The standard reflex is "recompute the aggregate inside each fold's training data". That closes the second failure - the validation period no longer feeds the number - and it is the right first move. It does not close the first. Within a training block spanning January to December, a single aggregate over the whole block still gives a January row a value influenced by December. That is harmless for the fold's own score, but it trains the model on a quantity that has no runtime counterpart: in production, scoring a row on 3 January, the only thing computable is an average over data up to 2 January. The consequence is a train-serve mismatch rather than an inflated fold score. The model relies on a stable, whole-period statistic; live it receives a noisier, shorter-history one, especially for cities that are new or thinly observed. Performance degrades in ways the offline number never predicted. ## The as-of construction The repair is to define the feature the way production must compute it: for a row dated `d` in city `c`, the value is an aggregate of city `c`'s rows dated strictly before `d`. An expanding average uses everything up to `d`; a trailing window uses a fixed recent stretch. Either way the column becomes reproducible at scoring time and contains no future. Three practical consequences follow. Early rows have few or no prior observations, so the feature needs a defined fallback - a global prior, a shrunk estimate, or an explicit missing indicator - and that fallback must be identical offline and online. The column becomes noisier, particularly at the start of a city's history, and the model's measured performance drops; that drop is the illusion leaving, not a regression. And the computation is heavier, since a single grouped mean becomes a per-row running statistic. ## Shrinkage and small groups When a city has three observations, its as-of average is mostly noise, and a model that leans on it will overfit those cities. Shrinking each city's estimate toward the global mean by an amount that depends on its count - heavy shrinkage when data is thin, almost none when it is plentiful - keeps the feature stable. The shrinkage strength is itself a hyperparameter and belongs inside whatever is refit per fold. ## Spotting the class of bug The question that catches all variants of it: *for each column, at what moment could this value have been computed, and does that moment precede the prediction?* Any statistic taken across the entire table - a mean, a count, a rank, a normalisation constant, a frequency encoding, a cluster assignment - fails that test unless it was explicitly built as-of. A fast empirical smoke test is to train on an early period and score a much later one after rebuilding features from the early period only. If the score collapses relative to the fold estimate, something in the pipeline was quietly reading across the whole timeline. A structural safeguard beats vigilance: build features through a path that only ever accepts data with a timestamp strictly earlier than the row being served. When the same code produces both the training column and the live one, whole-timeline aggregates stop being expressible.

  • Is recomputing the aggregate inside each fold's training data enough?
    It is necessary, not sufficient. It stops the validation period from feeding the number, which removes the inflated fold score. But a single aggregate over the whole training block still gives early rows a value shaped by later ones, and that value has no runtime counterpart. The model then relies offline on a statistic production cannot supply.
  • How do you handle rows where the as-of history is nearly empty?
    Define an explicit fallback and use the identical rule offline and online: a global prior, an estimate shrunk toward that prior by group count, or a missing indicator the model can learn from. Never silently backfill from a later value. Shrinkage strength is a hyperparameter and should be refit inside the fold like any other.
  • What quick test exposes a whole-timeline aggregate hiding in a feature pipeline?
    Train on an early period and score a much later one after rebuilding every feature from the early period alone. If that score collapses against the cross-validated estimate, some column was reading across the full timeline. Asking of each column when its value could first have been computed catches the rest by inspection.

saying these in an interview costs you the question

  • Believes ordered folds make any feature safe
  • Recomputes per fold and calls the problem solved
  • Backfills an early row's missing average from later data
  • Treats the score drop after the fix as a regression
  • Uses a raw group mean for a city with three observations

context