You backfill 90 days of ad training rows with a 7-day click count computed from today's full log — what breaks?
answer
- the row's own future is in the feature
- one aggregate, many timestamps
- leak lives inside a row, not across folds
- anchor the window on served_at
- strictly before the serve time
basics
~20 sEvery historical row's count includes clicks that happened after its own impression, so the feature is computed after the outcome it predicts. Offline metrics jump and the gain vanishes online, because serving can only see the past.
solid answer
~40 sThe backfill computed one aggregate over the whole log and attached it to rows from every day, so a row from day 3 carries a count that spans clicks from days 1 to 90. That is the future leaking into the feature, and it is a per-row defect — no train/validation split can rescue it, because the leak is inside each row rather than across a fold boundary. The offline metric rises, sometimes dramatically, then the online model performs like the old one, since at serve time a 7-day count can only cover the seven days before the request. The fix is an as-of computation: for each row, count clicks in `[served_at - 7 days, served_at)`, strictly before its own serve time, which also excludes the click that produced its own label.
code
sql · 8 linesSELECT i.impression_id,
COUNT(c.click_id) AS user_clicks_7d
FROM impressions i
LEFT JOIN clicks c
ON c.user_id = i.user_id
AND c.clicked_at < i.served_at
AND c.clicked_at >= i.served_at - INTERVAL '7' DAY
GROUP BY i.impression_idgo deeper
The rule to remember: a training row may only contain information that existed before that row's own event. A number computed today cannot be attached to a row from three months ago.
Explain the bounds precisely — the window ends strictly before the serve timestamp and starts seven days earlier — and why that also keeps the row's own attributed click out of its own feature.
Recognise the signature in production terms: a big offline gain that dies online, a feature whose offline distribution does not match what serving produces, and lift concentrated in the oldest backfilled rows.
The lasting question is how the platform makes as-of correctness the default rather than a review item, so that adding a feature cannot silently produce an unreproducible offline win.
## What the backfill actually did A backfill rebuilds historical training rows so a newly added feature exists for the past as well as the present. The tempting implementation is one aggregate over the log: group clicks by user, count them, join the result onto every row. It is one pass, it is fast, and it is wrong — because the aggregate is **as of the backfill job's run time**, not as of each row's own timestamp. So a row whose impression was served on day 3 receives a "7-day click count" assembled from a log that runs to day 90. For that row the feature is not a summary of the user's recent past; it is a summary of the user's entire future. And the single most predictive thing about that future is whether the user was, in fact, a clicker — which is exactly what the row's label says. ## Why the offline metric rises and the online one does not - Offline, the feature is **correlated with the label by construction**, so ranking metrics improve and the improvement looks like a genuine modelling win. - The lift is largest on the **oldest** rows, which have the most future folded into them, and smallest at the end of the backfill window. - Online, the serving path computes the same feature from data that ends at the request. It can only produce the honest version, so the model receives an input distribution it never trained on and behaves like the baseline — or worse, because it leaned on a signal that is now uninformative. The operational tell is a large offline gain that fails to reproduce in any live comparison, together with a feature whose offline distribution does not match what the serving path produces for the same users. ## The as-of rule For every derived aggregate on a training row, the window must be anchored on the row's own event time: 1. **Upper bound: strictly before `served_at`.** Not less-than-or-equal — an event stamped at the same instant as the serve cannot be shown to have preceded the decision. 2. **Lower bound: `served_at - 7 days`.** The same span the serving path will use, so training and serving compute the same thing. 3. **The row's own attributed click is automatically excluded**, because that click necessarily occurs at or after `served_at` and the strict upper bound cuts it out. This is the same discipline as the label join, pointed the other way: the label looks **forward** from the serve time within the attribution window, and every feature looks **backward** from it. | computation | window anchored on | includes the row's own label event | safe to serve | |---|---|---|---| | one aggregate over the whole log | the backfill job's run time | yes, and every later event too | no | | aggregate as of yesterday's midnight | a fixed daily boundary | no, if the boundary precedes the serve | yes, with stated staleness | | aggregate as of `served_at` | each row's own serve time | no | yes | ## The cheaper approximation, and its honest cost Recomputing a window per row is expensive at billions of rows, so pipelines often snapshot the aggregate at a coarse boundary — the value as of the most recent midnight before the impression — and join on that boundary. This is legitimate **provided the boundary precedes the serve time**, and provided the serving path uses a snapshot of the same age. The gap between "as of the last boundary" and "as of this instant" is staleness the model should be trained with rather than shielded from; what is never acceptable is a boundary that lands after the serve. ## How to catch it before it reaches a launch decision - **Compare distributions.** Take the feature as the serving path would compute it for a sample of recent requests, and compare against the same feature in the training set. A leaking backfill shows a visible shift, typically fatter in the high-count tail. - **Slice the offline metric by row age.** A genuine feature helps roughly evenly across the backfill window; a leaking one helps most where the most future was available. - **Re-run on a truncated log.** Rebuild the rows for day 3 using only data up to day 3. If the feature values change, the original was not as-of correct. - **Review the aggregate's window bounds as part of the change**, not after the metric comes back good. A suspiciously large offline gain is the classic symptom, and by the time it is celebrated the diagnosis is socially expensive. The general rule is worth stating plainly: **a training row may only contain what the serving path could have known at that row's timestamp.** Anything else is measuring the pipeline, not the model.
- The team says a strictly time-ordered train/validation split will catch this. Are they right?No. An ordered split protects against information crossing the boundary between folds, but here every single row already contains its own future before any split is made. Train and validation both look better, the relationship holds in both, and the split reports a healthy generalisation gap. The defect is in row assembly and only a row-level as-of rebuild — or a distribution comparison against what serving can compute — exposes it.
- Recomputing a rolling window per row is too expensive at billions of rows. What is the acceptable approximation?Snapshot the aggregate at a coarse boundary that precedes the serve — for example, its value as of the last midnight before the impression — and have the serving path read a snapshot of the same age. Correctness is preserved because the boundary is in the row's past; the cost is staleness, which is now identical in training and serving and therefore something the model learns to live with rather than a discrepancy.
It is like grading a stock picker's old calls using prices printed after the calls were made. The record looks brilliant, and the method that produced it predicts nothing tomorrow.
saying these in an interview costs you the question
- Computes the aggregate once over the whole log and joins it to every row.
- Says a time-ordered validation split will catch a leak of this kind.
- Uses a window ending at the backfill job's run time for historical rows.
- Treats a large unexplained offline gain as evidence of a good feature.
- Bounds the window at less-than-or-equal to the serve timestamp.
- Claims the counts were stale rather than too fresh.