skip to content

In a subscription churn table, why does a feature snapshotted at cancellation time rather than at the as-of stamp break the model live?

level: seniorimportance: must knowfreq 74%

answer

  1. the model is reading the answer
  2. cause or consequence, not correlation
  3. a mutable row has no event time
  4. neutral defaults at prediction time
  5. it fails silently, without errors

basics

~20 s

Because the feature records a consequence of the cancellation, not a cause. Offline it separates the classes almost perfectly; at prediction time the account has not cancelled yet, so the field holds a neutral value and the model scores confidently on nothing.

solid answer

~40 s

A feature is only eligible if its value could have been read at the as-of stamp. A field captured when the cancellation was recorded fails that test twice over: it sits after the boundary in time, and its value exists *because* the outcome happened. Offline the model looks extraordinary, because it is reading the answer. Live, every scored account is still active, so the field holds its default - status active, refund flag false, zero cancellation-flow visits - and the model loses the one signal it leaned on. Nothing errors; it simply scores everyone the same way. The eligibility test is not statistical but causal: could this value have been different because the outcome occurred? If yes, the field is a piece of the label, whatever its correlation looks like.

code

pseudocode · 11 lines
pseudocode
// leaky: the account record has no event time, so this reads today
features.plan_status = lookup(account_table, row.account_id).status

// eligible: only events observable at the row's own stamp
features.logins_28d = count(
    events where events.account_id == row.account_id
           and events.type       == LOGIN
           and events.event_time <= row.as_of
           and events.event_time >  row.as_of - 28 days)

row.label = churned(row.account_id, row.as_of, row.as_of + 30 days)

go deeper

for a junior

Remember the rule of thumb: a feature must be something you could have read at the moment of prediction, before the outcome was known.

for a middle

Explain both routes - a window that crosses the as-of stamp, and a field whose value exists only because the outcome happened - and why the second survives correct timestamps.

for a senior

Diagnose it: an implausible offline result, one dominant field, then a live model that scores flat and confident on neutral defaults without raising a single error.

for a principal

Make eligibility a reviewed contract per feature - source event, window, and whether the outcome could have changed the value - rather than a debugging step after the numbers look too good.

## Two leaks with the same shape Both failures put information from after the as-of stamp into a feature, but they arrive by different routes. 1. **Time travel.** The feature's aggregation window crosses the stamp - support tickets counted over the whole billing month when the stamp sits mid-month, or a trailing average computed to today rather than to the row's own date. The fix is mechanical: every window ends at `row.as_of`. 2. **Consequence features.** The window is correct but the field itself only takes its interesting value *because* the outcome happened. A refund flag, a cancellation-survey reason, a final-invoice amount, a 'winback offer sent' marker. These fields are effects of churn wearing the costume of predictors, and no amount of window discipline removes them, because they should not be in the table at all. ## The mutable row problem The most common route into the second kind is a dimension table read as it looks today. An account record with columns like `status`, `current_plan` or `cancelled_at` carries **no event time**: reading it returns the value as of the moment of the read, not as of the row's date. Join it to a two-year-old training row and the row learns today's outcome. The model then discovers a near-perfect rule - `status = cancelled` implies churn - which is true, trivially, and useless: at prediction time every account being scored is still active. The defence is to build features from dated events and join them with an explicit as-of predicate, so a value only enters a row if it was observable at that row's stamp. ## The cause-or-consequence audit Run each candidate field through three questions before it reaches the table: 1. **When was this value written?** If it has no timestamp, treat it as unavailable until it has one. 2. **Could it have been different because the outcome happened?** If the answer is yes, it is part of the label. 3. **What does it hold for an active account at scoring time?** If the honest answer is 'always the default', the model will get nothing from it in production. | Candidate feature | Offline behaviour | Value at prediction time | Verdict | |---|---|---|---| | Account status read from the live account table | near-perfect separation | always active | ineligible - carries the outcome | | Refund issued flag | very strong | almost always false | ineligible - a consequence | | Tickets opened between the stamp and the cancellation | very strong | none exist yet | ineligible - after the boundary | | Visits to the cancellation page in the 7 days before the stamp | strong | genuinely available | eligible - window ends at the stamp | | Logins in the trailing 28 days before the stamp | moderate | available | eligible | Note the fourth row: intent signals are not banned. A visit to the cancellation page is a legitimate and powerful feature precisely because it happens before the decision is executed and can be read at the stamp. The boundary is the stamp, not the topic. ## How it announces itself, and how it does not The tell is an offline result that nobody in the business believes - separation far above what a human analyst achieves - usually with one field dominating any importance ranking. The second tell arrives later and quieter: the live model does not throw, does not time out, and does not alert. It scores every active account on neutral inputs and produces a flat, confident, useless distribution. A leakage bug is a correctness bug that presents as silence. The standing defence is a written eligibility rule per feature - the source event, the window, and the answer to question two - reviewed when the feature is proposed rather than when the offline number looks too good. ## What this is not This is a question about which fields are **eligible** to be features at all. The machinery that enforces an as-of join at scale for a whole feature platform - how historical values are stored, indexed and backfilled so a training job can replay them - is a separate concern with its own design. Eligibility comes first: the best as-of join in the world will faithfully reproduce a consequence feature, dated correctly, and the model will still collapse in production.

  • A field passes the timestamp check but still looks suspiciously strong. What else do you test?
    Ask what writes it. A correct timestamp only proves the value existed by the stamp; it does not prove the value is independent of the outcome. An internal flag set by an agent who has already spoken to a departing account, or a queue assignment made after a retention team read a list, is dated before the cancellation and is still downstream of it. Trace the writer, not just the clock.
  • How would leakage of this kind show up in production monitoring?
    Mostly as absence. No errors, no latency change, no schema alarm: the field is present and holds its default, so the serving path is healthy by every technical signal. What moves is the prediction distribution - scores collapse towards one value and the flagged list stops resembling the offline evaluation - which is why prediction-side monitoring catches it and input validation does not.
  • Does excluding consequence features mean excluding anything related to cancelling?
    No. The boundary is temporal, not topical. A visit to the cancellation page, a downgrade browsed but not taken, a support contact about pricing - all of these happen before the decision is executed and are readable at the as-of stamp, so they are legitimate and often the strongest honest signals available.

saying these in an interview costs you the question

  • Any column in the account table is fair game as a feature
  • A feature with a huge importance score is proof the model works
  • Leakage would show up as an error or an alert in serving
  • If the timestamps are correct the feature cannot be leaking
  • Anything mentioning cancellation must be dropped from the table
  • A near-perfect offline result just means the problem was easy