skip to content

A nightly-rebuilt snapshot table reflects today's state, not decision-time state. How do you train on it safely?

level: principalimportance: nice to knowfreq 30%

answer

  1. The table has no memory
  2. Every mutable column is suspect
  3. Which fields can never change after creation?
  4. Rebuild the history, or start logging now

basics

~20 s

Treat every mutable column as leaked until proven otherwise. A nightly rebuild overwrites fields with post-outcome values, so train only on columns that cannot change after the decision, or reconstruct earlier values from an append-only change history.

solid answer

~50 s

The danger with a table rebuilt nightly is that it has no memory: a status, a balance or a disposition code shows the value it holds today, which for many rows is the value it took *because* of the outcome. There is no single leaky column to delete — the whole table is contaminated, and no individual row looks wrong. I triage the columns into three buckets. Write-once fields such as application date, region or original product type are safe. Mutable fields that have an append-only change history somewhere upstream can be reconstructed as they stood on the decision date. Mutable fields with no history are excluded and logged as a data-collection gap. Then I quantify: if the safe columns give a much weaker model, that weaker number is the honest baseline, and the gap becomes the business case for capturing history from now on.

go deeper

for a junior

Know the difference between a table storing the current value of a field and one storing every change to it. Only the second can tell you what was true at an earlier moment.

for a middle

Be ready to classify columns as write-once versus mutable, and to explain why a mutable status field in a latest-state table can quietly encode the very outcome you are predicting.

for a senior

Show a reconstruction plan: find the upstream change log, rebuild the mutable fields as of each decision date, and quantify the score gap against the naive snapshot model so the cost is visible.

for a principal

Own the call between shipping a weaker honest model, funding decision-time logging, and declining the project. Say how you would resource the instrumentation and what you commit to the sponsor in the meantime.

## Why a snapshot leaks A **snapshot table** stores the current value of each field for each entity. A rebuild replaces it wholesale, so what you read tonight is the state of the world tonight. A **history table** instead stores every change with the time it happened, so you can ask what a field held on any past date. If your training rows are drawn from a snapshot but your labels refer to events in the past, then every mutable column has had time to be overwritten by consequences of those events. A hospital record's discharge disposition code is a good example: written or amended after the episode concludes, so by the time the snapshot is rebuilt it may encode the readmission you are trying to predict. Nothing about the row looks odd. The value is genuine — it is simply as of the wrong moment. This is the hardest form of target leakage to argue about, because the usual detection instincts fail: - Inspecting rows fails. Each row is internally consistent and plausible. - Column-name heuristics fail. The offending fields are ordinary operational statuses. - A single-feature ablation may fail too, since the contamination is spread across many mutable columns rather than concentrated in one. You detect it from the *build semantics* of the table — is it a full rebuild, is there an effective-date or valid-from column, is there an upstream change log — not from the data. ## The three-bucket triage **Bucket one: write-once fields.** Values set at creation and never amended — date of birth, application channel, original loan term, admission date, the product bought. These are safe by construction, and they are the floor you can always build on. **Bucket two: mutable with recoverable history.** Many operational systems keep a change log or an audit trail even when the analytics snapshot does not. If it exists, you can reconstruct each field as it stood on the decision date and recover most of the modelling value. This is real engineering work, and it is usually the highest-return work on the project. **Bucket three: mutable with no history.** Exclude them. Not "use them carefully" — exclude them, and record why, because the next person to build a table here will otherwise pull them straight back in. ## The number you report Once you drop buckets two and three, the model is usually much weaker. The instinct to soften that is exactly the failure this question tests. Report two numbers with an explicit meaning for each: the snapshot-trained score, labelled as unachievable because it uses information the system will not have; and the honest score, labelled as what the business would get on day one. Never let the first number become the target anyone plans against. ## The strategic choice This is where the judgment lives, and there is no single right answer. Three defensible paths: **Ship the honest, weaker model now.** Justified when even the reduced model beats the current process — often a manual rule or nothing at all. It also starts generating the operational evidence that funds the next step. **Instrument first, model later.** Begin logging the exact feature values used at each decision, at the moment the decision is made. In a few months you own a clean, leakage-proof training set by construction. This is the durable fix, and it needs a sponsor because the payoff is delayed. **Decline, or narrow the scope.** If the safe fields carry almost no signal and the decision is high-stakes, the responsible answer is that this problem cannot be modelled from this data yet. Saying so early is cheaper than a year of unexplained production underperformance. Most mature teams run the first and second in parallel: a modest model shipped on write-once fields, plus decision-time logging switched on the same week. ## What to say in the room Frame the constraint before the technique. The organisation did not choose to lose history — the snapshot exists because it is cheap and answers the operational question "what is true now". Modelling asks a different question, "what was true then", and that requires a different storage decision. Presenting it that way turns an argument about a broken feature into a funding conversation about instrumentation, which is the conversation you actually want.

  • Why can't you spot snapshot leakage by inspecting a sample of rows?
    Because every row is plausible. The values are genuine, just as of the wrong moment, so there is no odd number or impossible combination to notice. You detect this from the table's build semantics — a full nightly rebuild, no effective-date column, no upstream change log — rather than from the data itself.
  • The history-free model is far weaker. How do you present that to a sponsor?
    Present both numbers with what each one means: the snapshot score is unachievable because it uses information the system will not have at decision time, and the honest score is what they would get on day one. Then price the gap — the cost of decision-time logging against the value of the extra performance — and let them choose.
  • Is a snapshot ever safe to train on?
    Yes, in two cases. If you restrict features to write-once fields such as an application date, an original term or a channel, nothing can have been overwritten by the outcome. And if each nightly snapshot is archived rather than replaced, you can select the version that predates each decision. The danger is specific to mutable fields in a single latest-state copy.

It is like asking a patient's current chart what their symptoms were on admission. Every entry is accurate, but the chart has been overwritten by everything that happened since.

saying these in an interview costs you the question

  • Assumes the snapshot is fine because every row is real data
  • Hunts for one leaky column instead of a leaky table
  • Reports the snapshot score as expected production performance
  • Treats a mutable status field like a stable attribute
  • Backfills earlier values by guessing from today's

context