A replay enriches each stored record with a customer tier read from a table that is overwritten nightly - what is wrong with the output?
answer
- the lookup was correct by coincidence
- one table, twelve months of records
- today's attribute on last March's row
- validity interval, looked up as-of
- overwritten in place cannot be undone
basics
~20 sEvery historical record is labelled with today's tier rather than the one in force at the time, so the replay quietly rewrites history. Reproducing the past needs reference data that retains its own change history, looked up as of each record's timestamp.
solid answer
~50 sThe live job was correct by coincidence: it read the reference table while the table was current, so each record got the tier it actually had. A replay breaks that coincidence and reads one table for a whole year of records. A customer promoted last week is now premium for every month of the replay, and totals per tier no longer reconcile with what was published at the time. The fix depends on the reference data, not the job. If it is append-only with a validity interval per version, do an as-of lookup on the record's own timestamp. If it is overwritten in place, the previous values are simply gone: the honest options are to reconstruct them from an audit trail if one exists, to restrict the replay to a period where the attribute did not change, or to state the restatement explicitly to consumers. Audit which enrichments are mutable **before** the replay, not after.
go deeper
Recognise that looking a value up today and attaching it to a record from months ago produces a wrong row that nothing will flag. The failure is a confident wrong answer, not a missing one.
Explain why the live job was correct only by coincidence, and classify reference data as immutable, versioned with validity intervals, or overwritten in place - the third cannot reproduce the past whatever the job does.
Audit every enrichment for mutability before committing to a replay, including thresholds and constants read from configuration, and choose deliberately between an as-of lookup, a narrowed replay and a declared restatement.
Set the expectation that reference data feeding pipelines is kept with validity intervals, and weigh that storage and modelling cost against the alternative: a platform whose historical numbers can never be regenerated, only approximated.
## The lookup that was accidentally correct Enrichment is the most ordinary thing a pipeline does: take a record, look something up by key, attach it. While the job runs live, the lookup is correct for a reason nobody writes down - the record and the reference table are being read at the same moment. The attribute attached is the attribute in force. Replay removes that coincidence. One table, read once, is now attached to twelve months of records, and the answer it gives is the answer for today. The result is not an error, a gap or a failure. It is a plausible, complete, confidently wrong dataset, and nothing in the run will flag it. ## Three kinds of reference data What a replay can honestly reproduce is decided entirely by how the reference data is kept: | Reference data | Replay behaviour | What you can do | |---|---|---| | Immutable once written (a code list that only gains entries) | reproduces the past exactly | join on key; nothing special needed | | Append-only with a validity interval per version | reproduces the past exactly | look up as of the record's own timestamp, not as of now | | Overwritten in place, newest value only | cannot reproduce the past at all | reconstruct from an audit trail, bound the replay, or declare the restatement | A last-updated column on an overwritten table does not help. It dates the value that survived; it does not hold the values that were destroyed. ## What the wrong attribute actually costs - **Numbers that no longer reconcile.** A monthly report regenerated from the replay disagrees with the one published that month, and neither side can say which is right without knowing the tier history. - **Silent reclassification.** Every customer whose attribute changed during the year is reclassified across their entire history in one direction - toward their current state. Aggregates shift systematically, not randomly, which is much harder to spot in a spot check. - **A comparison that cannot be trusted.** If the replay was produced to validate a logic fix, a stale lookup contaminates the comparison: differences from the reference data and differences from the fix are indistinguishable. ## The job's own configuration is reference data too The pattern is broader than a table. A threshold read from a settings file, a mapping constant, a list of excluded accounts, the rounding rule - each is a value in force at a time. If it changed in June, a replay applies today's value to January. Treat every externally supplied value the same way: either it is versioned and applied as of the record's time, or the divergence is stated. The lookup table is simply the largest and most visible instance. ## What varies between engines How the reference data reaches the workers changes what is even possible: - A job that loads a snapshot of the table at start and sends a copy to every worker process has exactly one version available for the whole run. Reproducing the past means loading the right snapshot, which requires a snapshot to exist. - A job that holds the reference data as state under its key, fed by a stream of changes to that table, is in a much better position: if the change stream is itself stored, the replay can read it interleaved with the records by timestamp, and the past reproduces faithfully. This is the only approach that is exact by construction rather than by preparation. - A job that queries an external service per record has no historical mode at all unless the service offers an as-of parameter, and per-record lookups over a year of history will also collide with the destination-rate problem from the other direction. ## Deciding, and saying so The decision is rarely purely technical. When the past genuinely cannot be reconstructed, the two honest outcomes are to narrow the replay to the period or the attributes where nothing changed, or to publish the replay as a **restatement** - numbers computed on today's classification, labelled as such, with the difference from the published figures quantified. What is not acceptable is shipping the replay into the same tables as the original output and letting consumers assume both were computed the same way. That is the failure this question is really testing for.
- The reference table is overwritten nightly and no history exists. What are the honest options?Three: reconstruct prior values from an audit trail or the change log of the source system if either exists; restrict the replay to a period or a set of attributes that demonstrably did not change; or publish the output as a restatement computed on today's classification, with the divergence from the original figures quantified. What is not acceptable is writing it into the same tables and letting consumers assume the method matched.
- How does this differ from the job's own configuration changing between the live run and the replay?It does not differ in kind. A threshold, a mapping constant or an exclusion list is a value that was in force at a time, exactly like a tier. If it changed since, the replay applies today's value to old records. Treat every externally supplied value as reference data: version it and apply it as of the record's time, or state the divergence.
Posting a year of held-back letters using this year's address book. Each one arrives where the person lives now, and the delivery records then state, quite confidently, that they lived there all along.
saying these in an interview costs you the question
- Reference tables are small, so re-reading them is safe
- The output must match the live run because the code is unchanged
- A last-updated column lets you reconstruct an overwritten table's past values
- Keeping the last few days of the lookup is enough for a year of history
- Treats the difference as acceptable drift and does not tell consumers