A job re-reads last month's records from its append-only source, but the enrichment attaches today's prices - what in the join caused that, and what makes the re-run reproducible?
answer
- the input rewound, the reference did not
- key-only match sees only now
- versions, retained versions, time predicate
- how far back is reproducible
- deleted keys and backdated corrections
basics
~20 sThe join takes the reference value current at processing time, and the copy holds only current values, so old records are enriched with today's numbers. Reproducibility needs a reference that retains each version with its validity span and a predicate on the record's own moment.
solid answer
~50 sRewinding the input reproduces the records; it does not rewind the reference. If the join matches on the key alone against a copy holding one current value per key, then every record - whatever its moment - is enriched with whatever is true today, so last month's orders come out priced at today's prices. Nothing errors, because the match succeeds. Three things together make the re-run reproduce the first run: the reference must retain each version with the span over which it was in force; the copy the job holds must retain those versions too, at least as far back as the oldest record that may be re-read; and the join predicate must select on the record's own moment rather than on the row marked current. Without all three, the output of this job is a function of when you ran it.
go deeper
Recall that rewinding the input brings back the records but not the reference, so a join on the key alone enriches old records with today's values.
Explain the three conditions together - retained versions in the reference, retained versions in the job's copy, and a predicate on the record's own moment - and why any one of them missing makes the output depend on when you ran it.
Demonstrate the operational edges: the retention horizon and the silent fallback past it, keys deleted since, a change feed that only has history from when it started, and a sample diff against the original output as the cheap check.
Own the promise: how far back this pipeline's output is reproducible, what retaining that history costs, and what is said to the people who will one day recompute a figure from beyond the horizon.
## What was actually replayed A rewind restores one of the join's two inputs. The records come back exactly as they were - same values, same moments, same keys. The reference did not come back: it is a live store that has been moving the whole time, and the job's copy of it holds whatever is true now. So the re-run is not the original computation at all. It is the old records joined against a new reference, and the join has no way to notice, because a key lookup succeeds either way. This is the most expensive failure on this leaf precisely because it is silent. There is no error, no dropped record, no gap in the output. There is a settlement file, a restated report or a repaired day's data in which every figure is internally consistent and historically wrong. ## The three conditions for a reproducible enrichment For the same records to produce the same enriched output whenever the job is run, all three of these must hold: 1. **The reference retains its history.** One row per key per span of validity, each carrying the moment it took effect and the moment it ceased to apply. If the reference overwrites in place, the information needed to reproduce the first run no longer exists anywhere, and no change to the job can recover it. 2. **The job's copy retains those versions too**, back at least as far as the oldest record that might be re-read. A copy kept current by a change feed accumulates versions naturally as changes arrive, but only from the moment the feed started; one kept by wholesale reload holds only what the reference currently shows. 3. **The predicate selects on the record's own moment** - the moment stamped by the source when the thing happened - rather than on a current flag or on the moment a worker reached the record. Matching on the processing moment is the same bug in a different costume: it moves whenever the job's speed moves. Miss any one and the output is a function of when the job ran. ## The retention horizon, and its cliff Condition 2 has a sharp edge that is worth naming explicitly. Version history is retained for some finite period, because it grows with change volume. That period is a promise: *output is reproducible for records no older than this*. Cross it and the behaviour is not an error but a quiet fallback - the oldest version still retained becomes the answer for every record older than it, so the further back you re-run, the more of the output is enriched from one arbitrary boundary version. A team that has a retention horizon and does not know its number is a team that will discover it during an audit. The same applies to deletes. A reference key retired since the records were produced is simply absent from a reloaded copy. Whether the enrichment then drops the record, emits it with empty fields, or fails is a design decision that has to be made deliberately rather than discovered on the re-run. ## What else moves between the runs | what the re-run restores | what it does not | |---|---| | the records, their keys and their moments | the state of the reference store | | the order within a key, where the source preserves it | the job's copy of the reference and its version history | | the enrichment logic, if unchanged | reference rows deleted or retired since | | the record's own moment | values corrected or backdated since the first run | That last row is the subtle one. Even a correctly written as-of match can disagree with the first run if a version was *corrected after the fact* - a price fixed retroactively. Both runs matched as-of and both were right by their own reference; the reference itself changed its mind about the past. Distinguishing when a version was true from when it was recorded is what makes that reproducible, and modelling it belongs to a data-modelling subject with its own owner. What belongs here is knowing that "as-of" alone does not guarantee agreement across runs unless the reference's history is itself immutable. ## Practical checks before a re-run - **Ask what the join predicate is.** If it has no clause over time, the re-run's numbers are today's numbers; stop there. - **Ask how far back the reference's version history goes**, and compare that to the oldest record being re-read. - **Ask what happens to keys that no longer exist**, and decide it rather than observe it. - **Ask which clock the predicate uses.** A predicate on the moment a worker reached the record looks like an as-of match and is not one. - **Compare a sample of re-run output against the original output**, on records you expect to be identical. A diff of a few thousand rows finds this in minutes; a reconciliation three weeks later finds it expensively. ## Where engines come into it Barely, and that is worth saying to an interviewer. Rewinding an input is available across this class in one form or another, and every engine can express a predicate over a validity span. The failure is not in the runtime: it is that the enrichment was written against a reference with no history and no time predicate, and that nothing in any engine will tell you so. The one runtime-flavoured wrinkle is *when* the copy is fixed - a runtime advancing one record at a time can apply a version change between two consecutive records, while one that collects arrivals for a short span holds the copy fixed for that span - and that difference changes which records land on which side of a change during a live run, not during a re-run.
- The copy is kept current by a change feed. Is the re-run safe?Only back to the point the feed started and only if the job kept the versions rather than overwriting each key as changes arrived. A feed reports changes from the moment it began, so nothing before that exists in the copy, and a copy that applies each change in place holds exactly one value per key - current, not historical.
- Both runs used a correct as-of predicate and still disagree. What can explain it?The reference's history changed between the runs - a version corrected or backdated after the first run read it. Both runs matched the version in force at the record's moment, according to the reference as it stood at the time. Making that reproducible needs the reference to distinguish when a version was true from when it was recorded.
- How would you detect this problem before anyone downstream does?Re-run a small, old slice and diff the enriched fields against the original output for records you expect to be identical. Any difference in an attribute that should be historically fixed is the signature. It costs minutes and catches the failure while it is still a sample rather than a restated report.
- What should the join do with a reference key that has since been deleted?Decide it explicitly: drop the record, emit it with the enrichment fields empty and a marker, or fail the run. All three are defensible; what is not defensible is discovering the behaviour during a re-run. Where the reference retains versions with their spans, a retired key usually still has a span covering the record's moment, which removes most of the problem.
saying these in an interview costs you the question
- Blames the source or the replay order rather than the join predicate
- Thinks rewinding the input also restores the reference store
- Assumes a change feed gives history from before the feed started
- Says an as-of predicate alone guarantees reproducibility, whatever the copy retains
- Has no answer for reference keys deleted since the records were produced
- Treats the version retention horizon as unlimited