A job crashed while writing its results and restarted from its last recovery point — why can the destination hold some rows twice?
answer
- recovery is reprocessing
- the destination was never told
- internally consistent, externally doubled
- append adds; a keyed write replaces
- advance the read position last
basics
~10 sRecovery is reprocessing: a restart re-reads input the job had already handled and writes it a second time. The destination remembers nothing of the first attempt, so a plain append adds those rows again.
solid answer
~60 sA restart resumes from a recovery point — a durable copy of everything the job would otherwise lose, written while the job kept running — and from the recorded read position stored with it, the marker saying how far into the input the job had durably got. Everything between that point and the crash is processed again, because recovery *is* reprocessing. The rows that already reached the destination were never part of what the job saved, so nothing the job restores knows about them. Whether the repeat is visible is decided entirely at the destination: an append adds a second copy, a write keyed by a value derived from the input record replaces the first copy, and a result staged privately and published in one final act never showed the abandoned attempt at all. How much is redone varies by model — a job holding nothing durable recomputes just the lost piece, a finite job usually re-runs the failed unit, a long-running stateful job rewinds to its last saved picture.
go deeper
Recall the chain: a crash means work after the last recovery point is done again, the same records are written again, and the destination has no memory of the first write. Duplicates are the normal outcome, not evidence of a bug.
Explain why the job's own restored accounting and the destination's contents are separate claims, and why advancing the recorded read position only after the output is durable trades silent loss for a visible duplicate.
Show that you can say, for a specific job, exactly what gets re-emitted after a Tuesday-afternoon restart — which model it recovers under, how far back it rewinds, and which destination absorbs the repeat and which does not.
Frame it as a property of each destination rather than of the pipeline: which effects must be repeat-proof, which can absorb duplicates, and what the organisation is actually buying when it pays to move one of them from the second class to the first.
## What a restart actually redoes Two durable things let a job resume. The first is a **recovery point**: a durable copy of everything the job would otherwise lose, written while the job keeps running, so a restart can resume from it instead of from the beginning of the input. The second is the **recorded read position**: the marker saying how far into the input the job has durably got, kept so a restart knows where to resume. When the job comes back, it resumes from both — and everything the job did after that point it now does again. That is not a defect in the mechanism; it *is* the mechanism. Recovery is reprocessing, and reprocessing means the same input records travel the same operators a second time. How much is redone is not one number, and it is the first thing to establish about any job: | Job shape | What a lost worker costs | What the destination may see | |---|---|---| | Holds nothing durable between records | The lost piece is rebuilt by re-running its recorded derivation — the engine's note of which inputs and which steps produced it | Only the output of the rebuilt piece, if that piece had already written anything | | A finite job over bounded input | The failed unit of work is re-run on another machine | Whatever that unit had written before it died | | Long-running, with accumulated state | The saved picture is restored and the input is rewound to match it | Everything written between the recovery point and the crash | Any of the three can produce a duplicate; they differ in how much. ## The destination sits outside the job's accounting The rows that left the process before the crash are not part of what the job saved. No recovery point mentions them, and restoring one cannot retract them. This is why two sentences that sound contradictory are routinely true at the same instant: - for a job that retains state, **the job's own accumulated totals can reflect each input record exactly once** after recovery, because the engine saved and restored them coherently; - and **the destination can hold a row twice**, because it is a different system that was written to and never told. A job that retains nothing has no internal accounting to be consistent about at all — for it, the only question that was ever interesting is the one at the destination. ## What decides whether the repeat is visible The engine cannot decide this. The write does: - **A plain append** — a second copy appears. Counts, sums and row totals downstream all move. - **A write keyed by a value computed from the input record** — the second write lands on the same key and replaces the first row, so the destination ends in the state it would have reached had the work run once. This holds only while the operation replaces; a statement that adds to a running total is not made safe by a stable key. - **A staged result promoted in a single step** — each unit writes to a private location and one final act makes the whole result visible, so a failed attempt's partial output is never visible to anyone. How cheap that final act is varies with the storage: where it is a directory move it costs nothing, and where publishing means copying many objects the atomic act has to be a single small pointer or manifest write instead. - **An effect that leaves the system** — a payment, a message to a third party, an alert. There is nothing to overwrite, and this is the case where duplicates hurt most. ## The rule that turns duplicates into the acceptable failure There are two orderings available, and only one of them is survivable. If the recorded read position is advanced **before** the output it covers is durable, a crash in between loses that output permanently and silently — the restart resumes past records it never actually wrote. If the position is advanced **after**, a crash in between repeats output. Every design in this area chooses the second and then works on making the repeat harmless, because a duplicate is visible and repairable and a silent hole is neither. ## How to answer this in an interview Say the three things in order: recovery re-reads, the destination was never told, and therefore the fix lives on the writing side — a key derived from the record, a single act that publishes output and position together, or deduplication pushed onto whoever reads. Then name which recovery model the job in the question is on, because the size of the repeat follows from it. A candidate who says "the engine handles it" has not understood that the engine's guarantee stops at its own process boundary.
- If the job's own accumulated totals are correct after recovery, why is that not enough?Because it is a claim about one system. The accumulated totals live inside the job and were saved and restored together; the rows at the destination left the process and were never part of that save. Internal coherence says nothing about what a destination already written to now holds — and a job that retains nothing has no internal accounting at all, only the destination question.
- Would recording the read position before writing the output avoid the duplicate?It would, and it replaces a duplicate with permanent silent loss. A crash between the two leaves the restart resuming past records whose output never became durable, and nothing downstream can tell those rows are missing. The accepted ordering is the other one: advance the recorded read position only after the output it covers is durable, then make the repeat harmless.
- Does it matter that the crash happened in one worker rather than the whole job?It changes the size of the repeat, not its nature. Losing one worker may cost only the units it held, while a whole-job restart from the last recovery point redoes everything after that point. Either way the destination sees whatever the lost work had already written, and neither case lets the job retract it.
You post a cheque, then your notebook is restored from yesterday's backup. The notebook now says you have not paid — but the recipient is holding the cheque. Re-sending on the notebook's word pays twice: the restored record and the effect already out in the world are two different systems, and only the recipient can refuse a second cheque.
saying these in an interview costs you the question
- Thinks the recovery point stores the output, so restoring it un-writes rows
- Believes the destination recognises rows it has already received
- Assumes a correct internal recovery implies a correct destination
- Says duplicates mean the engine's recovery is broken
- Thinks recording the read position first is the safer ordering
- Treats every job as recovering the same way, from a saved picture