Why is a job's recorded read position advanced only after the output it covers is durable, and what breaks if it is advanced first?
answer
- there is always a window
- the order picks your failure mode
- loss versus repetition
- only one of them is repairable
- durable at the destination, then advance
basics
~20 sThe ordering chooses which failure you get. Advancing the recorded read position first means a crash in the gap loses records nobody will ever reprocess. Writing output durably first means a crash re-does covered work — repetition, which the destination can be made to absorb.
solid answer
~50 sThe **recorded read position** is the durable marker of how far into the input the job has got, and a restart resumes from it. Between producing output and advancing that marker there is always a window, and the order decides what a crash inside it costs. Advance first, then crash before the output is durable: the restart resumes past records whose results were never written, so they are silently dropped — loss with no signal, and the input is usually gone by the time anyone notices the numbers. Write first, then crash before advancing: the restart resumes at the older position and redoes work that was already written, so the destination sees the same effect twice. Repetition is the failure you can engineer against; silent loss is not. That is the whole reason for the rule, and it is why most of these systems are at-least-once underneath — never fewer than once, possibly more.
go deeper
Remember the sequence: produce the output, make it durable, then record how far you have read. Never the reverse, because a crash after recording but before writing skips those records forever.
Explain the window between the two steps and name what a crash inside it costs under each order — silent loss one way, repetition the other — and why repetition is the one a destination can absorb.
Demonstrate that you check where the position is actually stored and whether it can disagree with the durable output, and that you can say how much input a restart of this job re-reads and what the destination does with it.
Argue the economics: a pipeline that can lose records quietly has no reconciliation story and no way to bound the error, so paying for repetition and absorbing it at the destination is the cheaper guarantee to own.
## Two orders, two different failures A job consumes input, produces output, and records how far it has got. The marker it records is the **recorded read position** — durable, so a restart knows where to resume, and it is the only thing standing between a restart and re-reading the input from the beginning. These steps cannot be made simultaneous by wishing. There is a window between "the output is durable" and "the position is advanced", and a crash can land inside it. The design question is which of the two you put first, because that choice picks the failure mode you will live with. | Order | Crash inside the window | Consequence | |---|---|---| | Advance the position, then write the output | the position already says those records are handled; the output never landed | **Silent loss.** The restart resumes past them, nothing re-reads them, and no counter anywhere records that anything is missing. | | Write the output durably, then advance the position | the output landed; the position still points behind it | **Repetition.** The restart re-reads those records and produces their output a second time. | ## Why only one of them is recoverable Repetition is a problem with a known set of answers at the destination: a write keyed so that a re-run computes the same key and overwrites rather than appends, a destination transaction that replaces a whole range, a final promotion step that makes one version of a result visible. The mechanics of making a repeat harmless belong to the destination side of this subject, but the point here is only that **they exist**. Silent loss has no answers. The records are not merely mis-processed; they are not reprocessed at all, and the input that held them may no longer offer them by the time an analyst notices a dip in a weekly number. There is no reconciliation you can run from inside the job, because the job's own position says everything is fine. This asymmetry — one failure mode repairable, the other not — is the entire justification for the rule, and it is the answer an interviewer is listening for. This is also what people mean when they say these systems are **at-least-once** underneath: the pipeline guarantees never fewer than once, and the count above one is what the destination is made to absorb. ## The ordering rule, stated precisely > The recorded read position may advance to cover a range of input **only once every effect that range produces is durable.** Two things follow that people routinely get wrong: 1. **"Durable" means durable at the destination, not inside the job.** Output buffered in a worker's memory, or staged in a location nothing has promoted yet, does not license the position to move. A common bug is advancing on the strength of a successful local write that a later step still has to publish. 2. **The position is per range, not per record.** Jobs advance it in units — whatever range a recovery point or a completed output covers — so the amount of repetition after a crash is bounded by that unit's size, not by one record. ## Where the position actually lives, and why that varies Runtimes in this family do not agree on where the marker is kept, and the difference matters when you are debugging a resume: - Written together with **a recovery point** — a durable copy of everything the job would otherwise lose, taken while it keeps running — so the state and the position it corresponds to are one artefact and cannot disagree. - Handed back to the source system, which keeps the marker on the job's behalf. The mechanics of that are the source system's own subject, but the ordering rule is unchanged. - Derived after the fact from what is visible at the destination, when the output itself records which range it covers. A job whose state and whose position are stored in two places, written at two moments, has reintroduced exactly the window the rule is about — one step further along. ## Closing the window rather than choosing a side The window can be removed, not just ordered, by **committing output and read position as one unit**: arranging that the write becomes visible and the position advances together, so no crash can leave one done and the other not. Whether that is available is a property of the destination, not of the engine — a destination with transactions or an atomic promotion step can do it; an arbitrary remote endpoint cannot. Where it is not available, you keep the ordering rule and make the repeat harmless instead. ## What this is not It is not a statement about the source system's own bookkeeping, and it is not the general vocabulary of delivery semantics. It is one ordering constraint inside the job: **durable first, position second** — because that trades an unfixable failure for a fixable one.
- What if the job writes output in several places before advancing the position?Then the position may advance only after the last of them is durable, and a crash in between leaves a partial effect that the re-run repeats. Either make every one of those writes absorb a repeat, or reduce them to a single promotion step that publishes everything at once, so there is one moment to order against.
- How much repetition does a restart cost under this rule?Whatever the position advance covers — the range of input consumed since it last moved. Advancing more often shrinks the repeated range and the catch-up after a restart, at the cost of more frequent durable writes. That frequency is the knob; the ordering is not.
- Does this rule apply to a finite job over stored files?The same logic applies with different bookkeeping: a unit is not recorded as done until its output is durable, which is why such jobs stage each unit's output privately and promote the result in one final act. If done-ness were recorded first, a crash would leave a unit permanently skipped.
saying these in an interview costs you the question
- Advances the recorded read position as soon as the records are read.
- Says the order does not matter because both cases lose the same work.
- Treats output buffered inside the job as durable enough to advance on.
- Claims repetition is the worse failure because duplicates corrupt the result.
- Thinks a faster save interval removes the window rather than shrinking the replay.