skip to content

A job crashes mid-run and restarts. Why does its recovery depend on the input being re-readable from a recorded position?

level: juniorimportance: must knowfreq 72%

answer

  1. recovery means doing work again
  2. you cannot invent lost records
  3. the input must still be there
  4. resume from a position you named
  5. a forgetful source caps recoverability

basics

~20 s

Recovery is reprocessing: a restart re-reads input it already consumed. So the source must retain those records and let a reader resume from a durably recorded position — a source that forgets each record after handing it over makes the lost work unrecoverable.

solid answer

~50 s

Every recovery mechanism in this class works by doing some work a second time; none of them can invent records a dead machine had already pulled in. Which work is redone varies with the job. A job that keeps nothing between records rebuilds only the lost piece from its **recorded derivation** — the engine's note of which inputs and which steps produced it. A finite job usually re-runs only the failed **unit of work**, the smallest thing the engine hands to one worker. A long-running job with accumulated state restores a saved copy of that state and continues forward from the position saved with it. All three then need the same thing from the source: the records that the redone work covers must still exist and must be reachable from a position the reader can name. That property belongs to the input, not to the engine, and no engine setting can supply it.

go deeper

for a junior

Recall the one-line chain: a crash means work is done again, doing it again means reading those records again, so the source must still have them and let you resume from a position you recorded.

for a middle

Explain that the three recovery models re-read different amounts — a lost piece rebuilt from its derivation, one failed unit, or everything since a saved copy of state — and that all three fail identically against a source that forgets.

for a senior

Show that you audit the input's replayability before you promise anything about a pipeline's behaviour across failures, and that you can name the segment of the path that is unrecoverable when the source offers no way back.

for a principal

Frame it as what the organisation is buying: a pipeline whose source cannot be rewound has an availability story but no correctness story, and closing that gap is a platform investment, not a job setting.

## Recovery is reprocessing, not restoration When a machine dies in the middle of a distributed run, nothing hands the job back the work that machine was doing. Underneath, every recovery mechanism in this class of system does the same thing: **some work is performed a second time**. The engine's only real choices are *how much* work is redone and *from where* it starts again. It cannot reconstruct records that a **worker** — one process on one machine, running some of the job's units and holding whatever that job keeps in memory — had already consumed and not finished with. That is why recoverability is a property of the **input** before it is a property of the engine. If the records the redone work needs are still obtainable, the engine has options. If they are not, the only honest outcome is a gap in the result, and no configuration anywhere in the cluster closes it. The input property this rests on is **a retained, re-readable input**: a source that keeps its records for some period and lets a reader start again from a position it names, as opposed to one that hands each record over once and forgets it. ## The three models, and what each of them re-reads There is no single recovery model in this family, and a sentence that is true of one is routinely false of the next. What they share is the dependence on re-readable input; what differs is how much input and from where. | Recovery model | What is redone after a loss | What it needs from the input | |---|---|---| | Recompute from the **recorded derivation of a piece** | just the lost piece, rebuilt by running its derivation again | the input records feeding that piece, re-readable at that position | | Re-run the failed **unit of work** | one unit, on another machine | that unit's slice of input — or a durable intermediate file that already stands in for it | | Restore a saved copy of accumulated state (**a recovery point** — a durable copy of what the job would otherwise lose, written while it keeps running) | everything consumed after that copy was written | the input from the position recorded alongside that copy, forward | Notice the third row is the expensive one, and it is the one people wrongly generalise. A job holding nothing between records has no saved picture to restore; an engine that materialises each phase to durable storage before the next phase reads it can often re-run a unit from those intermediate files and never touch the original source at all. Writing any one of the three as "how recovery works" is the classic error. ## What "re-readable" actually requires It is a stronger property than "the data still exists somewhere": - **A nameable position.** The reader must be able to say *start again here*, and that name must survive the restart. That is the **recorded read position** — the durable marker of how far into the input the job has got. - **Stability.** A second read at the same position must return the same records. A source whose contents are edited in place is re-readable and still useless, because the recomputed answer disagrees with the crashed run's. - **Availability for long enough.** The records must outlive the worst gap between a failure and the restart that repairs it. How long a given source keeps them is that source system's own subject, not the engine's. - **Re-reading without the source's permission.** A source that only ever moves forward, one hand-over per record, offers no way back regardless of how healthy it is. ## Sources that cannot be rewound A push feed, a socket, a sensor stream, an in-memory hand-off — these give the job one look at each record. A job over such an input can still be *available* (it can be restarted and keep running), but it cannot be *correct across a failure*: whatever was in flight is gone. The repair is not an engine setting; it is to place something durable and re-readable in front of the source and read from that instead. ## Where this stops A re-readable input makes recovery *possible*. It does not make it *invisible*: doing work again means the destination may see the same effect twice, and arranging that a repeat leaves no extra trace at the destination is a separate mechanism. Nor does re-readability say anything about the source system's own retention, ordering or delivery behaviour — those are properties of that system. What this question owns is the one precondition beneath all of it: **if the input cannot be re-read from a position the job recorded, the job has no recovery story at all.**

  • If a job keeps nothing between records, does it still need a recorded read position?
    A continuous one does: nothing else says where it got to, so without a durable position a restart either re-reads from the beginning or skips forward blindly. A finite job over stored files often does not need an explicit one — which units finished is already recorded, and the re-run of a unit re-reads its own slice.
  • Does re-readable also have to mean unchanging?
    Yes, if the results have to agree. A second read at the same position must return the same records; otherwise the recomputed piece differs from the one the crashed run produced, and two parts of the same output are derived from different input. A source that is appended to is fine; one edited in place is not.

A live broadcast you did not record cannot be listened to again; a tape can be wound back to a counter reading. Recovery needs the counter reading and the tape — the engine supplies neither.

saying these in an interview costs you the question

  • Says the engine can rebuild lost records without re-reading any input.
  • Thinks a saved copy of job state removes the need to re-read anything.
  • Assumes every source can be rewound, so replayability never needs designing for.
  • Treats restoring a snapshot as the only recovery model there is.
  • Believes restarting the process is itself recovery of the work.