A worker is lost five hours into a long run: when does that cost one re-run, and when must a whole region of the job rewind?
answer
- depends on where finished work lived
- durable elsewhere, or dead with the machine
- rebuilding can cascade upstream
- stateful jobs rewind a whole region
- cost set by recovery point age
basics
~20 sBlast radius follows where the lost results lived. Durable outside the machine: only the in-flight unit re-runs. On its local disk or in its memory: the producing units re-run too. Held as accumulated state: a whole connected region rewinds to its last recovery point.
solid answer
~50 sThere is no single answer, and saying so is most of the point. If each step materialises its output to storage outside any one machine, the dead worker takes only the units it was running, and those re-run elsewhere. If finished results were sitting in that worker's memory or on its local disk, a consuming unit can no longer fetch them, so completed units must run again — and rebuilding their inputs can cascade backwards until the chain reaches something still durable. If the job accumulates state between records, there is no independent unit to re-run at all: the lost worker's history cannot be reconstructed, so a whole connected region rewinds to the last **recovery point** — a durable copy of what the job would otherwise lose, written while it keeps running — with the input rewound to the read position that belongs to it. The shape of the job decides which, not the badge on the engine.
go deeper
Recall that one dead worker does not always cost one piece of work: what has to run again depends on where the results it had already produced were being kept.
Explain the three shapes and what distinguishes them, and be able to say why a finished unit can be forced to run again when nothing about that unit failed.
Bring numbers: how much input a rewind replayed, how far back a rebuild cascaded, and which boundary you would materialise at to cap it — including what that costs on every run that does not fail.
Own the trade directly: recovery time is bought with steady-state cost, and the right amount to buy depends on what a late result is worth to the business, not on what the runtime makes convenient.
## Blast radius is a property of the job, not a law of the engine "What does one dead worker cost?" has three honest answers, and which applies is decided by where the job's finished work was living at the moment the machine vanished. | job shape | what died with the machine | what has to run again | |---|---|---| | a finite job that materialises each step's output to storage outside any one machine | only the units in flight | those units, placed elsewhere | | a finite job holding finished results in worker memory or on the machine's local disk | the in-flight units *and* every finished result that lived there | the lost units, plus the units that produced the now-unreachable results, which can cascade upstream | | a long-running job that accumulates state between records | the portion of the state that worker held, which nobody else can reconstruct | a whole connected region, rewound to its last recovery point, with the input rewound to match | | a runtime serving an endless input as repeated small finite jobs — short bounded pieces, each run as a complete little job | the work inside the piece currently in flight | that piece, re-run; the bound is the piece, not the age of a saved picture | A candidate who names one row as "how recovery works" has told the interviewer which engine they learned on. A candidate who asks which row this job is in has answered the question. ## The second row: loss of an intermediate result This is the case people are least ready for. A downstream unit goes to fetch the partial output a now-dead worker had produced, and cannot. That producing unit had already completed, and the engine's own bookkeeping said so; the result is gone anyway. The engine therefore marks a *finished* unit as needing to run again — which means reading its input again — which, if that input was itself an intermediate result on another machine that has since died or been evicted, means rebuilding that too. The chain stops at the first thing still durable. In the good case that is one step back. Late in a long job with several redistributions behind it, it can be the original input. This is also why "the engine just rebuilds the lost piece from its **recorded derivation**" — the engine's note of which inputs and which steps produced that piece — is true about mechanism and misleading about cost. Rebuilding is cheap when the derivation is short and its inputs are still present. It is not cheap when the derivation reaches back through hours of work. ## The third row: rewinding a region Where the job accumulates state between records, a per-unit re-run is meaningless: the unit's output depends on everything it has seen, and that history died with the worker. The only sound move is to return to a moment when a coherent copy of the whole thing existed — the last recovery point — and replay the input from the recorded read position that belongs to it. How that coherent copy is captured across workers that never pause at the same instant is a mechanism of its own, owned by a sibling subject; what matters here is what it implies for blast radius: - **The recovery unit is not the retry unit.** One worker out of two hundred died, and two hundred workers rewind. Nothing smaller is coherent. - **The cost is set by the age of the recovery point, not by the size of the failure.** The job pays to replay everything that arrived since that point, across every worker. Taking recovery points more often shrinks the bill after a failure and raises the cost of every second the job runs; that trade is a separate subject. - **The input must reach back that far.** A rewind is only possible if the source still holds the records the job is about to re-read, which is its own precondition and its own subject. ## How much rewinds — and what genuinely varies - Some engines rewind **the entire job** on any worker loss: simple, coarse, and defensible when jobs are small. - Others rewind only a **connected region**: the set of steps linked by data in flight, bounded wherever results were materialised. That is strictly cheaper — but a region is only separable if the graph is actually cut somewhere. A pipeline that is one connected flow from source to destination has no smaller region to isolate, so region-level recovery buys it nothing. - Whether a lost worker is even replaced before recovery starts depends on the layer that supplies machines, and waiting for a replacement can dominate the recovery time on a busy cluster. ## How to answer it in a loop 1. Ask which row the job is in, and say why the row matters. 2. Name where finished work lives — durable shared storage, worker memory, local disk, accumulated state — because that is the actual variable. 3. Quantify: how much input would be replayed, and how long the last rewind took in practice. 4. Say what you would change: materialise at a boundary to cap the cascade, or take recovery points more often, and name what each costs on every successful run.
- Why can rebuilding one lost piece from its recorded derivation cascade backwards?Because the derivation names the inputs the piece was built from. If those inputs were themselves intermediate results held on machines that have since died or evicted them, they must be rebuilt first, and so on. The chain stops at the first thing still durable — one step back in the good case, the original input in the bad one. Materialising at a chosen boundary is how you cap it.
- What decides how much input a whole-region rewind has to reprocess?The age of the recovery point it rewinds to. Everything that arrived since then is replayed, by every worker in the region, regardless of how small the failure was. Writing recovery points more often shrinks that replay and charges the job continuously for the privilege; the capture mechanism and its tuning are a separate subject.
- Does every engine rewind only the affected region rather than the whole job?No, and it is not purely an engine choice. Some rewind the whole job by design. Others isolate a connected region, but only where the graph is cut at a boundary whose results were materialised — a pipeline connected end to end has no smaller region, so it rewinds entirely no matter what the runtime supports.
saying these in an interview costs you the question
- Assumes losing one worker always costs exactly the work that worker was doing
- Says a job restores from a saved picture regardless of whether it keeps anything between records
- Expects a lost intermediate result to be rebuilt without re-reading anything upstream
- Believes a copy saved on the dead machine's own disk still counts as a recovery point
- Treats a whole-region rewind as cheap because it only goes back to the last recovery point
- Claims a finite job must rewind its input source whenever any worker is lost