skip to content

Late in a long job, one worker process holding eight work slots dies — what is lost with it, and how does the process's size set that?

level: seniorimportance: should knowfreq 54%

answer

  1. everything in that address space
  2. all its slots, not one
  3. output others were still fetching
  4. size is the multiplier
  5. retry granularity is a different subject

basics

~20 s

Everything that lived in that process goes: all eight units of work in flight, from the beginning rather than from where they stopped, plus any intermediate output and cached data held only there. The process's size is the multiplier on all of it.

solid answer

~50 s

A worker process is the unit of loss. When one ends, the job loses the units of work in every one of its slots — each restarts from the start, since half-finished work has no portable form — along with any intermediate output it had produced and was still serving, anything it was holding in memory for reuse later, and for a job that remembers things per key, whatever portion it held since the last saved recovery point. The damage can reach past the process: other workers that were fetching its output now cannot. Size is the multiplier — eight slots means eight units lost at once, and a bigger pool means more held output gone — so fewer, larger processes fail less often and more expensively. Whether the job then retries the lost units, re-runs the round or restarts is a separate question about retry granularity.

go deeper

for a junior

Know that a worker process is the unit of loss: when one ends, every unit of work inside it is gone and starts again from the beginning.

for a middle

Enumerate what the process held beyond its in-flight work — produced output, cached results, retained per-key state — and connect the count of slots to the size of the loss.

for a senior

Show the propagation: consumers still fetching that output are affected too, so a late loss costs far more than the piece being computed, and say honestly what varies between engines.

for a principal

Set the shape policy from the tail rather than the average: decide which classes of job may run on the large shape given what one late loss costs the business.

## What lives inside one worker process A **worker process** is the process on a cluster machine that runs pieces of the job and owns the memory those pieces use. Over a long run it accumulates far more than the work currently in its slots. Understanding a process death means enumerating what was inside it: 1. **The units of work in flight** — one per occupied **work slot**, the concurrent units the process may run at once. A unit that was ninety per cent through its input is not ninety per cent saved; there is no portable half-finished form, so it starts again from the beginning. 2. **Intermediate output it had already produced** for a later step, which consumers may still be fetching. 3. **Anything materialised in its memory for reuse** by a later part of the job. A step written on the assumption that result was already computed now has to obtain it again. 4. **For a job that remembers things per key**, the portion of that retained state this process held, back to the last saved recovery point. All four are lost at once, because they lived in one address space and on one process's private scratch space. ## Size is the multiplier This is the part that belongs to the shape decision rather than to failure handling. At a fixed total capacity: | Shape | Units of work lost per death | Held output lost | Deaths expected | |---|---|---|---| | 4 processes x 16 slots | 16 | a quarter of the job's held output | fewer | | 16 processes x 4 slots | 4 | a sixteenth of it | more | The expected amount of redone work over a whole run is roughly the same to a first approximation — you lose more per event but have fewer events. What differs is the **variance and the shape of the damage**. One large loss late in a long run can force a visible stall, while a series of small ones is absorbed. That asymmetry is why long, expensive jobs on machines that come and go are usually run on the smaller shape, and short jobs on stable machines on the larger one. ## Why the damage can exceed the process The subtle point, and the one that distinguishes a senior answer, is that losing a worker is not always confined to that worker. Other workers that were fetching its intermediate output now find it unavailable. Depending on what the job can rebuild and from what, the effect can climb backwards into work that had already succeeded. The cost of a late loss is therefore not "the piece it was computing" but "everything downstream that still depended on what it held". ## Where designs differ Say what varies, because this family genuinely disagrees: - **Whether intermediate output dies with the process.** Where a **detached intermediate-output server** is in use — a separate long-lived process that keeps serving a finished worker's output after that worker is gone — the output survives the worker, and only in-flight work is lost. The oldest model in this family wrote intermediate results into storage outside the process altogether, so losing a process cost only the unit in it. Engines that keep output in the producing process's memory or private scratch lose all of it. - **How the lost work is rebuilt.** A finite job over re-readable input can rebuild a lost piece by running again the steps that produced it. A continuous job that remembers things per key resumes from its last saved recovery point and redoes everything after it. Neither is the universal answer. - **Whether the machine surviving helps.** If the process died but its machine did not, some designs can restart a process there and reuse state files left on local disk; others always rebuild from a remote copy. - **What "late" costs.** In a job whose recovery points are frequent, a late death costs little more than an early one. In a job with few or none, the cost climbs with elapsed time. ## Where this question stops The blast radius is the scope of this answer: everything that process held. What the job does next — retry only the lost units, re-run the whole round of work so nothing is half-finished, or restart from the last recovery point — is the separate subject of retry granularity, and the answer there depends on the engine and on the kind of job. An interviewer will usually accept, and often prefer, "that depends on the job's retry granularity" over a confident single rule. One boundary worth keeping straight while answering: this is the loss of a **worker**. The single **coordinating process** that plans the pieces, hands them out and tracks what finished is not a worker, and losing it fails differently — that is its own subject and should not be conflated with this one.

  • Does the job restart, re-run the round, or retry only the lost units?
    That is decided by the job's retry granularity, not by the worker shape, and the family disagrees: some engines re-run only the lost units, some re-run a whole round of work so nothing is left half-finished, and a job holding per-key state resumes from its last saved recovery point and redoes everything since. The blast radius tells you what was lost; the granularity tells you what is redone.
  • How can intermediate output outlive the worker process that produced it?
    Two ways exist in this class. A detached intermediate-output server — a separate long-lived process on the machine that keeps serving a worker's output after the worker is gone — decouples the output's lifetime from the producer's. Or the output is written somewhere outside the process entirely, which the oldest model in this family did by design. Neither is universal, so check before assuming finished work is safe.
  • Is the expected amount of redone work really the same for both shapes?
    To a first approximation over a long run, yes: larger processes lose more per death but die less often, and the products are comparable. The difference is in the tail. One large loss arriving late can stall a job visibly and push it past a deadline, while many small ones spread the cost smoothly. Choose by how much a bad single event costs you, not by the average.

saying these in an interview costs you the question

  • Says only the piece that was running is lost when a process dies
  • Assumes finished intermediate output always survives its producing process
  • Believes a half-finished unit of work resumes from where it stopped
  • Thinks losing a worker ends the job the way losing the coordinating process does
  • Claims a larger worker process makes each individual loss cheaper