skip to content

A long-running job loses the coordinating process on one machine and a worker process on another — how do the two losses differ?

level: seniorimportance: should knowfreq 58%

answer

  1. which process was lost decides everything
  2. a worker is one of many
  3. the coordinating process is a singleton
  4. plan and progress record exist once
  5. standby only where the engine offers one

basics

~20 s

Losing a worker process costs the pieces it was running and the intermediate output it still held; the rest of the job continues. Losing the coordinating process leaves nothing planning, handing out or tracking work, so the job stops.

solid answer

~50 s

A worker process — the process on a cluster machine that runs pieces of the job and owns their memory — is one of many, and the job is designed around losing one: the pieces it was running are produced again, and the surviving workers carry on. Whether its already-finished intermediate output is also lost depends on the cluster: some run a separate process that keeps serving a finished worker's output after that worker is gone, and some do not. The coordinating process is a singleton, and the plan, the assignment table and the record of what finished exist only in it. Some engines offer a standby that can take over from a progress record kept outside the process; where that is not offered or not arranged, its loss ends the job. Whether the job then re-runs one piece or restarts from a saved point is a separate subject.

go deeper

for a junior

Recall the asymmetry: a worker process is one of many and the job is built to lose one, while the coordinating process is the only one of its kind and the job does not go on without it.

for a middle

Explain what each loss actually takes with it — in-flight pieces and held intermediate output for a worker, the plan and the completion record for the coordinating process — and why nothing on the workers can rebuild the second.

for a senior

Show production judgment: name the intermediate-output question as what decides the cost of a lost worker, qualify the coordinating loss with the standby case, and hand the retry-or-restart decision to where it belongs instead of improvising a policy.

for a principal

The tradeoff to own is how much the platform should spend removing this single point of failure: a durable progress record and a standby cost latency and operational surface on every job, and are worth it only for jobs whose length or state make a restart genuinely expensive.

## Two losses, two blast radii The question is testing whether you treat processes in a distributed job as interchangeable. They are not: one of them is a singleton holding state nothing else has a copy of. | what was lost | what goes with it | what continues | |---|---|---| | one **worker process** — a process on a cluster machine that runs pieces of the job and owns their memory | the pieces it was running mid-flight, the memory those pieces held, and the intermediate output it was still serving to other workers | every other worker process; the job continues if the lost work can be produced again | | the **coordinating process** — the single process that plans the pieces, hands them out, tracks what finished and receives what the program brings back | the plan, the assignment table, the completion record, and anything the program had already brought back to one place | nothing that matters: in-flight pieces may run to the end, but nothing gathers their results or hands out the next round of work | A **round of work** here means one wave of pieces that must all run before the next wave can start — it is the unit the coordinating process is tracking when it dies. ## Why a lost worker process is usually survivable - **Lost-work recomputation.** The standard remedy is to rebuild a lost piece of intermediate output by running again the steps that originally produced it. The cost is machine-hours and elapsed time, not correctness. - **Or a saved recovery point.** A job that retains long-lived per-key state usually resumes that state from its most recent saved point rather than recomputing from the source, because recomputing hours of accumulated state from the input is not a trade anyone wants. - **What actually varies is the already-finished output.** If a detached intermediate-output server holds it, losing the worker does not by itself lose it, and the surviving work stands. If nothing does, the output went with the process and other workers that were still fetching it lose their source. This is the difference between a lost worker being a minor cost and a lost worker being expensive, and it is a property of how the cluster is arranged, not of the engine class. - Whether the job then re-runs just that piece, re-runs a whole round, or restarts entirely is a different subject with its own tradeoffs; stop your answer at the blast radius. ## Why the coordinating process is different in kind 1. **It holds no shard of anything.** Every other process in the job holds a slice of a divisible whole. This one holds the only copy of the plan and the only record of what has finished. 2. **The record of progress is in its memory unless something deliberately put it elsewhere.** Nothing on the workers can reconstruct which of ten thousand pieces had reported success. 3. **The workers do not elect a replacement.** Importing leader election from replicated stores is a common and wrong instinct here: worker processes have no copy of the plan to be elected custodian of. 4. **A standby exists on some engines, not as a property of the class.** Where one is offered, it works by keeping the job's progress in a durable store outside the process, so a replacement can read it and resume rather than start over. Where it is not offered, or not arranged for this job, the loss ends the run. 5. **Placement changes the exposure.** If the coordinating process runs inside the cluster beside the workers, it enjoys whatever the cluster's own machines enjoy. If it runs outside on the machine that submitted the job, the job's availability is that machine's availability — a closed laptop is a lost job. 6. **Restarting a dead process somewhere else is the platform's concern**, not the engine's, and is a separate subject from what the job loses. ## What to say, in order 1. Name which process was lost, because the answer is completely different for the two. 2. State the blast radius: everything that process held, and nothing else. 3. For the worker: say the lost pieces are produced again or resumed from a saved point, and note the intermediate-output question as the thing that decides how expensive it was. 4. For the coordinating process: say the job stops making progress, and qualify with the standby case rather than claiming it is always unrecoverable. 5. Resist finishing the sentence with a retry policy. Whether one piece or the whole job is redone is a decision with its own owner, and an interviewer who asked about blast radius did not ask about that.

  • The lost worker process had already finished several pieces. Is that output still usable?
    It depends on where it lived. If a detached intermediate-output server is running — a separate process that keeps serving a finished worker's output after the worker is gone — then yes, and the loss costs only the in-flight pieces. If the output lived in the worker process itself, it went with it, and any worker still fetching from it must wait for that output to be produced again.
  • Why do workers not simply elect a new coordinating process among themselves?
    Because election solves the wrong problem. The difficulty is not choosing who leads but that the plan, the assignment table and the completion record exist in exactly one place; no worker holds a copy to be elected custodian of. Engines that survive this loss do so by writing the job's progress to a durable store outside the process first, and election, where it appears at all, only decides who reads it.
  • Does a continuous job or a finite job come off better when this process is lost?
    Often the continuous one, because it is the kind of job most likely to have been arranged with saved recovery points and a standby, having been expected to run for months. A finite job usually just runs again from the start or from whatever intermediate output survived — annoying, but bounded by the job's own length rather than by months of accumulated state.

saying these in an interview costs you the question

  • Says the workers finish the job on their own once coordination is gone
  • Treats the loss of any one process as the same event
  • Assumes a finished worker's intermediate output is always still fetchable
  • Believes every engine keeps a standby coordinating process ready
  • Counts the coordinating process's loss as one piece of work to redo
  • Expects surviving workers to elect a replacement planner