What happens to pending futures when a ProcessPoolExecutor worker is killed?
answer
- One worker's death is not a local event
- The channel state cannot be resynchronized
- Every outstanding future gets the same exception
- Recovery lives outside the executor object
- Contrast: the other pool hangs instead
basics
~20 sEvery pending and running future is completed with BrokenProcessPool, and every later submit raises it too. The executor is permanently unusable, so recovery means building a new one and resubmitting the work that did not finish.
solid answer
~50 sWhen a `concurrent.futures.ProcessPoolExecutor` worker disappears abruptly — an out-of-memory kill, a segfault in a native extension, an external `kill` — the executor's management thread notices the dead process and declares the whole pool broken. Every future that was running or still queued is completed with `concurrent.futures.process.BrokenProcessPool` (a subclass of `BrokenExecutor`, which is a `RuntimeError`), and every subsequent `submit()` raises the same exception. There is no per-task recovery inside that executor: the pool is poisoned as a unit, because the executor cannot know which shared state the dead worker corrupted or which task killed it. Recovery is therefore structural — catch the exception at the batch level, construct a fresh executor, and resubmit the units of work that did not report a result. That only works if you tracked which units completed and made them safe to run twice.
code
python · 26 linesimport os
import signal
from concurrent.futures import ProcessPoolExecutor
from concurrent.futures.process import BrokenProcessPool
def die(_):
os.kill(os.getpid(), signal.SIGKILL)
def double(x):
return x * 2
if __name__ == "__main__":
executor = ProcessPoolExecutor(max_workers=2)
doomed = executor.submit(die, 1)
try:
doomed.result()
except BrokenProcessPool:
print("in-flight future poisoned")
try:
executor.submit(double, 21).result()
except BrokenProcessPool:
print("executor unusable; build a new one")
executor.shutdown()go deeper
Know that a process pool worker can die from outside your code and that the failure shows up as an exception on the futures, not as a silent wrong answer. Recognise the exception name when you see it in a log.
Explain that concurrent.futures.process.BrokenProcessPool is set on every outstanding future and on later submissions, and say why: the in-flight result is unknowable, the shared queues may hold a partial message, and the cause would likely kill a replacement too.
Show the recovery design you have actually run: checkpointed progress so a retry does not redo the whole batch, idempotent units, a fresh executor per attempt with bounded retries, bisection to isolate a poison record, and pool sizing driven by measured memory.
Own the blast radius. Decide which stages get their own isolated pool, whether a stuck job or a failed job is worse for this pipeline, what the retry budget is, and how a repeated pool break escalates to a human rather than burning a nightly window.
### The scenario An ad-auction bidder scores candidate bids in a `ProcessPoolExecutor` as part of a nightly batch that takes about 27 minutes. One record has a pathological payload; the worker handling it balloons and the kernel's out-of-memory killer takes that process out. What the operator sees is not one failed record — the entire remaining batch fails at once with `BrokenProcessPool`, including futures for records that had not been touched yet. ### Why one death poisons everything The executor keeps a management thread in the parent that watches the worker processes and the queues feeding them. A worker that exits without being asked to is an unrecoverable event from the executor's point of view, for three reasons. First, **the result is unknowable**. The dead worker was mid-task and sent nothing back, so that future can never be completed with a value. Second, **the channel may be corrupt**. Workers pull from a shared call queue and push to a shared result queue; a process killed mid-write can leave a partial message in the pipe, and there is no way to resynchronize a stream of pickled frames. Third, **the cause is unknown and probably recurs**. Whatever killed the worker — a memory limit, a native crash — will very likely kill the replacement, so silently rotating in a new process would produce an infinite kill loop rather than progress. So the executor takes the safe route: it marks itself broken, sets `BrokenProcessPool` on every outstanding future, and refuses further submissions. `BrokenProcessPool` subclasses `BrokenExecutor` which subclasses `RuntimeError`, so a bare `except RuntimeError` around your batch will catch it, though catching it by name is clearer. ### The contrast that makes the design legible `multiprocessing.Pool` chooses the opposite trade-off, and the difference is the sharpest way to remember both. When a `Pool` worker is killed, the pool quietly starts a replacement and keeps serving new tasks — but the task the dead worker was running is simply lost. Its `AsyncResult.get()` never returns. So the same crash gives you a loud, total failure with `ProcessPoolExecutor` and a silent hang with `multiprocessing.Pool`. Neither is "safe by default": one costs you the batch, the other costs you the ability to notice. In practice this is a strong argument for `ProcessPoolExecutor` in any pipeline where a stuck job is worse than a failed one, and for always passing a timeout to `AsyncResult.get()` if you stay with `Pool`. ### Building a system that survives it Because recovery cannot happen inside the broken executor, it has to be designed one level up. **Track completion outside the executor.** If work is checkpointed as it completes — records marked done in a store, output written per shard — then a crash 20 minutes into a 27-minute run costs you the incomplete tail rather than the whole run. Without that, retry means redoing everything. **Make tasks idempotent.** A killed worker died at an unknown point, possibly after a side effect. Retrying is only safe if running the unit twice is equivalent to running it once. **Retry with a fresh executor, and bound the retries.** Wrap the batch, catch `BrokenProcessPool`, build a new executor, resubmit the outstanding units. Cap the attempts, because a poison record reproduces the crash every time. Bisecting the retry — halving the batch until the offending unit is isolated — turns "the batch keeps dying" into "record X is the problem", and lets the rest of the work land. **Attack the cause, not just the symptom.** The commonest cause by far is memory. Size the pool by memory footprint rather than by core count, cap the per-task working set, and reject or route oversized inputs before they reach a worker. Where workers accumulate memory over time rather than blowing up on one record, `max_tasks_per_child` recycles them before they get there. A segfault instead means a native extension bug: pin the version, isolate the call, and consider running that specific work in a short-lived executor whose death costs nothing. **Isolate the risky stage.** A pipeline that runs risky native work in its own executor, separate from the executor doing safe pure-Python work, contains the blast radius: only the risky stage's futures are poisoned. ### What to say in an interview Name the exception and its class, state plainly that the pool is dead as a unit rather than per task, and then move immediately to the recovery design: checkpointed progress, idempotent units, a fresh executor on retry, bounded attempts with bisection for poison inputs, and pool sizing driven by memory. The contrast with `multiprocessing.Pool`'s silent loss of the in-flight task is the detail that shows you have actually operated both.
- How does multiprocessing.Pool behave differently when a worker is killed?It replaces the dead worker and carries on accepting tasks, but the task that worker was running is lost silently — its `AsyncResult.get()` blocks forever because no result is ever filed against it. So the same crash surfaces as a loud batch-wide failure under `ProcessPoolExecutor` and as a hang under `Pool`. If you use `Pool`, always pass a timeout to `get()` so a lost task becomes an error rather than a stall.
- Why not just have the executor replace the dead worker and retry the task?Because the executor cannot tell why the worker died. If the task itself caused it — a payload that exhausts memory, an input that segfaults a native library — a retry reproduces the crash, and automatic replacement turns one bad record into an unbounded kill loop. It also cannot verify the shared call and result queues are intact after a process was killed mid-write. Retrying is a policy decision that needs application knowledge, so it belongs to the caller.
- How would you find which input killed the pool when the batch is large?Bisect. Catch `BrokenProcessPool`, split the outstanding units in half, and resubmit each half to a fresh executor; the half that dies contains the poison record, and a handful of rounds isolates it while the healthy half completes. Bound the total attempts so a systemic failure does not turn into an endless split, and log the surviving candidate set at every round so the search is reconstructable afterwards.
- What sizing mistake most often causes this in production?Sizing the pool by core count when the constraint is memory. Each worker is a full interpreter with its own copy of the working set, so a machine with plenty of cores can still be pushed into the out-of-memory killer by the peak footprint of concurrent tasks. Size by measured per-worker peak against the available memory, leave headroom for the parent, and cap or route oversized inputs before they reach a worker.
It is a blown fuse rather than a tripped appliance: you do not swap the toaster, you reset the box — and you find out why it drew that much current before switching back on.
saying these in an interview costs you the question
- Thinking only the running future fails and the queue continues
- Believing the executor transparently replaces the dead worker
- Retrying on the same executor object after it is broken
- Blaming application code when the kernel issued an out-of-memory kill
- Retrying a poison input unboundedly with no bisection
- Sizing a process pool by cores when memory is the binding limit