skip to content

Worker Crashes and Recovery

What happens when a worker dies: a segfault or the OOM killer poisons a pool, pending futures raise BrokenProcessPool, and leaked children outlive the parent. Interviewers want the recovery story.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How does an exception raised inside a multiprocessing.Pool worker reach the parent?

level: middleimportance: must knowfreq 60%

answer

  1. The child cannot share a call stack
  2. Something must survive the pickle boundary
  3. Tracebacks are not picklable; text is
  4. Two paths: re-raised, or a callback
  5. The callback runs on the pool's own thread

basics

~20 s

The worker catches the exception, pickles it together with a text rendering of the child's traceback, and sends it back. The parent re-raises it out of AsyncResult.get(), or hands it to error_callback if you supplied one.

solid answer

~40 s

A pool worker never lets an exception escape into the void: it catches it, wraps it so the child's traceback survives as text, pickles the pair and sends it over the result pipe. In the parent, the result-handler thread stores it against that task. `AsyncResult.get()` then re-raises the original exception in your calling thread, with the remote traceback attached as its `__cause__` so the printed traceback shows both sides; `AsyncResult.successful()` returns `False` once the task is done. If you passed `error_callback` to `apply_async()` or `map_async()`, that callable is invoked with the exception instance instead — but it runs on the pool's internal result-handler thread, not yours, so it must be quick and must not raise. The blocking `map()` simply re-raises the first exception it meets and discards the rest of the batch's results.

code

python · 20 lines
python
import multiprocessing as mp


def score(bid):
    raise ValueError(f"bad bid {bid}")


def on_error(exc):
    print("error_callback:", type(exc).__name__, exc)


if __name__ == "__main__":
    with mp.Pool(1) as pool:
        result = pool.apply_async(score, (7,), error_callback=on_error)
        try:
            result.get()
        except ValueError as exc:
            print("get() re-raised:", exc)
            print("cause:", type(exc.__cause__).__name__)
        print("successful():", result.successful())

go deeper

for a junior

Know that a failure inside a pool worker is not silently swallowed: it comes back and is re-raised where you ask for the result. Always ask for the result rather than assuming a submitted task succeeded.

for a middle

Explain the round trip: the worker catches, the traceback becomes text because frames are not picklable, and the parent re-raises out of AsyncResult.get() or calls error_callback. Know that Pool.map() gives you the first exception and nothing else.

for a senior

Demonstrate the failure modes you have hit: an unpicklable exception hanging get() forever, a heavy error_callback throttling the result-handler thread, and the batch design that reports per-item outcomes instead of losing a whole run to one bad record.

for a principal

Own the error contract at the boundary: what a task is allowed to raise, whether failures are values or exceptions, whether a batch is all-or-nothing or per-item, and how the failure signal reaches monitoring rather than dying in a worker's stderr.

### The problem being solved A worker in a `multiprocessing.Pool` runs in a separate operating-system process, so an exception raised in your task function cannot propagate up a shared call stack — there is no shared stack. The pool's job is to turn that remote failure into a local one that looks as much like an ordinary exception as the process boundary permits. ### What the worker does The worker loop calls your function inside a `try`. On failure it captures the exception object *and* formats the child's traceback into a string, because traceback objects are not picklable — they reference live frames. It then pickles the pair and writes it to the result pipe, flagged as a failure rather than a result. The worker then goes back to waiting for the next task: one failing task does not kill a `multiprocessing.Pool` worker, and it does not disturb the other tasks in flight. ### What the parent does The parent runs an internal result-handler thread that reads the pipe and files each incoming result against the task that produced it. Two things then happen depending on how you asked for the result: * **You asked synchronously.** `AsyncResult.get()` (and the blocking `map()` / `apply()` built on it) re-raises the unpickled exception in the thread that called it. The remote traceback is attached as the exception's `__cause__`, so a printed traceback shows the child's frames under a header, then the parent's frames. `AsyncResult.successful()` reports `False`, and raises if the task has not finished yet — so check `ready()` first if you are polling. * **You supplied `error_callback`.** `Pool.apply_async()` and `Pool.map_async()` both take `callback` and `error_callback`. On failure the pool calls `error_callback(exception_instance)` and does *not* call `callback`. This is the fire-and-forget path: it lets you log or count failures without anyone ever calling `get()`. The critical operational detail about `error_callback` is **which thread runs it**: the pool's own result-handler thread, not yours. That thread is the single pump feeding every result back to the parent. A slow callback throttles the whole pool, a blocking one stalls it, and an exception raised inside the callback propagates into that thread and can tear the result machinery down. Keep it to appending to a list or emitting a log line. ### The two edges that bite **The exception must round-trip through pickle.** Pickle reconstructs an exception by calling its class with the arguments recorded in `args`. A custom exception whose `__init__` takes extra required parameters, but which does not pass them to `super().__init__()` or define `__reduce__`, will pickle fine in the child and then fail to *unpickle* in the parent. That failure happens inside the result-handler thread, which dies with a traceback printed from a thread nobody is watching — and then `get()` blocks forever, because the result it is waiting for will never be filed. The symptom in production is a hang, not an error, which is why this is worth knowing before you meet it. The fix is to keep exception constructors pickle-friendly: accept the same arguments you pass to `super().__init__()`, and prefer plain messages over rich objects that may themselves be unpicklable. **Only the exception survives, not the frames.** You get the exception instance and a *string* of the child's traceback. Live locals, `__traceback__` frames from the child, open resources — none of that crosses. If your debugging depends on the failing frame's locals, capture what you need in the child and put it in the exception's message or attributes. ### Batch semantics `Pool.map()` re-raises the *first* exception it encounters and gives you nothing else — the successful results of the same batch are discarded along with it. When partial success matters, submit tasks individually with `apply_async()` and inspect each `AsyncResult`, or have the task function itself catch its exceptions and return a tagged value such as `("error", repr(exc))`, so every item reports its own outcome and the batch always completes. ### Where this stops working All of this assumes the worker process is *alive* to catch and report. A task that segfaults the interpreter or gets killed by an out-of-memory kill never reaches the `except` clause, so nothing is sent back at all — a different failure mode entirely, with a different recovery story.

  • Which thread runs error_callback, and why does that matter?
    The pool's internal result-handler thread in the parent — the single pump that files every incoming result. So the callback must be fast and must not raise: a blocking callback stalls result delivery for every task in the pool, and an exception inside it propagates into that thread and can break the pool's result machinery. Append to a list or log a line; do the real work elsewhere.
  • Why can a failing task make AsyncResult.get() hang forever instead of raising?
    Because the exception has to be unpickled in the parent before it can be re-raised. A custom exception whose constructor requires arguments it never passes to `super().__init__()` fails to reconstruct, and that failure happens on the result-handler thread. That thread dies, no result is ever filed against the task, and `get()` waits forever. Always pass a timeout to `get()` in production, and keep exception constructors picklable.
  • How do you get partial results out of a batch where one item fails?
    `Pool.map()` re-raises the first exception and discards the whole batch's results, so either submit with `apply_async()` and inspect each `AsyncResult` independently, or make the task function catch its own exceptions and return a tagged outcome such as a success/failure pair. The second approach keeps every item reporting for itself and means the batch always runs to completion.

saying these in an interview costs you the question

  • Claiming exceptions are lost or only printed in the child
  • Expecting the child's live traceback frames to arrive intact
  • Doing slow or blocking work inside error_callback
  • Assuming Pool.map returns the results that did succeed
  • Thinking a failing task kills or replaces the pool worker
  • Writing custom exceptions whose constructors cannot be unpickled

context

open as a page

What happens to pending futures when a ProcessPoolExecutor worker is killed?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Every pending and running future is completed with BrokenProcessPool, and every later submit raises it too. The executor is permanently unusable, so recovery means building a new one and resubmitting the work that did not finish.

open as a page

What does a negative multiprocessing.Process.exitcode mean?

level: juniorimportance: should knowfreq 40%

basics

~20 s

A negative exitcode of -N means the child was killed by signal N and ran no cleanup. None means it has not finished, 0 means a clean exit, and a positive number is the status the child chose.

open as a page

When is maxtasksperchild on a multiprocessing.Pool worth setting?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Set it when workers accumulate what they should not keep across tasks — leaked memory or stale per-process cached state. The worker retires after that many tasks and a fresh one replaces it; each recycle costs a process start.

open as a page