skip to content

Why can a with-block around multiprocessing.Pool lose pending results?

level: seniorimportance: nice to knowfreq 25%

answer

  1. The exit method is one line
  2. Not every shutdown verb waits
  3. Blocking calls hide the problem
  4. Fire-and-forget submissions are the ones at risk
  5. Two calls in order give you a drain

basics

~10 s

Pool.exit calls terminate(), not close() then join(). Leaving the with-block while apply_async or map_async submissions are still outstanding kills the workers immediately, so those tasks never run and their AsyncResult objects are never fulfilled.

solid answer

~40 s

`multiprocessing.Pool.__exit__` is one line: `self.terminate()`. That is the right default for exception safety — workers are always reaped, even on an error path — but it is not a graceful drain. Blocking calls are safe because they finish inside the block: `map`, `starmap` and any `AsyncResult.get()` you make *before* the dedent. The trap is fire-and-forget: submitting with `apply_async` or `map_async`, exiting the block, then calling `.get()` on the results. The workers were killed, those results are never set, and a `.get()` with no timeout blocks forever. If you need every submission to complete, drain inside the block or shut down explicitly with `pool.close()` followed by `pool.join()` — `close()` stops new submissions, `join()` waits for the queued work. Note `join()` before either raises `ValueError`.

code

python · 19 lines
python
import multiprocessing as mp
import time


def load(batch_id):
    time.sleep(0.5)
    return batch_id * 2


if __name__ == "__main__":
    with mp.Pool(2) as pool:
        pending = [pool.apply_async(load, (i,)) for i in range(6)]
    print("ready after with-block:", [r.ready() for r in pending])

    pool = mp.Pool(2)
    pending = [pool.apply_async(load, (i,)) for i in range(6)]
    pool.close()
    pool.join()
    print("ready after close+join:", [r.ready() for r in pending])

go deeper

for a junior

Remember that leaving the pool's with-block does not politely wait for work to finish. Retrieve every result before the block ends and this never bites you.

for a middle

Be able to name the three shutdown verbs and what each one does: stop accepting work, wait for workers to exit, kill immediately — and say which one the context manager runs.

for a senior

Diagnose the symptom: a silent hang with no children left and the parent parked waiting on a result. Then argue the structural fix, and default to a timeout on every result retrieval in production code.

for a principal

Decide the house rule. Guaranteed cleanup on the exception path versus guaranteed completion on the success path is a real tension; pick one shape for the codebase, write it down, and make abandoned-work cases explicit rather than accidental.

This one bites people who reason by analogy from other pool APIs, where leaving a context manager means "wait for the work". `multiprocessing.Pool` does the opposite, and the source is unambiguous: ```python def __exit__(self, exc_type, exc_val, exc_tb): self.terminate() ``` **Three shutdown verbs, three meanings.** `close()` tells the pool to accept no further submissions. Work already queued still runs to completion. It returns immediately and waits for nothing. `join()` blocks until every worker process has exited. It is only legal after `close()` or `terminate()` — calling it on a running pool raises `ValueError("Pool is still running")`. `close()` then `join()` is therefore the drain: no new work, then wait for the rest. `terminate()` kills the workers now. Queued tasks are discarded, in-flight tasks are cut off mid-execution, and the `AsyncResult` objects for anything unfinished are simply never completed. That is what the `with` block does on the way out, on both the normal and the exception path. **Why the trap is invisible.** The common shape hides it: ```python with multiprocessing.Pool(4) as pool: results = pool.map(load, batches) # blocks here — safe ``` `map` does not return until every result is in, so by the time `__exit__` runs there is nothing left to lose. The `with` block is genuinely correct here, and it is correct for `starmap`, and for `imap` fully consumed inside the block. Now change one word: ```python with multiprocessing.Pool(4) as pool: pending = [pool.apply_async(load, (b,)) for b in batches] values = [r.get() for r in pending] # hangs forever ``` The submissions returned instantly, the block ended, `terminate()` killed the workers, and the results will never arrive. `ApplyResult.get()` with no timeout waits on an event that will never be set, so the program hangs with no error and no CPU use — the worst failure mode there is. With a timeout you at least get an exception, and that exception is worth knowing: `AsyncResult.get(timeout=...)` raises **`multiprocessing.TimeoutError`**, which derives from `multiprocessing.ProcessError` and is **not** the builtin `TimeoutError`. An `except TimeoutError:` around it does not catch it. **The three correct shapes.** First, collect inside the block. Keep the `with`, and make sure every result is retrieved before the dedent — call `.get()` on each `AsyncResult`, or use a blocking form. This keeps the exception-safety the context manager exists for. Second, shut down explicitly when you want a drain: construct the pool, submit, `pool.close()`, `pool.join()`, then read the results. Wrap it in `try`/`finally` with `terminate()` in the failure path if you also want guaranteed cleanup on an exception. Third, when abandoning work is genuinely the intent — a cancelled run, a caller that timed out — the `with` block is exactly right, because killing outstanding work immediately is what you want. **Why the default is defensible.** A `Pool` owns real OS processes. If the body of the `with` raises, waiting for a possibly-wedged worker would turn one bug into a hung process, and worker processes leaked from a failed run are worse than discarded results. `terminate()` guarantees the resource is released. The API's mistake, if there is one, is that the same verb runs on the success path too, so "it worked in testing" is a function of whether your test drained inside the block. **Diagnosing it in the wild.** A run that hangs after producing partial output, with all children gone and the parent parked in a `wait`, is this bug. A stack dump of the parent shows a frame inside `ApplyResult.get`, and the children have already exited. The fix is structural, not a tuning knob: move the collection inside the block or replace the `with` with `close()` plus `join()`. A defensive habit worth adopting regardless: always pass a timeout to `AsyncResult.get()` in production code. It converts an unbounded hang into a catchable `multiprocessing.TimeoutError` you can log with context, and costs nothing when the code is right.

  • What is the difference between Pool.close() and Pool.terminate()?
    `close()` stops the pool accepting new submissions and lets already-queued work finish; it returns immediately and waits for nothing. `terminate()` kills the worker processes now, discarding queued tasks and cutting off in-flight ones. Neither waits — `join()` is what blocks until the workers have exited, and it is legal only after one of the two has been called.
  • So should you stop using with-blocks for multiprocessing.Pool?
    No. The context manager guarantees the worker processes are reaped even when the body raises, which is exactly what you want from a resource holding real OS processes. Keep it, and make sure every result is retrieved before the dedent. Reach for explicit close() plus join() only when submissions must outlive the block.
  • A run hangs with no output and no CPU use after the pool block. How do you confirm this is the cause?
    Check that no worker children remain while the parent is still alive — the pool was terminated. Then dump the parent's stack: a frame inside `ApplyResult.get` waiting on an event confirms it. The fix is structural rather than a tuning change: collect results inside the block, or replace the with-statement with close() followed by join().

It is a lab closing at the stroke of five: everything mid-bench is binned, not finished. If you want the day's work, you stop accepting samples first and then wait for the benches to clear.

saying these in an interview costs you the question

  • Assuming the with-block waits for outstanding tasks to finish
  • Calling AsyncResult.get() with no timeout after the pool exited
  • Thinking terminate() lets in-flight tasks run to completion
  • Catching the builtin TimeoutError around AsyncResult.get(timeout=...)
  • Believing close() blocks until queued work is done
  • Calling join() on a running pool without close() or terminate() first

context