skip to content

Why does multiprocessing.Queue deadlock when you join the child before draining it?

level: middleimportance: must knowfreq 55%

answer

  1. put() buffers, it does not deliver
  2. A feeder thread writes into an OS pipe
  3. The pipe has a fixed small capacity
  4. The child cannot exit until it flushes
  5. Drain to a sentinel, then join the process

basics

~20 s

A multiprocessing.Queue is a bounded OS pipe fed by a background thread in the writing process. If the pipe fills, that thread blocks until the reader drains it, and the child cannot exit — so the parent's join() waits forever. Drain first, then join.

solid answer

~50 s

`multiprocessing.Queue` is not a shared container; it is an OS pipe plus a per-process buffer and a **feeder thread** that pickles queued items and writes them into the pipe. `put()` returns as soon as the item is buffered, so a child can "finish" its loop with megabytes still unflushed. The pipe has a fixed OS-level capacity — tens of kilobytes on many systems — and once it is full the feeder thread blocks on the write. A child process will not exit until its feeder thread has flushed, so `Process.join()` in a parent that has not yet called `get()` blocks forever: the child waits for the parent to read, the parent waits for the child to exit. The fix is ordering, not tuning: **drain the queue to completion first, then join.** Send one sentinel per worker so the reader knows when to stop, and never rely on `qsize()` or `empty()`, which are approximate.

code

python · 14 lines
python
import multiprocessing as mp

def producer(q):
    q.put(b"x" * 10_000_000)

if __name__ == "__main__":
    q = mp.Queue()
    p = mp.Process(target=producer, args=(q,))
    p.start()
    p.join(timeout=2)
    print("alive after join timeout:", p.is_alive())  # True
    payload = q.get()      # draining unblocks the feeder thread
    p.join()
    print("joined; bytes received:", len(payload))

go deeper

for a junior

Remember the ordering rule: read everything out of a multiprocessing.Queue before joining the process that wrote to it, and send an explicit end marker so the reader knows when to stop.

for a middle

Explain the machinery — the feeder thread, the bounded OS pipe, and why a child cannot exit with unflushed data — and show the sentinel-per-producer drain loop as the fix.

for a senior

Diagnose it live: a hung parent in join, a child alive with no CPU, and a queue nobody is reading. Then argue for removing the hand-rolled queue entirely in favour of a pool or executor that owns the result path.

for a principal

Frame it as a design smell. Bounded channels with no dedicated drainer are a deadlock class, not a bug; decide where backpressure belongs in the pipeline and whether the team should be writing raw queues at all.

### What a multiprocessing.Queue actually is The name invites the wrong model. `queue.Queue` is a container in one process guarded by a lock. `multiprocessing.Queue` is a **transport**: an operating-system pipe, plus a small in-process buffer on the writing side, plus a background **feeder thread** created lazily on the first `put()`. The write path is: `put(obj)` appends `obj` to the writer's internal buffer and returns almost immediately. The feeder thread later pops it, pickles it to bytes, and writes those bytes into the pipe. The read path is: `get()` reads bytes from the pipe in the receiving process and unpickles them. Two consequences follow immediately, and both are the whole question: 1. **`put()` returning does not mean the data has gone anywhere.** It means the item is queued for the feeder thread. 2. **The pipe is bounded.** Its capacity is an OS property, commonly around 64 KB, and it is not a Python setting you can raise. When it is full, the feeder thread's write blocks until the other side reads. ### The deadlock, step by step A child does `q.put(big_payload)` and returns from its target function. The interpreter now runs the child's shutdown, which joins the queue's feeder thread so buffered data is not silently thrown away. The feeder is stuck mid-write because the pipe is full. So the child never exits. Meanwhile the parent calls `p.join()` before any `get()`. `join()` waits for the child to exit. Nobody is reading the pipe, so the pipe never drains, so the feeder never finishes, so the child never exits, so `join()` never returns. Classic circular wait: **parent waits for exit, child waits for a reader.** The cruel part is scale sensitivity. With three small integers the whole payload fits in the pipe buffer, the feeder flushes instantly, and `join()`-then-`get()` works perfectly — in the test. The same code deadlocks in production the day a payload crosses the buffer size. This is why the bug reaches production so often. ```python q = mp.Queue() p = mp.Process(target=producer, args=(q,)) p.start() p.join() # hangs once the payload exceeds the pipe buffer print(q.get()) # never reached ``` ### The fix: ordering, plus a termination signal The correct shape is always **drain, then join**: ```python p.start() for item in iter(q.get, None): # None is the sentinel handle(item) p.join() ``` The reader needs to know when to stop, and the queue cannot tell it. `qsize()` is approximate and raises `NotImplementedError` on some platforms; `empty()` is a snapshot that is stale the moment it returns. Neither is a termination condition. Use one of: * **A sentinel per producer.** Each worker puts a unique end marker such as `None` when done; the reader counts markers until it has seen one from every worker. Note *per producer*: one sentinel with four workers stops the reader after the first worker finishes. * **`multiprocessing.JoinableQueue`.** Consumers call `task_done()` for each item and the producer side calls `join()` **on the queue** to wait until every item has been marked done. Be careful with the vocabulary in an interview: `Queue.join()` on a JoinableQueue means "all items processed", while `Process.join()` means "that process has exited". They are different waits and mixing them up is its own deadlock. * **Do not hand-roll it at all.** `multiprocessing.Pool` and `concurrent.futures.ProcessPoolExecutor` already own the result queues and their draining; `Pool.imap` or an executor's `map` gives you results as an iterator with none of this exposed. ### The escape hatch and why it is a last resort `Queue.cancel_join_thread()` tells the queue not to block process exit on flushing. It removes the hang by **discarding buffered data** — the child exits and whatever had not reached the pipe is gone. That is acceptable for a best-effort log queue on a forced shutdown and is wrong almost everywhere else, because it converts a loud deadlock into silent data loss. Its counterpart `join_thread()` explicitly waits for the flush after `close()`. ### Related shapes worth recognizing The same bounded-pipe reasoning explains the mirror-image hang: a parent that fills a queue for workers, then waits for workers that are themselves blocked writing results back into a full result queue. Two bounded channels with no drainer on either side is a textbook deadlock, and the fix is the same — a dedicated reader that keeps draining, or a pool that does it for you. It also explains why a child that raises after putting a large item can leave a parent hung: the exception unwinds, the shutdown still tries to flush the feeder, and the pipe is still full. The rule to carry away: **with `multiprocessing.Queue`, the reader must run before the join, and the writer's `put()` is a promise, not a delivery.**

  • Why can't you just call qsize() until it returns zero and then join?
    Because `qsize()` is approximate: it is not reliable across processes, it does not account for bytes still in flight inside the pipe or the writer's buffer, and on some platforms it raises `NotImplementedError`. `empty()` has the same problem — it is a stale snapshot the instant it returns. Termination has to be signalled in the data stream, with a sentinel per producer or with `JoinableQueue.task_done()` accounting.
  • What does Queue.cancel_join_thread() actually change, and when would you accept it?
    It stops the queue from blocking process exit while the feeder thread flushes, so the process can terminate with items still buffered — those items are lost. It removes the hang by discarding data, which is only defensible for best-effort traffic such as diagnostics on a forced shutdown. For anything whose loss matters, fix the ordering instead: drain the queue, then join the process.
  • How is JoinableQueue.join() different from Process.join()?
    `Process.join()` waits for an operating-system process to exit. `JoinableQueue.join()` waits until every item put on that queue has been matched by a `task_done()` call from a consumer — it is about work accounting, not process lifetime. Consumers on a JoinableQueue are usually daemon processes you never join individually; you wait on the queue and then let them be torn down.
  • Why does this bug so often pass tests and fail in production?
    Because the deadlock only appears once the payload exceeds the operating system's pipe capacity, commonly tens of kilobytes. A test that puts a handful of small objects fits entirely in the buffer, the feeder thread flushes immediately, the child exits, and join-then-get looks correct. The first large real payload crosses the threshold and the identical code hangs.

The queue is an outbound mail slot with a clerk feeding it. If the slot is jammed full the clerk cannot go home, and standing at the door waiting for the clerk to leave without ever emptying the slot means nobody moves.

saying these in an interview costs you the question

  • Believing put() has delivered the data when it returns
  • Calling Process.join() before draining the queue
  • Using qsize() or empty() as a termination condition
  • Reaching for cancel_join_thread() to silence the hang
  • Confusing JoinableQueue.join() with Process.join()
  • Assuming the pipe capacity can be raised from Python

context