skip to content

Why can writing to a subprocess.Popen pipe deadlock, and what does Popen.communicate() do about it?

level: middleimportance: should knowfreq 48%

answer

  1. The buffer between the two processes is small
  2. Both sides waiting, neither moving, no error
  3. wait() before read() is the documented trap
  4. One call services every stream at once
  5. Safety is paid for in parent memory

basics

~20 s

An OS pipe has a small fixed kernel buffer. If the parent keeps writing to the child's stdin without draining its stdout, the child blocks on a full output pipe, stops reading input, and both sides wait forever. Popen.communicate() services both directions at once, so it cannot deadlock that way.

solid answer

~50 s

Each pipe created by `stdin=subprocess.PIPE` or `stdout=subprocess.PIPE` is an OS pipe with a fixed kernel buffer - on Linux typically 64 KiB. Once the child has filled its stdout buffer, its next write blocks, so it stops reading its stdin; meanwhile the parent is blocked writing more input into a stdin buffer the child is no longer draining. Neither side can move: a classic mutual wait. The same hazard appears in the simpler shape `proc.wait()` followed by `proc.stdout.read()`, because `wait()` returns only when the child exits and the child cannot exit while its output pipe is full. `Popen.communicate()` is the fix: it writes the input, closes stdin, and reads stdout and stderr concurrently - with a selector on POSIX and helper threads on Windows - until EOF, then waits and returns the pair. Its cost is that everything is buffered in memory.

code

python · 12 lines
python
import subprocess
import sys

child = "import sys; sys.stdout.write(sys.stdin.read().upper())"
with subprocess.Popen(
    [sys.executable, "-c", child],
    stdin=subprocess.PIPE,
    stdout=subprocess.PIPE,
    text=True,
) as proc:
    out, _ = proc.communicate("shard 7 rebuilt\n", timeout=30)
print(proc.returncode, out.strip())

go deeper

for a junior

Know the safe default: prefer subprocess.run, and if you do use Popen with pipes, call communicate() rather than reading and writing the streams yourself. Recognise a hung job with no error output as a possible pipe deadlock.

for a middle

Explain the buffer mechanism concretely - a fixed-size kernel pipe, a blocked writer, a child that consequently stops reading - and why communicate() servicing every stream at once removes it. Name the wait()-then-read() trap.

for a senior

Diagnose it in production: a stalled job burning no CPU, both processes blocked on write, seen from the process state or a stack dump. Show the streaming alternative and the memory ceiling that makes communicate() unsuitable for very large output.

for a principal

Decide when a pipeline of processes is the wrong architecture at all - when the data volume argues for files, an object store, or a queue between stages instead of pipes whose failure mode is a silent hang no health check catches.

This is the question that separates people who have used `subprocess.run()` from people who have used `subprocess.Popen` in anger, and the mechanism is entirely about operating-system pipes rather than anything Python-specific. ## The mechanism When you pass `stdin=subprocess.PIPE` or `stdout=subprocess.PIPE`, the operating system creates a pipe: a kernel buffer of fixed size with a reader at one end and a writer at the other (64 KiB is the usual Linux figure). Writing to a full pipe blocks the writer until a reader consumes some bytes. Reading an empty pipe blocks the reader until a writer produces some, or until every write end is closed, which surfaces as EOF. Now take a parent that pushes a large document into a child's stdin and plans to read the transformed result afterwards. The child reads, transforms, writes. Its output pipe fills up after 64 KiB because nobody in the parent is reading it yet. Its next write blocks. Blocked in a write, the child never reads more stdin. The parent's stdin pipe fills up too, and the parent's next write blocks. Both processes are now waiting on the other, forever, with no error and no CPU use - the job simply stops, which makes it a nasty production incident rather than a crash. A more innocent-looking variant needs only one pipe: ```python proc = subprocess.Popen(cmd, stdout=subprocess.PIPE) proc.wait() # documented deadlock hazard data = proc.stdout.read() ``` `wait()` returns when the child exits, but a child that has produced more output than the buffer holds cannot exit - it is blocked writing. The module's documentation calls this out explicitly and tells you to use `communicate()` instead. ## What communicate() does `Popen.communicate(input=None, timeout=None)` handles every open pipe at once. It writes `input` to stdin (if there is one) and then closes stdin so the child sees EOF, reads stdout and stderr until both hit EOF, waits for the child to exit, and returns the `(stdout, stderr)` pair. On POSIX it multiplexes with a selector; on Windows it runs a helper thread per stream. Because no direction is ever left unserviced while another blocks, the deadlock above cannot occur. `subprocess.run()` uses `communicate()` internally, which is exactly why the high-level API is deadlock-free and why most code should stay there. The price is memory: `communicate()` accumulates the entire output in the parent before returning anything. For a child emitting gigabytes, redirect `stdout=` to a real file object or a temporary file, or stream the output yourself. ## Streaming safely If you want incremental output - progress lines from a long rebuild, for example - iterate the child's stdout while writing nothing: ```python with subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True) as proc: for line in proc.stdout: handle(line) ``` Two details make that safe. First, only one stream is piped: `stderr=subprocess.STDOUT` merges the error stream into the same pipe, so there is no second pipe to starve. Reading stdout to EOF while stderr fills its own separate pipe is the third common way to hang. Second, `Popen` used as a context manager closes the pipes and waits for the child on exit, so you do not leave the child running and the descriptors open. If you must write and read at the same time and cannot buffer everything, you need real concurrency - a reader thread per stream, or non-blocking descriptors driven by a selector - which is precisely the machinery `communicate()` already contains. Reaching for `communicate()` is nearly always the better answer. ## Two more traps `communicate(timeout=...)` raises `subprocess.TimeoutExpired` and, importantly, does **not** kill the child; the documented pattern is to catch it, call `proc.kill()`, and call `communicate()` again to reap the process and collect what output there is. And if you write to `proc.stdin` by hand, remember to flush and then close it: a child that is reading until EOF will wait indefinitely for a stdin you never closed, and that hang has nothing to do with buffer sizes at all.

  • You pipe both stdout and stderr separately and read only stdout to EOF. What goes wrong?
    The child can fill the stderr pipe and block writing to it, so it never finishes and never closes stdout - your read never reaches EOF. Either merge the streams with `stderr=subprocess.STDOUT`, drain both concurrently, or use `communicate()`, which already does.
  • communicate(timeout=5) raised TimeoutExpired. What state is the child in?
    Still running - `communicate()` does not kill it. The documented recovery is to call `proc.kill()` and then `communicate()` again, which reaps the process and returns whatever output was collected. Skipping the second call leaves you with a child that has been signalled but not waited for.
  • When would you not use communicate() at all?
    When the output is too large to hold in memory, or when you need it incrementally. Then redirect `stdout=` to a file object, or iterate `proc.stdout` while writing nothing to the child - and merge stderr into the same pipe so there is only one stream to service.

Two people passing notes through a narrow letterbox that holds only a few slips: once it is stuffed full in one direction, neither can post the next note, and both stand there waiting politely forever.

saying these in an interview costs you the question

  • Believes a pipe buffers unlimited data
  • Calls proc.wait() before reading a piped stdout
  • Says communicate is just a convenience wrapper for read
  • Reads only stdout while stderr has its own pipe
  • Assumes TimeoutExpired from communicate kills the child
  • Forgets to close the child's stdin so it never sees EOF

context