What do you use instead of Popen.communicate() when a child streams gigabytes to stdout?
answer
- Keep draining, stop accumulating
- A file descriptor has no fixed capacity
- Process it as it arrives instead
- One reader per stream, never one wait
- Peak memory equals the whole output
basics
~20 sStop capturing into memory. Point the child's stdout at an open file object, or keep subprocess.PIPE and drain it incrementally — a loop over Popen.stdout, or one reader thread per stream when the streams must stay separate.
solid answer
~50 s`Popen.communicate` is deadlock-safe but memory-hungry: it accumulates every byte of `stdout` and `stderr` and returns them as single objects, so a child producing gigabytes drives the parent's RSS to that size and risks the out-of-memory killer. The alternatives keep the drain but drop the accumulation. Passing an open file object as `stdout` hands the child a regular file descriptor — no bounded buffer, no reader needed, and the data is on disk where you can process it afterwards. Keeping `subprocess.PIPE` and iterating `Popen.stdout` line by line gives incremental processing at constant memory, and is the right shape when you want progress as it happens; merge `stderr` with `subprocess.STDOUT` so there is only one pipe to keep empty. When the two streams must stay separate, start one `threading.Thread` per stream so neither can starve the other, then join the threads and call `Popen.wait`.
code
python · 15 linesimport pathlib
import subprocess
import sys
import tempfile
child = "import sys; sys.stdout.write('row\\n' * 6800)"
with tempfile.TemporaryDirectory() as tmp:
report = pathlib.Path(tmp) / "report.txt"
with report.open("wb") as sink:
done = subprocess.run(
[sys.executable, "-c", child],
stdout=sink,
stderr=subprocess.STDOUT,
)
print(done.returncode, report.stat().st_size)go deeper
Recall that capturing output means holding all of it in memory, and that a child writing to a file object instead of subprocess.PIPE costs the parent nothing.
Explain the three shapes and their tradeoffs: file redirection for batch output, a merged pipe iterated line by line for incremental work, and reader threads when the streams must stay separate.
Show the failure you are avoiding — an OOM-killed parent with no traceback — and cover the details that bite in production: child-side block buffering, keeping stderr drained too, and retaining a timeout alongside the memory fix.
Own the boundary condition: output size is set by input you do not control, so the safe design bounds it structurally — write to disk or stream and aggregate — rather than sizing the container to today's largest run.
## What communicate costs `Popen.communicate` solves the deadlock by reading both pipes concurrently, but it solves it by *buffering*: it appends every chunk to a list and returns the joined result. Peak memory is therefore the full output — briefly closer to twice it, while the chunks and the joined object both exist. That is fine for a version banner and wrong for a data extract. A parent that quietly allocates several gigabytes is a parent that gets killed by the kernel's out-of-memory killer, and in a container it is a restart with no Python traceback at all — the process simply vanishes. So the goal is to keep the property that made `communicate` safe, *something is always draining the pipe*, while removing the property that makes it expensive, *everything is kept*. ## Option 1: never use a pipe — redirect to a file `stdout=open(path, "wb")` passes a real file descriptor to the child. A regular file has no fixed capacity, so the child can never block on a full buffer and the parent needs no reader at all; it can call `Popen.wait` safely. Memory stays flat, the parent can be doing something else entirely, and you get a durable artefact you can re-read, checksum or ship. The costs are disk space and latency: nothing is available to your code until the child finishes, or until you tail the file yourself. Add `stderr=subprocess.STDOUT` to merge, or a second file to keep the streams apart. This is the simplest correct answer for batch work and the one many candidates never reach for. ## Option 2: keep the pipe and consume it incrementally When you want to act on output as it arrives, iterate the stream: ```python with subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT, text=True) as proc: for line in proc.stdout: handle(line) proc.wait() ``` Memory is one line at a time, and the pipe never stays full because the loop keeps taking from it. Two details matter. The merge into a single stream is what makes it safe — with `stderr` on a second unread pipe you are back in the deadlock. And the child's own output buffering decides your latency: a program writing to a pipe usually switches to block buffering, so lines can arrive in 4–8 KiB clumps or not at all until exit. For a Python child, `-u` or `PYTHONUNBUFFERED` fixes it; for other programs you may need their own flush flag or a pseudo-terminal. ## Option 3: a reader thread per stream If you must keep `stdout` and `stderr` apart *and* process incrementally, give each pipe its own `threading.Thread` whose only job is to read that stream to end-of-file. Neither stream can starve the other, and the main thread can wait on the child and then join the readers. This is precisely what `Popen.communicate` does internally on Windows — you are re-implementing it with your own per-chunk handling instead of accumulation. Keep the thread bodies trivial: read, hand off, close. Anything that can block in a reader thread reintroduces the original hang. ## A concrete failure A ticket-triage bot shells out to a classifier and captures its report with `capture_output=True`. On a 6,800-row batch the report is a few hundred kilobytes and everything is fine; as the queue grows, the same call is asked to hold a multi-gigabyte dump. The container is OOM-killed mid-run, the supervisor restarts the bot, and because the exception path was never reached the bot serves the previously cached scores instead — a stale cached value presented as a fresh classification, with a green dashboard above it. The fix is not a bigger container. It is to stop deciding memory usage by the size of someone else's output: write the report to a file, or stream it row by row and aggregate as you go. ## Choosing Ask two questions. *Do I need the bytes at all?* If not, `subprocess.DEVNULL`. *Do I need them incrementally?* If not, redirect to a file and read it afterwards; if yes, stream a merged pipe, or run reader threads when the streams must stay distinct. Reach for `Popen.communicate` only when the output is small and bounded by something you control — and remember that "small" is a property of today's input, not of the code. Whichever shape you choose, still set a timeout: the memory strategy and the hang strategy are independent, and you need both.
- Why can iterating Popen.stdout show nothing for a long time and then a burst of lines?Because the child, not the parent, decides when to flush. Most programs use line buffering on a terminal but switch to block buffering when stdout is a pipe, so output arrives in 4–8 KiB chunks. A Python child can be run with `-u` or with `PYTHONUNBUFFERED` set; other programs need their own unbuffered flag, or a pseudo-terminal to convince them they are interactive. It is not a Python-side buffering problem.
- If you redirect stdout to a file, do you still need to worry about the pipe deadlock?Not for that stream — a regular file descriptor has no fixed capacity, so the child cannot block writing to it and `Popen.wait` is safe. But check every stream: if `stderr` is still `subprocess.PIPE` and nobody reads it, a chatty child can fill that pipe and stall exactly as before. Send `stderr` to the same file, to its own file, or to `subprocess.DEVNULL`.
- When would you use asyncio.create_subprocess_exec instead of threads for this?When the parent is already an asyncio program, or when it supervises many children at once. The event loop can await reads from every child's streams without a thread apiece, so scaling to dozens of concurrent commands costs no extra stack. For a synchronous program running one or two children, reader threads are simpler and carry no risk of blocking the loop with accidental synchronous work.
Reading a firehose into a bucket works until the bucket is smaller than the flow. You either point the hose at a drain, or you keep a cup under it and empty the cup continuously.
saying these in an interview costs you the question
- Reaches for capture_output regardless of output size
- Thinks communicate streams rather than accumulates
- Solves memory pressure by enlarging the container
- Redirects stdout to a file but leaves stderr piped
- Blames Python buffering for a child's block buffering
- Drops the timeout once memory is handled