skip to content

A renderer service spawns a subprocess.Popen per job and never calls poll() or wait(); ps fills with <defunct> children. What is leaking, and how do you fix it?

level: seniorimportance: should knowfreq 38%

answer

  1. The dead child awaits a headcount
  2. Not memory: a slot and a PID
  3. Signals cannot clear it
  4. One non-blocking call is enough
  5. Completion order, not launch order

basics

~20 s

Nothing leaks memory: a finished child keeps a process-table entry and its PID until the parent collects its exit status. Calling poll(), wait() or communicate() on the subprocess.Popen object reaps it; a parent that never does eventually cannot spawn.

solid answer

~50 s

A `<defunct>` entry is a zombie — a child that has already exited, whose exit status the kernel holds until its parent collects it. It holds no memory and no CPU, only a process-table slot and a PID, and those are finite, so a long-lived parent that never reaps eventually fails to spawn with an `OSError`. The reaping calls are `Popen.wait` (blocking), `Popen.poll` (non-blocking: `None` while the child runs, the code once it has finished) and `Popen.communicate`; `subprocess.run` cannot leak, because it waits internally. The usual bug is an ordering assumption — the parent keeps its `subprocess.Popen` objects in a list and waits on them in launch order, so a job that finished third stays defunct while the parent blocks on the first. Sweep all outstanding children with `poll()` instead, or give each one a thread that waits on it.

code

python · 10 lines
python
import subprocess, sys, time

running = [
    subprocess.Popen([sys.executable, "-c", "import time; time.sleep(0.2)"])
    for _ in range(4)
]
while running:
    running = [child for child in running if child.poll() is None]
    time.sleep(0.05)
print("all four children reaped")

go deeper

for a junior

Know that a finished child process stays listed as defunct until its parent asks for its exit status, and that calling wait() or poll() on the subprocess.Popen object is what asks.

for a middle

Explain what the entry actually holds — a PID and a process-table slot, not memory — and why exhausting those makes the next spawn fail with an OSError far from the code that caused it.

for a senior

Demonstrate the diagnosis on a live service: a monotonically growing defunct count under one parent PID, ResourceWarning under development mode, and a reaping design that collects in completion order instead of launch order.

for a principal

Own the pattern across services: where process spawning is allowed at all, whether a supervisor or a pool owns child lifetimes, and how a slow-burn resource exhaustion like this gets an alert rather than being masked by periodic restarts.

## What a zombie actually is When a process exits, the kernel tears down its memory, closes its descriptors and stops scheduling it — but it cannot discard the process entirely, because nobody has yet asked how it ended. The exit status has to be kept somewhere until the parent collects it, and that somewhere is a residual process-table entry holding the PID and the status. `ps` shows such an entry in state `Z` with the label `<defunct>`. It is not stuck, it is not consuming CPU or memory, and it cannot be killed: it is already dead, and a signal to a dead process does nothing. The only thing that clears it is the parent calling one of the wait family. So the leak is not memory. It is PIDs and process-table slots, both of which are bounded system-wide and per-user. A service that spawns a child per unit of work and never reaps will run normally for hours and then start failing to spawn anything at all, surfacing in Python as an `OSError` from the very next `subprocess.Popen` call — a failure that looks nothing like its cause. ## The Python-level mapping Every reaping mechanism in `subprocess` is one of three calls on the `subprocess.Popen` object: * `Popen.wait` blocks until that specific child finishes, collects the status, and returns it. * `Popen.poll` does not block: it returns `None` if the child is still running and the return code if it has finished — and the act of returning a code *is* the reap. * `Popen.communicate` exchanges data with the child and waits for it as part of the same call. `subprocess.run` is safe by construction, because it waits before returning; only long-lived code holding `subprocess.Popen` objects can leak. Underneath, all of these end up in the same place as `os.waitpid`, which you can also call directly — `os.waitpid(pid, os.WNOHANG)` is the raw non-blocking form for PIDs you obtained without `subprocess`. ## The ordering assumption that causes it A concrete shape from an invoice-PDF renderer: the service farms each document out to a child, and because a four-person team wrote it around a batch of four workers, the parent collects results by walking its list of children in launch order and calling `wait()` on each in turn. That looks like reaping and mostly is — until rendering time stops correlating with submission order. One pathological invoice makes the first child run for minutes; the other three finish in seconds and sit defunct the whole time, because the parent is blocked on a child that has not exited. Under load the batch loop never gets far enough down the list to reap anything, the defunct count grows monotonically, and the service dies of PID exhaustion while appearing, in its own logs, to be working normally. The fix is to stop assuming any completion order. Sweep every outstanding child with a non-blocking `poll()` each pass and drop the ones that report a code; or hand each child to its own thread that blocks in `wait()` and reports back. Both reap in completion order rather than launch order. ```python running = [child for child in running if child.poll() is None] ``` ## Diagnosing it on a live service Count the defunct entries and see whether the count only ever grows — a steady number is normal churn, a monotonic climb is the leak. Confirm the parent: zombies are always children of a specific process, and the parent PID column names the process that owes the wait. Inside Python, run the service under development mode (`-X dev`) or with warnings as errors: destroying a `subprocess.Popen` whose child is still running emits a `ResourceWarning` naming the process, which is exactly the smell of a child dropped without being waited for. Finally, remember that restarting the parent clears the whole backlog at once, because those entries are reparented and collected by the init process — which is why this bug is so often "fixed" by a restart and then returns. ## The rules that keep it away Prefer `subprocess.run` whenever you can afford to block, since it cannot leak. When you do need `subprocess.Popen`, treat the object as owning a resource: every one you create must eventually reach a `poll()` that returns a code, a `wait()`, or a `communicate()`, and there should be exactly one place in the code responsible for that. Never write a loop that waits on children in the order you started them; and never keep a list of finished `subprocess.Popen` objects around for their output without having collected their status first.

  • How do you reap children without blocking a service's main loop?
    Sweep the outstanding children with poll() on each pass and discard those that return a code — poll() never blocks and the returned code is the reap. For PIDs you did not get from subprocess, os.waitpid(pid, os.WNOHANG) is the raw equivalent. The alternative is a thread per child that blocks in wait(); either way you collect in completion order rather than launch order.
  • Can subprocess.run leave a zombie behind?
    No. run waits for the child before it returns, including on its failure paths, so by the time you hold a subprocess.CompletedProcess the status has already been collected. Zombies come from long-lived code holding subprocess.Popen objects it never waits on — which is the price of the extra control Popen gives you.
  • Why does sending SIGKILL to a defunct process not remove it?
    Because it is already dead. The process image is gone; what remains is a kernel bookkeeping entry holding the exit status for a parent that has not collected it. Signals are delivered to running processes, so nothing is there to receive one. Only the parent waiting — or the parent exiting, after which the entry is reparented and collected — clears it.

A zombie is a death certificate nobody has collected: the person is gone, but the registry keeps the file open until the next of kin signs for it.

saying these in an interview costs you the question

  • Thinks a zombie still uses CPU or memory
  • Tries to clear defunct entries by killing them
  • Assumes garbage collection reaps child processes
  • Waits on children strictly in launch order and calls that reaping
  • Confuses a zombie with an orphan that is still doing work
  • Believes only a parent restart can clear the entries

context