skip to content

How does a Python program at PID 1 reap the orphans it inherits?

level: seniorimportance: should knowfreq 35%

answer

  1. A dead process still holds a slot
  2. Orphans land on the first process
  3. Non-blocking, and any child at all
  4. Loop until it reports nothing new
  5. Deliveries coalesce, one signal many deaths

basics

~20 s

Orphans are reparented to PID 1, and a dead child stays a zombie until its parent collects it. A Python program there must loop on os.waitpid(-1, os.WNOHANG) until it reports nothing more, on signal.SIGCHLD or each main-loop pass.

solid answer

~40 s

A terminated process keeps a slot in the process table — a zombie — holding its PID and exit status until its parent waits for it. When a process dies with children still running, those children are reparented to PID 1, so a program running there inherits grandchildren it never spawned and becomes responsible for collecting them. The reaping loop is `os.waitpid(-1, os.WNOHANG)` called repeatedly: a non-zero PID means one was collected and you should go round again, `(0, 0)` means children exist but none has exited, and `ChildProcessError` means there are no children left. Drive it from a `signal.SIGCHLD` handler, remembering that signals coalesce so one delivery can cover several deaths, or simply run it each pass of the main loop. Without it the PID table fills and process creation eventually fails.

code

python · 19 lines
python
import os
import time

for _ in range(3):
    if os.fork() == 0:
        os._exit(7)

reaped = []
while True:
    try:
        pid, status = os.waitpid(-1, os.WNOHANG)
    except ChildProcessError:
        break
    if pid == 0:
        time.sleep(0.01)
        continue
    reaped.append((pid, os.waitstatus_to_exitcode(status)))

print(len(reaped), "children reaped:", reaped)

go deeper

for a junior

Recall what a defunct process is: an exited process whose status nobody has collected yet. Know that it cannot be killed, because it is already dead, and that its parent has to wait for it.

for a middle

Explain reparenting and the mechanics of os.waitpid with os.WNOHANG: the meaning of a -1 PID, the (0, 0) return, and the ChildProcessError that ends the loop. Be able to write the loop from memory.

for a senior

Show the operational picture — a slow PID leak that only surfaces as failed process creation hours later — and the design choices between a SIGCHLD handler and a main-loop tick, including the race with code that waits for its own children.

for a principal

Argue about ownership: whether your services should implement init duties at all, or delegate them to a dedicated first process, and what you standardise so every team is not rediscovering orphan reaping on its own.

### Why zombies exist at all When a process exits, the kernel cannot forget it immediately: somebody may still want the exit status. So it keeps a minimal entry — no memory, no file descriptors, just the PID and how the process ended. That entry is a zombie, shown as `<defunct>` in a process listing. It disappears when the parent collects the status with a wait call. A zombie costs almost nothing individually; what it costs is a PID, and PIDs are a finite, per-namespace resource. ### Why PID 1 gets them When a process dies while its own children are still alive, those children are orphaned and reparented — they get a new parent, and that parent is PID 1. A program running at PID 1 therefore acquires children it never created and whose lifecycle it knows nothing about. A traditional init exists partly to do this one chore forever. An interpreter dropped into that slot inherits the chore without inheriting the code. The failure is quiet and slow. Consider a geocoding batch runner sitting at PID 1, spawning short-lived worker processes to keep up with a 1,200-request-per-minute peak. A bad deployment makes each worker die within milliseconds of start — a circular import raising `ImportError` before the worker ever does any work. The runner notices only that throughput dropped. Meanwhile every dead worker, and every helper those workers had started, has been reparented and is sitting defunct. Nothing logs an error, memory looks fine, and then hours later process creation starts failing with `BlockingIOError` or `OSError` because the namespace has run out of PIDs. The original crash and the eventual outage look like unrelated incidents. ### The reaping loop The primitive is `os.waitpid(pid, options)`. Two arguments make it an init loop: * `-1` as the PID means *any* child, not a specific one — the point, since orphans are children you did not choose. * `os.WNOHANG` makes the call non-blocking, so a program with real work to do can reap without stalling. The three outcomes have to be handled distinctly: * a returned `(pid, status)` with a non-zero PID — one was collected; loop again immediately, because there may be more; * `(0, 0)` — there are children, but none has exited yet; stop looping for now; * `ChildProcessError` — there are no children at all; stop looping. `os.waitstatus_to_exitcode(status)` turns the raw status into the familiar exit code, returning a negative number when the process died from a signal. ### When to run it Two designs, and both are defensible: **On `signal.SIGCHLD`.** Install a handler that runs the loop. The critical detail is that standard signals do not queue: three children dying at nearly the same moment can produce a single delivery. A handler that reaps exactly one child leaks the rest, which is why the handler must loop until it sees `(0, 0)` or the exception. Remember too that CPython runs the handler in the main thread at a bytecode boundary, so it must stay short. **In the main loop.** If the program already has a periodic tick, calling the same non-blocking loop once per pass is simpler and avoids signal re-entrancy entirely. Latency of a few hundred milliseconds in collecting a zombie is irrelevant. ### The interaction that bites `os.waitpid(-1, ...)` collects *any* child, including ones another part of the program is waiting for. If your reaper wins the race, a `subprocess.Popen` object's own wait may find the status already consumed, and higher-level machinery that expects to see its worker die can be left confused. The usual discipline is to keep the two worlds separate: let the library that spawned a process collect that process, and have the reaper tolerate the fact that some statuses will be taken from under it. Track the PIDs you deliberately manage, and treat everything else the loop returns as an inherited orphan you simply discard. ### Verifying it works Fork a couple of children that exit immediately, then run the loop and count what it collects. In a live process, a count of collected orphans exported as a metric is a cheap early warning: a rising count of children you never spawned means something below you is crashing, long before the PID table runs dry.

  • Why must a signal.SIGCHLD handler loop instead of reaping a single child?
    Standard signals are not queued. If several children exit while one delivery is already pending, the kernel collapses them into one, so a handler that calls `os.waitpid` once collects one status and leaves the others defunct forever. The handler must loop until `os.waitpid(-1, os.WNOHANG)` returns `(0, 0)` or raises `ChildProcessError`.
  • What is the risk of a global reaper in a program that also manages its own child processes?
    `os.waitpid(-1, ...)` collects any child, including one that another component intends to wait for. Whoever calls first consumes the status, and the other side sees its wait fail or hang on bookkeeping that will never resolve. Keep a set of PIDs you deliberately own, let the code that spawned them collect them, and have the reaper discard everything else.
  • How would you detect this problem in a running system before it causes an outage?
    Count defunct entries and compare children against the PIDs you spawned yourself. Exporting the number of orphans the reap loop collected turns a silent leak into a visible signal, and a rising count means something beneath you is crashing. The hard failure is process creation failing once the namespace exhausts its PIDs, which arrives long after the cause.

saying these in an interview costs you the question

  • Thinks zombies free themselves after a timeout
  • Tries to remove a zombie by signalling it
  • Reaps exactly one child per SIGCHLD delivery
  • Uses a blocking wait in the main loop
  • Assumes only processes you spawned become your children
  • Confuses a zombie with a runaway process consuming memory

context