skip to content

How should a Python PID 1 forward signal.SIGTERM to the children it started?

level: middleimportance: should knowfreq 30%

answer

  1. Signals do not cascade down a tree
  2. The handler only records the request
  3. Pass it on, then wait on a clock
  4. Escalate to the unhandleable one
  5. Own group per child covers grandchildren

basics

~20 s

A signal sent to PID 1 reaches only that process; children see nothing. Record the request in the handler, then have the main loop forward it with subprocess.Popen.send_signal, wait on a shared deadline, and escalate to signal.SIGKILL.

solid answer

~40 s

Signalling a PID delivers to that process alone, so a supervisor's stop request lands on the Python process at PID 1 and its worker processes never hear about it. Forwarding is explicit work. Install a handler for `signal.SIGTERM` that only records the request, and in the main loop call `subprocess.Popen.send_signal` — or `os.kill(pid, signal.SIGTERM)` for a raw PID — on each child you own. Then wait with a deadline: `subprocess.Popen.wait(timeout=...)` and, on `subprocess.TimeoutExpired`, `subprocess.Popen.kill()` to escalate. If a child spawns a tree of its own, start it in its own process group with Popen's start_new_session argument and signal the group with `os.killpg`, so the grandchildren are covered too. Only exit once the children are collected, or the exit status you report is a lie about what actually stopped.

code

python · 23 lines
python
import os
import signal
import subprocess
import sys

child = subprocess.Popen([sys.executable, "-c", "import time; time.sleep(30)"])
pending = []


def on_signal(signum, frame):
    pending.append(signum)


signal.signal(signal.SIGTERM, on_signal)
os.kill(os.getpid(), signal.SIGTERM)

if pending:
    child.send_signal(pending[0])
try:
    print("child exit status:", child.wait(timeout=5))
except subprocess.TimeoutExpired:
    child.kill()
    print("child killed:", child.wait())

go deeper

for a junior

Recall that a signal goes to one process only. If your program starts other processes, it has to tell them to stop itself; nothing does that for you.

for a middle

Explain the sequence: record in the handler, send with subprocess.Popen.send_signal or os.kill, wait with a timeout, escalate with kill. Know why the handler must stay small and why start_new_session changes the reach of os.killpg.

for a senior

Show judgement about the clock — one shared deadline rather than per-child timeouts, intake stopped before forwarding, a truthful exit status recorded — and be able to describe the silent data loss when children are abandoned mid-work.

for a principal

Decide where this responsibility lives: a shared supervisor component, a dedicated init process, or duplicated in every service. Own the grace-window budget and how forwarding behaviour is verified rather than assumed.

### Delivery is per-process, not per-tree The first thing to be clear about is that sending a signal to a PID delivers it to exactly that process. There is no cascade to children. The reason a keyboard interrupt appears to stop a whole pipeline is a different mechanism entirely — the terminal signals a foreground process *group*, and everything in the group is a target. A supervisor stopping an isolated environment has no such group semantics to lean on: it signals the first process and nothing else. So a Python program at PID 1 that spawned worker processes is in a peculiar position. It may have a perfectly good handler installed, it may shut its own loop down cleanly, and its children will still be running when it exits — at which point they are orphaned, reparented, and killed only when the environment itself is torn down. Anything they were part-way through is lost, and the shutdown looked clean from the outside. ### The forwarding pattern Four steps, in order. **1. Record the request.** The handler installed for `signal.SIGTERM` should do as little as possible — set a `threading.Event` or append to a list. CPython runs handlers as ordinary Python in the main thread at a bytecode boundary, so a handler that starts signalling and waiting can be re-entered by a second delivery halfway through. **2. Pass it on.** For a child held as a `subprocess.Popen` object, `subprocess.Popen.send_signal(signal.SIGTERM)` — or the equivalent `subprocess.Popen.terminate()` — is the direct route. For a bare PID from `os.fork`, use `os.kill(pid, signal.SIGTERM)`. Forward the signal you actually received rather than hardcoding one, so an interrupt and a terminate request stay distinguishable to the child. **3. Wait with a deadline.** `subprocess.Popen.wait(timeout=...)` blocks until the child exits or raises `subprocess.TimeoutExpired`. Compute the deadline once for the whole shutdown and give each child what remains of it, rather than granting every child a fresh full timeout — with a handful of workers, per-child timeouts silently multiply the total and blow past whatever grace period the outside supervisor allows. **4. Escalate.** On timeout, `subprocess.Popen.kill()` sends `signal.SIGKILL`, which the child cannot handle or ignore. Then collect it, so it does not linger defunct. ### Process groups and grandchildren Forwarding to direct children is not enough when a child spawns its own tree. The tool for that is the process group. Starting a child with Popen's start_new_session argument puts it in a new session and a new process group whose ID equals the child's PID, so `os.killpg(child.pid, signal.SIGTERM)` reaches the child and everything it started. The reason not to simply signal your *own* group is that you are in it: `os.killpg(os.getpgid(0), signal.SIGTERM)` re-delivers the signal to yourself and, with a handler installed that forwards to the group, you have written an infinite loop. ### The failure this prevents Take a geocoding batch runner sitting at PID 1 and driving worker processes to sustain a 1,200-request-per-minute peak. Each worker holds a batch of in-flight requests and writes results at the end of a batch. A stop request arrives, the runner exits cleanly, and every worker is killed mid-batch by the environment teardown. The runner's own logs say it shut down gracefully. The real result is a peak-hour window of dropped work with no error anywhere, discovered later as a gap in the output. Forwarding turns that into each worker finishing or abandoning its batch deliberately, on a bounded clock. ### Details worth getting right * **Report a truthful status.** `subprocess.Popen.wait` returns a negative number when the child died from a signal — `-15` for a terminate, `-9` for a kill. Logging that distinguishes a child that cooperated from one you had to force. * **Do not double-handle.** If a child is in its own group and you both `os.killpg` the group and `send_signal` the child, the child gets the signal twice; harmless for a well-written child, confusing for one that counts them. * **Second interrupt means impatience.** A common convention is that the first signal starts a graceful drain and a second one escalates immediately, which is easy to express when the handler is only recording requests. * **Ordering.** Stop accepting new work first, then forward, then wait. Forwarding before you have stopped intake means a worker can be handed a new item after being asked to stop.

  • Why not just signal your own process group so every descendant gets it at once?
    Because you are a member of that group. `os.killpg(os.getpgid(0), signal.SIGTERM)` delivers the signal to the sender as well, so a handler that forwards by signalling its own group re-triggers itself indefinitely. Give each child its own session and group with Popen's start_new_session argument, then signal `os.killpg(child.pid, ...)`, which covers the child's descendants without including you.
  • How do you bound the total shutdown time when there are several children?
    Compute one deadline when the request arrives and give each `subprocess.Popen.wait` call whatever remains of it. Per-child timeouts add up — five children at ten seconds each is fifty seconds, long past most grace periods — and the outside supervisor will force-kill you mid-drain. With a shared deadline, whatever has not stopped by the end gets `subprocess.Popen.kill()`.
  • What does subprocess.Popen.wait return for a child that died from a signal?
    A negative number whose magnitude is the signal number: `-15` for a terminate, `-9` for a forced kill. A non-negative value is an ordinary exit code. Logging that difference tells you whether a child drained on request or had to be escalated, which is exactly the data you need to tune a grace window.

saying these in an interview costs you the question

  • Assumes children inherit a signal sent to the parent
  • Exits immediately after forwarding, without waiting
  • Gives each child a fresh full timeout
  • Signals its own process group and loops forever
  • Does the whole drain inside the signal handler
  • Never escalates, so one stuck child hangs the shutdown

context