How should a Python PID 1 forward signal.SIGTERM to the children it started?
answer
- Signals do not cascade down a tree
- The handler only records the request
- Pass it on, then wait on a clock
- Escalate to the unhandleable one
- Own group per child covers grandchildren
basics
~20 sA signal sent to PID 1 reaches only that process; children see nothing. Record the request in the handler, then have the main loop forward it with subprocess.Popen.send_signal, wait on a shared deadline, and escalate to signal.SIGKILL.
solid answer
~40 sSignalling a PID delivers to that process alone, so a supervisor's stop request lands on the Python process at PID 1 and its worker processes never hear about it. Forwarding is explicit work. Install a handler for `signal.SIGTERM` that only records the request, and in the main loop call `subprocess.Popen.send_signal` — or `os.kill(pid, signal.SIGTERM)` for a raw PID — on each child you own. Then wait with a deadline: `subprocess.Popen.wait(timeout=...)` and, on `subprocess.TimeoutExpired`, `subprocess.Popen.kill()` to escalate. If a child spawns a tree of its own, start it in its own process group with Popen's start_new_session argument and signal the group with `os.killpg`, so the grandchildren are covered too. Only exit once the children are collected, or the exit status you report is a lie about what actually stopped.
code
python · 23 linesimport os
import signal
import subprocess
import sys
child = subprocess.Popen([sys.executable, "-c", "import time; time.sleep(30)"])
pending = []
def on_signal(signum, frame):
pending.append(signum)
signal.signal(signal.SIGTERM, on_signal)
os.kill(os.getpid(), signal.SIGTERM)
if pending:
child.send_signal(pending[0])
try:
print("child exit status:", child.wait(timeout=5))
except subprocess.TimeoutExpired:
child.kill()
print("child killed:", child.wait())go deeper
Recall that a signal goes to one process only. If your program starts other processes, it has to tell them to stop itself; nothing does that for you.
Explain the sequence: record in the handler, send with subprocess.Popen.send_signal or os.kill, wait with a timeout, escalate with kill. Know why the handler must stay small and why start_new_session changes the reach of os.killpg.
Show judgement about the clock — one shared deadline rather than per-child timeouts, intake stopped before forwarding, a truthful exit status recorded — and be able to describe the silent data loss when children are abandoned mid-work.
Decide where this responsibility lives: a shared supervisor component, a dedicated init process, or duplicated in every service. Own the grace-window budget and how forwarding behaviour is verified rather than assumed.
### Delivery is per-process, not per-tree The first thing to be clear about is that sending a signal to a PID delivers it to exactly that process. There is no cascade to children. The reason a keyboard interrupt appears to stop a whole pipeline is a different mechanism entirely — the terminal signals a foreground process *group*, and everything in the group is a target. A supervisor stopping an isolated environment has no such group semantics to lean on: it signals the first process and nothing else. So a Python program at PID 1 that spawned worker processes is in a peculiar position. It may have a perfectly good handler installed, it may shut its own loop down cleanly, and its children will still be running when it exits — at which point they are orphaned, reparented, and killed only when the environment itself is torn down. Anything they were part-way through is lost, and the shutdown looked clean from the outside. ### The forwarding pattern Four steps, in order. **1. Record the request.** The handler installed for `signal.SIGTERM` should do as little as possible — set a `threading.Event` or append to a list. CPython runs handlers as ordinary Python in the main thread at a bytecode boundary, so a handler that starts signalling and waiting can be re-entered by a second delivery halfway through. **2. Pass it on.** For a child held as a `subprocess.Popen` object, `subprocess.Popen.send_signal(signal.SIGTERM)` — or the equivalent `subprocess.Popen.terminate()` — is the direct route. For a bare PID from `os.fork`, use `os.kill(pid, signal.SIGTERM)`. Forward the signal you actually received rather than hardcoding one, so an interrupt and a terminate request stay distinguishable to the child. **3. Wait with a deadline.** `subprocess.Popen.wait(timeout=...)` blocks until the child exits or raises `subprocess.TimeoutExpired`. Compute the deadline once for the whole shutdown and give each child what remains of it, rather than granting every child a fresh full timeout — with a handful of workers, per-child timeouts silently multiply the total and blow past whatever grace period the outside supervisor allows. **4. Escalate.** On timeout, `subprocess.Popen.kill()` sends `signal.SIGKILL`, which the child cannot handle or ignore. Then collect it, so it does not linger defunct. ### Process groups and grandchildren Forwarding to direct children is not enough when a child spawns its own tree. The tool for that is the process group. Starting a child with Popen's start_new_session argument puts it in a new session and a new process group whose ID equals the child's PID, so `os.killpg(child.pid, signal.SIGTERM)` reaches the child and everything it started. The reason not to simply signal your *own* group is that you are in it: `os.killpg(os.getpgid(0), signal.SIGTERM)` re-delivers the signal to yourself and, with a handler installed that forwards to the group, you have written an infinite loop. ### The failure this prevents Take a geocoding batch runner sitting at PID 1 and driving worker processes to sustain a 1,200-request-per-minute peak. Each worker holds a batch of in-flight requests and writes results at the end of a batch. A stop request arrives, the runner exits cleanly, and every worker is killed mid-batch by the environment teardown. The runner's own logs say it shut down gracefully. The real result is a peak-hour window of dropped work with no error anywhere, discovered later as a gap in the output. Forwarding turns that into each worker finishing or abandoning its batch deliberately, on a bounded clock. ### Details worth getting right * **Report a truthful status.** `subprocess.Popen.wait` returns a negative number when the child died from a signal — `-15` for a terminate, `-9` for a kill. Logging that distinguishes a child that cooperated from one you had to force. * **Do not double-handle.** If a child is in its own group and you both `os.killpg` the group and `send_signal` the child, the child gets the signal twice; harmless for a well-written child, confusing for one that counts them. * **Second interrupt means impatience.** A common convention is that the first signal starts a graceful drain and a second one escalates immediately, which is easy to express when the handler is only recording requests. * **Ordering.** Stop accepting new work first, then forward, then wait. Forwarding before you have stopped intake means a worker can be handed a new item after being asked to stop.
- Why not just signal your own process group so every descendant gets it at once?Because you are a member of that group. `os.killpg(os.getpgid(0), signal.SIGTERM)` delivers the signal to the sender as well, so a handler that forwards by signalling its own group re-triggers itself indefinitely. Give each child its own session and group with Popen's start_new_session argument, then signal `os.killpg(child.pid, ...)`, which covers the child's descendants without including you.
- How do you bound the total shutdown time when there are several children?Compute one deadline when the request arrives and give each `subprocess.Popen.wait` call whatever remains of it. Per-child timeouts add up — five children at ten seconds each is fifty seconds, long past most grace periods — and the outside supervisor will force-kill you mid-drain. With a shared deadline, whatever has not stopped by the end gets `subprocess.Popen.kill()`.
- What does subprocess.Popen.wait return for a child that died from a signal?A negative number whose magnitude is the signal number: `-15` for a terminate, `-9` for a forced kill. A non-negative value is an ordinary exit code. Logging that difference tells you whether a child drained on request or had to be escalated, which is exactly the data you need to tune a grace window.
saying these in an interview costs you the question
- Assumes children inherit a signal sent to the parent
- Exits immediately after forwarding, without waiting
- Gives each child a fresh full timeout
- Signals its own process group and loops forever
- Does the whole drain inside the signal handler
- Never escalates, so one stuck child hangs the shutdown