skip to content

Group Termination and Orphans

Terminating a child signals only that child, so a shell wrapper leaves its grandchildren running. Putting the child in its own session and signalling the whole group is the fix interviewers want.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How do subprocess.Popen.terminate() and Popen.kill() differ on Unix?

level: juniorimportance: must knowfreq 62%

answer

  1. Two different signals, not two strengths
  2. One of them can be refused
  3. The other never runs child code
  4. SIGTERM versus SIGKILL
  5. Both aim at one pid only

basics

~20 s

Popen.terminate() sends SIGTERM, which the child may catch and use to shut down cleanly. Popen.kill() sends SIGKILL, which no process can catch, block or ignore, so the child dies immediately with no cleanup. Both signal only the direct child.

solid answer

~50 s

On Unix both are thin wrappers over `Popen.send_signal()`, which calls `os.kill()` on the child's pid: `terminate()` sends `signal.SIGTERM`, `kill()` sends `signal.SIGKILL`. The difference is not speed but agency. SIGTERM is a request the child can handle with `signal.signal()` — flush a file, drop a lock, exit with its own status — or ignore entirely, so `terminate()` is never a guarantee. SIGKILL is enforced by the kernel and delivers no code to the child at all: no handler, no `finally` block, no `atexit` callback, no buffer flush, so half-written files and stale lock files are exactly what you should expect. On Windows there are no POSIX signals and both methods perform the same abrupt process termination, so the graceful path does not exist there. Neither call reaches anything but the direct child pid, and neither reaps it — you still call `Popen.wait()`.

code

python · 17 lines
python
import subprocess, sys, time

CHILD = """
import signal, sys, time
signal.signal(signal.SIGTERM, lambda *a: (print("child: cleaning up"), sys.exit(0)))
time.sleep(30)
"""

p = subprocess.Popen([sys.executable, "-c", CHILD])
time.sleep(0.5)
p.terminate()
print("after terminate:", p.wait())

q = subprocess.Popen([sys.executable, "-c", CHILD])
time.sleep(0.5)
q.kill()
print("after kill:", q.wait())

go deeper

for a junior

Recall the pairing: terminate means SIGTERM and can be handled, kill means SIGKILL and cannot. Be able to say why you would try the first one before the second.

for a middle

Explain the mechanics: both route through Popen.send_signal to os.kill on one pid, SIGKILL is refused by signal.signal, and no cleanup code of any kind runs when it is delivered.

for a senior

Show the production judgement: a grace period between the two, an explicit check that the child actually died, and a clear account of what state is left behind when you do escalate.

for a principal

Own the policy question — how long a shutdown budget the platform grants children, what cleanup is allowed to depend on SIGTERM at all, and how that changes on Windows where the graceful path does not exist.

`subprocess.Popen.terminate()` and `subprocess.Popen.kill()` read like two strengths of one operation. On Unix they are two different signals, and what separates them is not how fast the child dies but whether the child gets to run any of its own code on the way out. ## What the two calls actually do Both are one-line wrappers around `Popen.send_signal()`. `send_signal()` first checks whether the child has already been reaped (if so it does nothing, because the pid may since have been recycled by the kernel onto an unrelated process), then calls `os.kill(self.pid, sig)`. `terminate()` passes `signal.SIGTERM`; `kill()` passes `signal.SIGKILL`. That is the whole implementation. Anything you can say about the two methods is a statement about the two signals. ## SIGTERM is a request SIGTERM is the conventional "please stop" signal, and it is deliverable to user code. A program can install a handler with `signal.signal(signal.SIGTERM, handler)` and use it to close a database connection, rename a temp file into place, remove a pid file, and exit with a status of its own choosing. It can also set the disposition to `signal.SIG_IGN` and ignore the signal outright, or block it while it is inside a critical section and take delivery later. That is the point of SIGTERM and also its limitation: `terminate()` is a polite request that a well-behaved program honours, not an assurance that the process is gone. A child that is wedged, that has ignored the signal, or that is blocked in a call the handler cannot interrupt will still be there afterwards. Code that sends SIGTERM and immediately assumes success is the single most common bug in this area. ## SIGKILL is not deliverable to user code SIGKILL is special-cased by the kernel: it cannot be caught, cannot be blocked and cannot be ignored. `signal.signal(signal.SIGKILL, ...)` raises `OSError` rather than installing anything. When SIGKILL is delivered, the kernel destroys the process. Nothing of the child runs: no signal handler, no `finally` block, no `atexit` callback, no `__del__`, no flush of buffered writes still sitting in the process's own memory. The practical consequences are entirely predictable, and an interviewer wants to hear you name them: a file the child was midway through writing stays truncated, a lock file or temp directory it would have removed stays on disk, an in-memory batch it had not yet written is lost. Those are the costs you are choosing to accept when you escalate. The one state SIGKILL cannot cut through is a process stuck in an uninterruptible kernel wait inside a driver — it dies as soon as that call returns, and not before. ## Both reach only the direct child Neither method touches anything except the one pid that `Popen.pid` holds. If the child is itself a wrapper — a shell started with `shell=True`, a launcher script, a supervisor — its own children receive nothing and keep running after the wrapper dies. Reaching a whole tree is a separate mechanism (a process group), not a stronger signal. Neither method reaps the child either. After signalling you still call `Popen.wait()`, both to let the kernel hand back the exit status and to know when the process is actually gone rather than merely signalled. ## Windows Windows has no POSIX signal delivery. There, `terminate()` and `kill()` are the same method: both ask the platform to terminate the process abruptly with an exit code. A Windows child therefore never gets the graceful path, and portable code that relies on a clean shutdown has to arrange one itself — an IPC message, a sentinel file, or a protocol on the child's stdin — rather than assuming SIGTERM semantics. ## Which one to use The default should always be `terminate()` first, then a bounded grace period in which you wait for the child, then `kill()` if it is still alive. Reaching for `kill()` immediately trades a corrupt output file for a couple of saved seconds, which is almost never the trade you want. Reaching for `terminate()` and never checking leaves you with a process you believe is dead and is not. When you need something other than those two signals — SIGINT to emulate Ctrl-C, SIGHUP to ask a daemon to reload, SIGUSR1 for an application-defined action — use `Popen.send_signal()` directly with the constant from the `signal` module. `terminate()` and `kill()` are just the two shortcuts the standard library chose to name.

  • Why is calling Popen.kill() straight away, with no SIGTERM first, a poor default?
    Because SIGKILL runs none of the child's code. Anything the child would have done on the way out — completing a write, renaming a temp file into place, releasing a lock, removing a pid file — does not happen, so you trade a corrupted or half-finished artefact for a few saved seconds. SIGKILL is the right tool only after a bounded grace period has proved the child will not stop on its own.
  • Do Popen.terminate() and Popen.kill() behave the same way on Windows?
    Yes. Windows has no POSIX signals, so both call the same abrupt platform termination and the child exits with a fixed code without running any cleanup. The graceful/forceful distinction is Unix-only, so portable shutdown code cannot rely on a SIGTERM handler in the child and needs its own protocol — a message on stdin, a sentinel file, or an IPC request — to ask a Windows child to stop cleanly.
  • What does Popen.send_signal() give you that terminate() and kill() do not?
    Any signal you like. `send_signal(signal.SIGINT)` emulates Ctrl-C so a child that already handles KeyboardInterrupt shuts down through its normal path; SIGHUP is the conventional "reload your config" signal for daemons; SIGUSR1 and SIGUSR2 are reserved for application-defined meanings. `terminate()` and `kill()` are simply the two shortcuts the standard library named, and they route through `send_signal()` themselves.

SIGTERM is knocking on the door and asking someone to leave; SIGKILL is the building being demolished around them. Only the first gives them time to take anything with them.

saying these in an interview costs you the question

  • Says the child can install a handler for SIGKILL
  • Thinks terminate() and kill() differ only in speed
  • Assumes either call kills the whole process tree
  • Believes terminate() guarantees the child has exited
  • Expects temp files to be cleaned up after SIGKILL
  • Assumes SIGTERM semantics hold on Windows

context

open as a page

Why does Popen.terminate() leave a shell=True command's real program running?

level: middleimportance: must knowfreq 58%

basics

~20 s

With shell=True the direct child is /bin/sh, so Popen.pid is the shell's pid and terminate() signals the shell, not the program it started. Start the child with start_new_session=True and signal the whole group with os.killpg(os.getpgid(p.pid), signal.SIGTERM).

open as a page

Popen.terminate() sometimes fails to stop a translation-memory updater's export child; how would you build a shutdown that always ends it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Signal the child's whole process group with SIGTERM, wait a bounded grace period, then send SIGKILL to the same group and wait again. Start the child with start_new_session=True so that group exists and excludes your own process.

open as a page

What happens to subprocess.Popen children when the Python parent exits first?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

They keep running. A child whose parent dies is an orphan and is reparented by the kernel to PID 1, or to the nearest ancestor marked as a subreaper. Python never kills them for you, so cleanup has to be explicit.

open as a page