skip to content

Why does Popen.terminate() leave a shell=True command's real program running?

level: middleimportance: must knowfreq 58%

answer

  1. Popen knows only one pid
  2. The command string is not the child
  3. Something between you and the program
  4. A new session at spawn time
  5. os.setsid then os.killpg on the group

basics

~20 s

With shell=True the direct child is /bin/sh, so Popen.pid is the shell's pid and terminate() signals the shell, not the program it started. Start the child with start_new_session=True and signal the whole group with os.killpg(os.getpgid(p.pid), signal.SIGTERM).

solid answer

~40 s

`subprocess.Popen` knows exactly one pid: the process it forked. With `shell=True` that process is `/bin/sh -c "..."`, so `terminate()` sends SIGTERM to the shell. The shell often forks the real command as its own child, and killing the shell leaves that grandchild running — reparented, holding whatever ports, locks and pipe ends it had. Two fixes. Where you control the command, drop `shell=True` and pass an argument list, so the direct child *is* the program. Where you genuinely need a shell or the child spawns its own workers, start it with `start_new_session=True` (which calls `os.setsid()` in the child after the fork), then signal the whole group with `os.killpg(os.getpgid(p.pid), signal.SIGTERM)`. The new session matters: without it the child shares your process group, and `killpg` would signal your own Python process too.

code

python · 10 lines
python
import os, signal, subprocess, time

p = subprocess.Popen("sleep 30 & wait", shell=True, start_new_session=True)
pgid = os.getpgid(p.pid)          # read it before the child is reaped
time.sleep(0.3)

p.terminate()                     # SIGTERM reaches the shell only
print("shell returncode:", p.wait())

os.killpg(pgid, signal.SIGKILL)   # the backgrounded sleep is still alive

go deeper

for a junior

Remember that with shell=True the child Python started is the shell, not your program, so terminate() may hit the wrapper. Prefer passing an argument list where you can.

for a middle

Explain the mechanics: Popen tracks one pid, start_new_session=True calls os.setsid() in the child, and os.killpg with os.getpgid signals the group the child leads.

for a senior

Demonstrate that you design for it at spawn time — the group is created when the child starts, the pgid is cached, ProcessLookupError is handled, and shutdown is verified rather than assumed.

for a principal

Own the platform choice: whether services supervise their own children at all or delegate tree lifecycle to a process supervisor or container runtime, and what that means for portability to Windows.

The gap here is between what `subprocess` tracks and what your command actually starts. `subprocess.Popen` records one number, `Popen.pid`, and it is the pid of the process the module forked and exec'd — nothing else. `Popen.terminate()` and `Popen.kill()` send their signal to that pid alone. Anything that process starts afterwards is invisible to Python. ## Why shell=True is the classic trigger With `shell=True`, `Popen` does not run your command. It runs `/bin/sh -c "your command"`, and the shell runs your command. So `Popen.pid` is the shell's pid, and the program you care about is a grandchild. What makes this genuinely nasty is that the behaviour is not consistent. Most shells optimise the simplest case: for `sh -c 'sleep 30'` — one program, no pipeline, no redirection, no `&&`, no traps — the shell may `exec` the program in place, replacing itself, so the pid you hold really is the program's and `terminate()` works. Add a pipe, a background `&`, a `cd x && cmd`, or a second command, and the shell forks instead and stays alive as the parent. Your shutdown then works in development, works in the test that used a bare command, and fails in production where the command grew a redirect. Never build termination on the assumption that the shell exec'd. The same problem exists without `shell=True` whenever the direct child is a wrapper: a launcher script that ends in a call to the real binary, a program that forks worker processes, a virtual-environment shim. `shell=True` is merely the most common instance. ## What happens to the survivor When the shell dies and the real program does not, that program becomes an orphan and is reparented (see the orphan question). It keeps running with everything it inherited: its listening socket, its file locks, its open write end of any pipe you created. It is now unreachable through your `Popen` object, because Python never knew its pid. ## The fix: a process group you can signal A process group is a set of processes that can be signalled as a unit with `os.killpg(pgid, sig)`. To use one you need the child to be in a group that contains the child's descendants and does *not* contain you. `subprocess.Popen(..., start_new_session=True)` arranges that: in the child, after the fork and before the exec, the module calls `os.setsid()`. The child becomes the leader of a new session and a new process group whose id equals its pid, and it detaches from the controlling terminal. Every process it forks afterwards inherits that group unless it deliberately leaves. Shutdown then looks like this: ```python import os, signal, subprocess p = subprocess.Popen(cmd, shell=True, start_new_session=True) pgid = os.getpgid(p.pid) os.killpg(pgid, signal.SIGTERM) ``` Read the group id with `os.getpgid(p.pid)` rather than assuming it equals `p.pid`, and read it *early* — once the child has exited and been reaped by `Popen.wait()`, `os.getpgid()` raises `ProcessLookupError`, and worse, the pid may have been recycled onto an unrelated process. Caching the pgid right after the `Popen` call and signalling that cached value is the safe pattern. `os.killpg()` itself raises `ProcessLookupError` when the group is already empty, which is a normal outcome to catch and ignore. Since Python 3.11 there is also a `process_group` argument to `Popen`: `process_group=0` puts the child into a new process group of its own (a `setpgid` in the child) without creating a new session or detaching the terminal. Use it when you want a signalable group but still want the child to receive terminal signals. On 3.10 and earlier, `start_new_session` is the only built-in option. ## Why the new session or group is not optional If you skip it, the child stays in *your* process group. `os.killpg(os.getpgid(p.pid), signal.SIGTERM)` then resolves to your own group, and you SIGTERM your own Python process along with every other child in it. That is a spectacular bug to ship, and it is the reason the group must be created deliberately at spawn time — you cannot retrofit it after the fact from the parent. ## Side effects to know about `start_new_session=True` detaches the child from the controlling terminal, so it no longer receives the terminal's SIGINT when someone presses Ctrl-C, or SIGHUP when the terminal goes away. That is often exactly what you want for a managed worker — you become solely responsible for its lifecycle — but it means signal delivery is now entirely your job. And on Windows none of this applies: there are no POSIX process groups, `start_new_session` and `os.killpg` are not the mechanism, and tree termination is handled through platform-specific job objects instead. ## The cheaper fix first Before reaching for groups, ask whether you need a shell at all. `subprocess.Popen(["tool", "--flag", value])` makes the tool the direct child, so `terminate()` reaches it, and it also removes an entire class of quoting problems. Reserve `shell=True` plus a process group for the cases where the command really is a shell pipeline or the child really does spawn its own tree.

  • Why can't you just call os.killpg(os.getpgid(p.pid), SIGTERM) without start_new_session=True?
    Because without it the child inherits your process group. `os.getpgid(p.pid)` then returns your own group id, and `os.killpg()` delivers SIGTERM to your Python process and every other child in that group — you kill yourself along with the target. The group has to be created in the child at spawn time, which is what start_new_session (or process_group=0 on 3.11+) does.
  • When does /bin/sh exec the command instead of forking it, and can you rely on that?
    Most shells optimise the trivial case — a single program with no pipeline, redirection, background `&` or operator — by exec'ing it in place, so the shell's pid becomes the program's. You cannot rely on it: whether the optimisation applies depends on the shell, its build and the exact command text, so a command that gains a redirect months later silently breaks a shutdown path that assumed the pid was the program's.
  • What does start_new_session=True change about the child besides making it signalable as a group?
    It calls os.setsid() in the child, so the child leads a new session with no controlling terminal. It no longer receives terminal-generated signals: Ctrl-C in your terminal sends SIGINT to your foreground group and the child is not in it, and it does not get SIGHUP when the terminal closes. Every signal it receives now has to come from you.
  • How do you avoid signalling a recycled pid after the child has already exited?
    Cache the group id immediately after Popen returns, and only signal while the child is known unreaped — check Popen.poll() is None, and never call os.getpgid(p.pid) after Popen.wait() has reaped it, because that pid may already belong to an unrelated process. Wrap the signal in a try/except for ProcessLookupError, which simply means the group is already empty.

saying these in an interview costs you the question

  • Thinks Popen.pid is the pid of the named program
  • Says terminate() signals the whole process tree
  • Calls os.killpg without a new session, hitting the parent
  • Assumes the grandchild dies when the shell dies
  • Assumes the shell always execs the command
  • Reads os.getpgid(p.pid) after the child was reaped

context