skip to content

Why is a subprocess.Popen child's returncode -9 rather than 137 when a signal ends it?

level: middleimportance: should knowfreq 52%

answer

  1. The kernel returns more than one field
  2. The library folds it into one integer
  3. Sign carries the kind of ending
  4. 128 plus N belongs to a shell
  5. Two macros decode a raw status

basics

~20 s

Python's subprocess module decodes the raw wait status itself: a non-negative value is the child's own exit status, and -N means signal N killed it, so SIGKILL reports -9. The 128+N form is a shell convention, not Python's.

solid answer

~50 s

When a child ends, the kernel hands the parent a packed wait status encoding two different endings: the process exited with a status, or a signal terminated it. Raw, you take that apart with `os.WIFEXITED` plus `os.WEXITSTATUS`, or `os.WIFSIGNALED` plus `os.WTERMSIG`. The `subprocess` module does that decoding and collapses it into one integer — the `returncode` attribute of the `subprocess.Popen` object — where non-negative means a normal exit with that status and negative means death by signal `-returncode`. A child killed with SIGKILL reports -9; one that took SIGTERM reports -15. The 128+N numbers people expect are what a shell prints for a signalled child; Python never produces them. So `if code != 0` is not enough when you care whether you killed the child or it failed on its own: test the sign, and `signal.Signals(-code).name` gives a readable name.

code

python · 6 lines
python
import os, signal, subprocess, sys

child = subprocess.Popen([sys.executable, "-c", "import time; time.sleep(30)"])
os.kill(child.pid, signal.SIGKILL)
status = child.wait()
print(status, signal.Signals(-status).name)

go deeper

for a junior

Remember that a negative value from a child process is not an error code the program chose: it names the signal that killed it, with -9 for SIGKILL and -15 for SIGTERM.

for a middle

Explain the packed wait status and the two decode paths — os.WIFEXITED with os.WEXITSTATUS versus os.WIFSIGNALED with os.WTERMSIG — and why the subprocess module can collapse both into one signed integer.

for a senior

Show what you do differently when a child reports -9 versus a positive status: memory limits and supervisor timeouts for the former, input and dependency problems for the latter, and a native-crash bug report for -11.

for a principal

Own how process-outcome information reaches the team: which statuses page someone, which retry, and whether cross-platform workers can rely on a signal-based convention at all when Windows reports plain exit codes.

## The kernel's answer is not a single number A process on a POSIX system can end in two structurally different ways. It can finish under its own control and hand the kernel an exit status — the integer given to `exit()`, or returned from `main`, or carried by `SystemExit` in Python. Or it can be terminated from outside by a signal it did not handle: SIGKILL from an out-of-memory killer, SIGTERM from an orchestrator, SIGSEGV from a native crash inside an extension module. When the parent collects the result, the kernel does not give it "the exit code". It gives back a packed *wait status* in which several fields are encoded, and the parent has to ask which kind of ending it was before the number inside means anything. At the raw level, Python exposes exactly the C macros for that job. `os.WIFEXITED(status)` is true when the child exited normally, and then `os.WEXITSTATUS(status)` is the status it passed. `os.WIFSIGNALED(status)` is true when a signal killed it, and then `os.WTERMSIG(status)` is that signal's number, with `os.WCOREDUMP(status)` telling you whether the kernel wrote a core image. These are the only correct way to read a status that came out of `os.waitpid`; applying `os.WEXITSTATUS` to a status whose child was signalled produces a meaningless number. ## What subprocess does with it The `subprocess` module hides that two-field structure behind a single Python integer, stored as the `returncode` attribute of a `subprocess.Popen` object and as `returncode` on the `subprocess.CompletedProcess` returned by `subprocess.run`. Its encoding is simple and worth memorising: * `0` — the child exited normally with status 0, the success convention. * a positive `N` — the child exited normally with status `N`. * a negative `-N` — the child was killed by signal `N`; nothing about its own logic is in that number. * `None` on a `subprocess.Popen` object — the child has not been waited for yet, so no status has been collected. So -9 means signal 9, SIGKILL. -15 means SIGTERM. -11 means SIGSEGV, which on a Python child almost always means a segmentation fault inside a C extension rather than anything the interpreter did. `signal.Signals(-code).name` converts the number to the symbolic name for a log line, and it is worth guarding that conversion so it only runs when the code is actually negative. ## Where 137 comes from, and why Python does not print it People expect 137 because that is the number they have seen from a shell or a container runtime. Those report a signalled child by adding 128 to the signal number — 128+9 = 137, 128+15 = 143 — because a shell has only one byte of exit status to pass upward and needs to fold both endings into it. Python has no such constraint: it hands you a full Python integer and can afford to make signal death a distinct range rather than an overlapping one. The two conventions describe the same event; only one of them is what your Python code will see. Mapping between them by hand (`137 - 128`) when you must match a log line from elsewhere is fine, but do not expect the runtime to do it. ## Why the distinction matters in practice On a service that spawns work, "the child failed" and "something killed the child" call for entirely different responses. A positive status is the program's own report and usually means bad input, a missing dependency, or a genuine tool error — retrying it unchanged will fail again. A negative status means an external force intervened: -9 is very often the kernel's OOM killer or a supervisor losing patience, so the correct response is to look at memory limits and timeouts, not at the input. -11 points at a native crash and belongs in a bug report against the extension. Code that only tests `if code != 0` throws all of that away and produces alerts that cannot be triaged. ## Platform caveat The negative convention is POSIX-specific, because signals are. On Windows there are no signals in this sense; the value is the process exit code the operating system reports, which is a plain unsigned number and can be large. Cross-platform code that branches on a negative status needs to accept that on Windows the branch simply never fires, and that the equivalent information — a process that was forcibly ended — arrives as an ordinary exit code instead.

  • How would you turn a negative return code into a readable signal name in a log message?
    Guard on the sign first, then pass the negated value through the signal.Signals enum: for a code of -9, signal.Signals(9).name is "SIGKILL". Doing it unguarded on a positive code either raises ValueError or, worse, names an unrelated signal for a small positive exit status, so the branch on negativity is part of the idiom rather than an optimisation.
  • What does os.WCOREDUMP tell you that the single collapsed integer cannot?
    Whether the kernel wrote a core image for the signalled child, which is the difference between having a crash to debug and only knowing that one happened. It is meaningful only when os.WIFSIGNALED is true, and it needs the raw status from os.waitpid — the subprocess module's single integer has already discarded that bit.
  • Can a child on Windows ever report a negative return code?
    No. The negative convention encodes a POSIX signal number, and Windows has no signal-terminated wait status. There the value is the process exit code the operating system reports, and a forcibly ended process shows whatever exit code the terminating call supplied. Cross-platform code should treat the negative branch as a POSIX-only path.

saying these in an interview costs you the question

  • Reads -9 as an error code the program itself returned
  • Expects 137 because that is what a shell displays
  • Only tests for non-zero and never distinguishes signal death
  • Passes a subprocess return code to os.WTERMSIG instead of a raw status
  • Thinks a negative code means the child failed to start

context