skip to content

What does exit code 137 from a Python CI step mean?

level: seniorimportance: nice to knowfreq 25%

answer

  1. One byte, two different outcomes
  2. Subtract a round number first
  3. Nine is the one you cannot catch
  4. No traceback is itself the evidence
  5. Keep your own codes well below it

basics

~20 s

137 is the shell's 128+N encoding for a process killed by signal 9, usually the kernel out-of-memory killer or a container memory limit. Python did not choose it: no sys.exit() call produced 137, the runner's shell reported the kill that way.

solid answer

~50 s

A shell or CI runner reports a process that was **killed by a signal** as `128 + signal number`, because a single byte has to carry both outcomes. So 137 is 128 + 9 (SIGKILL) — in practice the out-of-memory killer or a container hitting its memory cap; 143 is 128 + 15 (SIGTERM), a supervisor or a job timeout asking the process to stop; 139 is 128 + 11 (SIGSEGV), usually a crashing native extension rather than Python code. The distinction matters because these are not application failures at all: no traceback will be in the log, and retrying without raising the memory limit will just reproduce it. The OS wait status keeps the two cases apart properly, and 128+N is the shell's flattening of that. The corollary: keep your own exit codes well under 126, so a deliberate failure is never mistaken for a kill.

code

python · 12 lines
python
import os

pid = os.fork()
if pid == 0:
    os.kill(os.getpid(), 9)

_, status = os.waitpid(pid, 0)
if os.WIFSIGNALED(status):
    sig = os.WTERMSIG(status)
    print(f"killed by signal {sig}; a shell would report {128 + sig}")
else:
    print("exited with", os.WEXITSTATUS(status))

go deeper

for a junior

Recognize the pattern rather than memorizing a table: a status above 128 in a shell or CI log means the process was killed by a signal, and the signal number is the value minus 128. 137 is the common one.

for a middle

Decode the common values and explain where they come from: the shell flattens killed-by-signal into 128+N. Know that 137 is SIGKILL, that Python never produces it itself, and that the absence of a traceback is the confirming evidence.

for a senior

Drive the diagnosis from the number. Separate a memory kill from a shutdown signal, point to the evidence outside the process log, and explain why retries and broader exception handling cannot help against a signal the process never receives.

for a principal

Own exit codes as an operational contract: which codes your tools may use, what a supervisor or pipeline should do with each, and why the range must stay below 126. Make sure alerting distinguishes a resource kill from an application failure rather than treating every non-zero step the same.

## Why one number has to carry two different outcomes A process can end in two fundamentally different ways: it can finish and hand back a status of its own choosing, or it can be killed by a signal it never got to respond to. The operating system keeps these apart cleanly — the wait status a parent retrieves records which case occurred and either the exit code or the signal number. A shell, though, exposes only a single number in `$?`, and a CI runner surfaces only that number as the step's result. The convention that squeezes both into one byte is **128 + N**: a process killed by signal N is reported as 128 plus N. So the numbers you actually meet in a build log decode like this: | reported | signal | what it usually means | |---|---|---| | 137 | 9, SIGKILL | out-of-memory kill, or a container memory limit; unkillable-process cleanup | | 143 | 15, SIGTERM | a supervisor, job timeout or orchestrator asked the process to stop | | 139 | 11, SIGSEGV | a memory fault, nearly always inside a native extension or the interpreter itself | | 134 | 6, SIGABRT | an abort from C-level code, often a failed assertion in a compiled library | | 130 | 2, SIGINT | an interrupt, typically Ctrl-C or a cancelled interactive run | The two that dominate real incidents are 137 and 143, and the difference between them is the whole diagnosis: 137 means something took the process away with no chance to react, while 143 means something asked politely and the process either declined or ran out of grace period. ## Reading 137 correctly The most common cause by far is memory. A step that trains, parses or holds a large working set grows past the memory the runner or container allows, and the kernel's out-of-memory killer picks it off. The signature is unmistakable once you know it: **no Python traceback in the log**. An out-of-memory condition Python itself notices raises `MemoryError` and prints a traceback; a kill at 137 leaves nothing, because the process was not running any more by the time anything could be written. Output truncated mid-line is another tell — the buffered tail was never flushed. The practical consequences follow directly. Retrying the step unchanged reproduces it, because nothing about the workload changed. Adding `try`/`except` around the code does nothing, because SIGKILL cannot be caught or handled. The fixes are all outside the program: raise the memory limit, shrink the working set, stream instead of loading, or split the job. 143 reads differently. Something sent SIGTERM — a job timeout, a deploy rolling the container, an operator stopping the service. Here the program *could* have responded, and if 143 shows up on a service that was supposed to shut down cleanly, the question is why it did not finish in the grace period before the follow-up kill. ## What Python can and cannot produce Python's own exit statuses are small and deliberate: 0 for success, 1 for an uncaught exception, and whatever integer you pass to `sys.exit()`. The interpreter never chooses 137 for you. That leads to the trap worth naming in an interview: `sys.exit(137)` is legal, and in a CI log it is **indistinguishable** from a genuine SIGKILL. An engineer who reads 137 will start hunting a memory problem that does not exist. The same applies to any code at or above 128, and to 126 and 127, which shells reserve for "found but not executable" and "command not found". The discipline is therefore to keep application exit codes in the range 1–125, and to give them meaning: for example 1 for an unexpected failure, 2 for bad usage, and a small set of distinct codes for conditions a supervisor should treat differently — say, a retryable upstream error versus a fatal configuration error. Remember too that only the low eight bits of the status survive, so a code above 255 wraps and can land in signal territory by accident: `sys.exit(400)` reports as 144, which reads like 128 + 16. ## Seeing the real distinction from Python When your own code supervises children, do not work with the flattened number at all — the wait status carries both facts, and the `os` module exposes the accessors for it: ```python import os pid = os.fork() if pid == 0: os.kill(os.getpid(), 9) _, status = os.waitpid(pid, 0) if os.WIFSIGNALED(status): sig = os.WTERMSIG(status) print(f"killed by signal {sig}; a shell would report {128 + sig}") else: print("exited with", os.WEXITSTATUS(status)) ``` That prints `killed by signal 9; a shell would report 137`, which is the convention made explicit: the 128 offset is bookkeeping the shell adds, not something the dying process reported. ## The interview-shaped answer Given "the step exited 137", a strong answer moves in three beats. First, decode it: 128 + 9, so the process was SIGKILLed and did not choose its own status. Second, name the likely cause and the confirming evidence: a memory limit or the OOM killer, confirmed by the absence of a traceback, truncated output, and the runner's or kernel's memory records. Third, state what will and will not help: raising the limit or reducing the working set will; a retry, a broader `except`, or a `finally` block will not, because nothing in the process ran after the signal arrived.

  • How do you tell a 137 caused by memory from one caused by a job timeout?
    By what else the environment recorded. A memory kill leaves an out-of-memory entry in the kernel or container runtime's records and typically arrives partway through a heavy phase; a timeout kill usually sends SIGTERM first, so you would see 143, and only escalates to 137 if the process ignored it past the grace period. In both cases the process log itself is silent, which is why the surrounding evidence is the diagnosis.
  • Why can a Python program not catch the condition that produces 137?
    Because SIGKILL is not deliverable to the process at all — the kernel removes it without running any handler, so there is no point at which Python code could execute. No `except` clause, `finally` block or registered shutdown hook helps. A SIGTERM (143) is different: it can be handled, which is exactly why a service that wants a clean stop installs a handler for it and finishes in-flight work before exiting.
  • What guidance would you give a team about which exit codes their tools may use?
    Keep them in 1–125 and give them documented meanings, since 126 and 127 are shell conventions and anything at or above 128 collides with the signal encoding. Reserve 1 for unexpected failures and assign small distinct codes to conditions a supervisor should act on differently, such as retryable versus fatal. Also remember codes are truncated to eight bits, so anything above 255 wraps unpredictably.

An exit code is a note the process leaves on the way out; 137 is not a note at all but the coroner's tag, recording that something else ended the process before it could write anything.

saying these in an interview costs you the question

  • Reads 137 as an application-defined error code
  • Says Python returned 137 from a sys.exit call
  • Cannot map 137 back to 128 plus signal 9
  • Suggests catching the kill with a try block
  • Assumes 143 means the code failed rather than a stop request
  • Uses exit codes above 128 for application conditions

context