skip to content

Why does a containerized Python process exit with status 137 and print no traceback?

level: juniorimportance: should knowfreq 48%

answer

  1. The number is 128 plus something
  2. The kernel decided, not the interpreter
  3. One signal you cannot install a handler for
  4. No finally, no atexit, no flush
  5. Same number for a stop-timeout escalation

basics

~20 s

Status 137 is 128 + 9: the kernel delivered SIGKILL, most often when the process crossed its container memory limit. SIGKILL cannot be caught, so no except, finally or atexit code runs and nothing is printed.

solid answer

~50 s

A wait status of 137 means the process was terminated by signal 9, because the shell and container runtimes report a signalled death as 128 plus the signal number. In a memory-limited container the usual sender is the kernel: the cgroup memory controller could not charge another page, reclaim failed, and the kernel OOM-killed a process in that cgroup. SIGKILL is special in POSIX -- it cannot be caught, blocked or ignored, and `signal.signal` refuses to install a handler for it -- so CPython never gets control back. No exception is raised, no `finally` or `with` cleanup runs, no `atexit` callback fires, and buffered stdout is never flushed, which is why the last log lines are often missing too. Note that 137 only means *SIGKILL*, not necessarily out-of-memory: an orchestrator escalating from SIGTERM to SIGKILL after a stop timeout produces the same number.

code

python · 9 lines
python
import signal
import subprocess
import sys

child = subprocess.run(
    [sys.executable, "-c", "import os, signal; os.kill(os.getpid(), signal.SIGKILL)"]
)
print("Python sees:", child.returncode)      # -9
print("a shell reports:", 128 + signal.SIGKILL)  # 137

go deeper

for a junior

Recall that 137 is 128 + 9, that 9 is SIGKILL, and that a kill by signal explains why there is no traceback. Knowing to look at the container's exit reason rather than only the application log is most of the answer at this level.

for a middle

Explain the mechanics: CPython runs signal callbacks between bytecodes, SIGKILL never reaches that path, so no exception, no finally, no atexit, no flush. Be able to say why the last log line is missing and how 137 differs from 143.

for a senior

Show that you diagnose rather than guess: prove it was an OOM kill from the cgroup oom_kill counter or the kernel log rather than assuming, and design services that survive sudden death with idempotent, redeliverable work and no cleanup parked in finally.

for a principal

Own the platform consequence: because the failure is uncatchable, reliability has to come from headroom against peak rather than average, bounded concurrency, and redelivery semantics. Decide what memory-related evidence every service must emit before it dies so incidents are diagnosable at all.

## What the number actually is Unix does not have an exit code for "killed by a signal"; it has a wait status that says either *exited with code N* or *terminated by signal N*. Shells and container runtimes flatten those two cases into one number by reporting a signalled death as **128 + signal**. SIGKILL is 9 on Linux, so you see 137. Nothing in your program chose that number -- there is no `sys.exit(137)` anywhere. From inside Python the same event looks different: - `subprocess.run` reports a negative return code, -9, and - if you decode a raw wait status yourself, `os.WIFSIGNALED` is true and `os.WTERMSIG` is 9. ## Who did the killing When a container runs under a memory limit, the kernel's **cgroup memory controller** charges every page the process faults in against that limit. When a charge cannot be satisfied and reclaim (dropping page cache, swapping if allowed) does not free enough, the kernel invokes the **OOM killer** scoped to that cgroup and sends SIGKILL to a process inside it. Selection, timing and delivery are all kernel-side; the interpreter is a bystander that is simply stopped between two bytecode instructions. If the process killed is the container's PID 1, the whole container ends. ## Why Python prints nothing CPython's signal handling is deliberately **cooperative**. The C-level handler does almost nothing except set a flag; the actual Python callback runs later, in the eval loop, between bytecodes. SIGKILL never reaches that machinery at all, because POSIX forbids catching, blocking or ignoring it -- `signal.signal(signal.SIGKILL, handler)` raises `OSError`. The consequences are total: - no `MemoryError` is raised, - no `except` block runs, - no `finally` runs, - no context manager's `__exit__` runs, - no `atexit` callback fires, - no `__del__` runs, and - the logging module never gets its shutdown flush. `faulthandler` does not rescue you either: it exists to dump a traceback on fatal faults such as SIGSEGV or SIGABRT, and SIGKILL is not one of the signals it can arm. ## The missing last log line When stdout is a pipe rather than a terminal it is **block-buffered**, so several kilobytes of already-written log lines can be sitting in the buffer when the kill lands. The logs then point at an earlier place in the code than where the process actually died, which sends people hunting the wrong function. Running with `PYTHONUNBUFFERED=1` or `python -u`, or using a logging handler that flushes each record, makes the trail honest -- at some throughput cost. ## 137 is not a synonym for out-of-memory Any SIGKILL from any sender produces it. The most common non-OOM source is a **shutdown sequence**: the platform sends SIGTERM, waits a grace period, and escalates to SIGKILL when the process has not exited. A Python service that ignores SIGTERM, or that is blocked in a long call that never checks for the signal, will show 137 on every single deploy and nothing is wrong with its memory at all. Tell the two apart with evidence rather than assumption: - the cgroup's `memory.events` file carries an `oom_kill` counter that increments only for real OOM kills, - the kernel ring buffer records a "Killed process" line, and - the container runtime records an OOM-killed exit reason. A clean SIGTERM death, by contrast, shows 143 (128 + 15) and *is* catchable -- Python can install a SIGTERM handler and shut down gracefully. ## Reading the evidence from each vantage point The same death looks different depending on who is watching. - A shell or an orchestrator's status field shows 137. - A Python parent that spawned the process sees a return code of -9 from `subprocess.run`, or `os.WIFSIGNALED` true with `os.WTERMSIG` returning 9 if it decodes a raw wait status itself. - The kernel log shows the cgroup that ran out, the chosen victim's process id and command, and its **resident size at the moment of death** -- which is often the single most useful number in the whole incident, because it tells you what the process was actually holding rather than what it was averaging. Knowing all three views matters because in most deployments you can only reach one of them, and a supervisor that logs only "child exited with -9" is reporting the same event as a dashboard that says OOMKilled. ## What follows for design Because there is no in-process hook, everything useful must be **pre-emptive**. - Log or export the *peak* memory rather than the average, since the kill is driven by an instantaneous high-water mark that a 30-second metrics scrape will usually miss. - Cap concurrency so peak footprint is bounded by construction. - And treat sudden death as a design constraint rather than an exceptional case: work items must be **idempotent and redeliverable**, because any cleanup you were planning to do in a `finally` block is not going to happen.

  • Does status 137 always mean the process ran out of memory?
    No. It means SIGKILL, whoever sent it. The common non-OOM cause is a stop sequence: the platform sends SIGTERM, waits out a grace period, and escalates to SIGKILL because the process did not exit. Confirm a real out-of-memory event from the cgroup's oom_kill counter in memory.events, the kernel log line naming the killed process, or the runtime's OOM-killed exit reason -- not from the number alone.
  • Why is the last log line before the kill often missing?
    When stdout is a pipe rather than a terminal, it is block-buffered, so up to several kilobytes of emitted lines sit unflushed. SIGKILL gives no opportunity to flush them, so the visible log ends earlier than the code actually reached. Setting PYTHONUNBUFFERED=1 or running python -u, or using a logging handler that flushes per record, makes the tail trustworthy at some throughput cost.
  • How does status 143 differ from 137, and why does the difference matter operationally?
    143 is 128 + 15, a SIGTERM death. SIGTERM can be handled: Python can install a handler with signal.signal, stop accepting work, finish in-flight items and exit cleanly, so 143 usually indicates an orderly shutdown. 137 means the process was cut off mid-instruction with no cleanup at all. Seeing 137 at every deploy usually means the SIGTERM path is missing or too slow.

SIGTERM is a request to leave the building that you can answer; SIGKILL is the power being cut at the breaker -- there is no chance to save your work or leave a note.

saying these in an interview costs you the question

  • Says the process raised MemoryError and the traceback was swallowed
  • Claims try/except MemoryError would have caught this kill
  • Thinks 137 is an exit code the application chose
  • Believes finally blocks or atexit handlers still run on SIGKILL
  • Assumes 137 always means out-of-memory, never a stop-timeout SIGKILL
  • Suggests installing a SIGKILL handler to log the cause

context