Exit status 137 and no traceback: how do you tell a kernel OOM kill of a Python process from a MemoryError?
answer
- One of the two leaves a traceback
- Signal nine cannot be caught
- 128 plus the signal number
- Overcommit means the allocation usually succeeds
- The evidence lives in the kernel log
basics
~10 sA MemoryError is an ordinary Python exception, so the interpreter is alive to print a traceback and run handlers. A kernel OOM kill is SIGKILL from outside: no traceback, no cleanup, exit status 137.
solid answer
~50 sThe two look nothing alike once you know where to look. `MemoryError` means **one allocation request failed** while the interpreter was still running — you get a normal traceback naming the line, `finally` blocks and `except` handlers run, and the process can in principle continue. A kernel out-of-memory kill is `SIGKILL` delivered from outside the process: it cannot be caught, blocked or handled, so there is **no traceback, no `finally`, no `atexit` hook**, the log simply stops mid-sentence, and the shell reports status **137** (128 + 9). Confirmation lives outside the application: the kernel log records which process it chose and its resident size at the time. On Linux the OOM kill is the more common of the two, because memory overcommit means `malloc` usually succeeds and the shortfall is only discovered when pages are actually touched — so the process is killed rather than told no.
code
python · 6 linestry:
buf = bytearray(10 ** 18) # far larger than any address space
except MemoryError:
print("the allocation was refused and Python raised MemoryError")
else:
print(len(buf))go deeper
Be ready to recognise that a missing traceback means the process was killed rather than raising an exception, and that MemoryError is a normal exception you can see in a traceback.
Explain the mechanics: SIGKILL is uncatchable so no handler or finally runs, status 137 is 128 plus signal 9, and overcommit is why an allocation usually succeeds instead of raising.
Show how you actually work the case — confirming the kill from the exit status and the kernel log, telling a spike apart from steady retention, and instrumenting resident size ahead of time because the process gets no chance to report.
Own the operating posture: choose the memory watermark at which a service takes itself out of rotation, decide whether workloads are sized to fail fast or to be restarted, and make exit-status capture a platform default rather than a per-team habit.
## Two different failures that both mean "out of memory" **`MemoryError`** is raised by the interpreter when a specific allocation request cannot be satisfied — the allocator asked the operating system for memory and was refused. It is a normal exception: it has a traceback pointing at the line that asked for too much, it unwinds the stack, `finally` blocks run, and code above may catch it. **A kernel OOM kill** is not a Python event at all. When the system (or a memory cgroup) is out of memory, the kernel picks a victim process and sends it `SIGKILL`. Signal 9 cannot be caught, blocked or ignored — not by a `signal` handler, not by `atexit`, not by a `try`/`finally`. The process stops executing instantly. There is no traceback because there is no interpreter left to print one. ## The evidence that distinguishes them | Evidence | `MemoryError` | OOM kill | |---|---|---| | Traceback in your log | yes, naming the line | none; the log stops mid-stream | | `finally` / `atexit` / cleanup | runs | does not run | | Exit status seen by a shell or supervisor | normal (often 1) | **137** = 128 + 9 | | Status from a `subprocess` handle | positive | **-9** (negative signal number) | | Record outside the process | none | kernel log entry naming the victim | The negative-versus-137 detail trips people up: the child-process API reports a signal death as the **negative** signal number, while shells and orchestrators report **128 + signal**. Both describe the same `SIGKILL`. ## Why the kill is the common case on Linux Linux overcommits: a large allocation typically *succeeds* on paper because the kernel hands back address space it has not actually reserved. The shortfall only becomes real when the process touches the pages. So the classic "Python raises `MemoryError` when you run out of RAM" mental model is mostly wrong on a modern Linux box — you far more often get killed than told no. `MemoryError` in the wild usually comes from something more specific: an absurd size request that fails immediately, a per-process address-space limit, or a 32-bit address space. ## Working a real case Take a nightly translation-memory updater that walks a dependency graph of 17 services, pulling each one's segment corpus and merging it. It runs fine for the first few nights, then the container disappears at around hour two with no traceback and status 137, and the last log line is mid-batch. There is no exception to catch, so the investigation is about *shape*, not stack: 1. **Confirm the kill.** Status 137 plus a kernel log entry naming the interpreter as the victim settles it — without both, you are guessing. A supervisor that records only "exited" is not evidence. 2. **Establish whether growth is monotonic.** If resident size climbs steadily across batches and never falls, this is retention, not a spike — something is holding references for the process's lifetime. 3. **Look for lifetime mismatches.** The classic one here is a cache parked in a **mutable default argument**: `def merge(segments, cache={})` evaluates that dict once at `def` time, so every call shares it, it is reachable for as long as the function object exists, and nothing ever evicts. Every night it holds the whole corpus. 4. **Distinguish spike from leak.** A single oversized batch that dies at the same point every run is a spike — stream it or chunk it. Steady climb that dies later each time you give it more headroom is retention — more headroom only buys minutes. ## What to install before the next kill Since the process gets no chance to say goodbye, the instrumentation has to be *ahead* of the failure: - **Log resident size periodically** with a batch counter, so the growth curve is already in your logs when the kill happens. - **Record the exit status** in whatever supervises the process, so 137 is visible rather than inferred. - **Handle `SIGTERM`.** Many orchestrators send `SIGTERM` first and `SIGKILL` only after a grace period; a `SIGTERM` handler that flushes state and logs the current batch turns some hard kills into clean shutdowns. It cannot help against an in-kernel OOM kill, which is `SIGKILL` with no warning. - **Fail your own health check** at a memory watermark you choose, so the process is restarted on your terms instead of the kernel's. ## Why catching `MemoryError` rarely helps Even the catchable case is a poor recovery point. The handler itself runs in a memory-starved process, and building an error message, formatting a traceback or logging all allocate — so recovery frequently fails again. The defensible pattern is narrow: catch it where you can immediately drop the one huge object you just tried to build, and otherwise let the process die and be restarted cleanly.
- Why can a Python process not log anything on its way out of a kernel OOM kill?Because the kill is `SIGKILL`, which the kernel delivers without giving the process a chance to run any code. Signal 9 cannot be caught, blocked or ignored, so no handler installed through the `signal` module fires, no `finally` block runs, and no `atexit` hook runs. Anything you want to know afterwards must have been logged before the kill, or recorded outside the process by the kernel or the supervisor.
- Is catching MemoryError and continuing a reasonable strategy in a long-running service?Rarely. The handler runs inside a process that is already out of memory, and formatting a message, building a traceback or logging all allocate, so recovery often fails a second time. The narrow case that works is catching it right where a single huge object was being constructed, dropping that object and refusing the request. Otherwise the honest design is to let the process exit and be restarted by a supervisor.
- How would you tell a memory spike apart from steady retention when all you have is repeated kills?Log resident size against a work counter and look at the curve. A spike shows a flat baseline with a jump at the same point in the workload every run, and dies at the same place no matter how much headroom you add. Retention shows a monotonic climb that never drops between units of work, and extra headroom only moves the kill later. The two call for different fixes: streaming versus finding what holds the references.
A MemoryError is the bank declining your card — you are still standing there and can react. An OOM kill is the building being demolished around you: no notice, no last words, and the only record is the city's paperwork.
saying these in an interview costs you the question
- Expecting a traceback from a process killed by SIGKILL
- Believing Python always raises MemoryError when RAM runs out
- Installing a SIGKILL handler to log the crash
- Reading exit status 137 as an application error code
- Assuming MemoryError means the whole machine is out of memory
- Treating catch-and-continue on MemoryError as a general recovery