skip to content

questions

3

When does CPython raise a catchable MemoryError instead of the container being OOM-killed?

level: middleimportance: must knowfreq 41%

answer

  1. One is an exception, one is a signal
  2. Address space is cheap, pages are charged later
  3. The limit bites when you touch, not when you ask
  4. The handler itself needs memory to run
  5. Move the ceiling inside the process to regain control

basics

~20 s

MemoryError is raised only when an allocation request fails immediately, so the interpreter regains control. A cgroup limit is charged when pages are touched, not when they are requested, so the kernel kills the process instead -- there is nothing to catch.

solid answer

~50 s

`MemoryError` is an ordinary Python exception raised when the C allocator behind an object refuses a request -- it returns NULL and CPython converts that into an exception at the allocation site. Under default Linux overcommit that almost never happens for realistic sizes: the kernel hands out address space cheaply and charges physical pages only when they are first touched. A cgroup memory limit is enforced at that charge point, so the process dies *while writing to memory it already believes it owns*, by an uncatchable SIGKILL. You do see a real `MemoryError` when a request is refused up front: an absurd single allocation, an address-space ceiling set for the process, or a host configured to refuse overcommit. Even then, recovering is fragile because the handler itself needs memory. The reliable way to get a catchable failure is to bound your own buffers and concurrency below the platform limit.

code

python · 4 lines
python
try:
    buf = bytearray(1 << 62)
except MemoryError:
    print("MemoryError: the allocator refused the request up front")

go deeper

for a junior

Recall that MemoryError is a normal exception with a traceback while an OOM kill leaves no traceback at all, and that the two need completely different investigations. Knowing which one you are looking at is the first useful step.

for a middle

Explain the mechanism: overcommit makes the allocation succeed, the cgroup charges pages on first touch, so the failure lands in the kernel rather than at a Python allocation site. Name the cases that still produce a real MemoryError.

for a senior

Demonstrate that you design the ceiling instead of hoping for an exception: bounded buffers, capped in-flight work, size limits and streaming, so pressure surfaces as an error you wrote with headroom to shed load. Explain why catching MemoryError is unreliable recovery.

for a principal

Own the policy question: how much headroom over peak each workload is given, whether workloads that cannot bound their footprint get isolated, and what the platform contract is when a limit is hit -- restart-and-redeliver versus in-process degradation.

## Two different failures that people treat as one `MemoryError` is raised *inside* your process: some allocation returned NULL, CPython noticed, and raised an exception at that exact line. The stack unwinds, `except` blocks run, `finally` blocks run, and the process is still alive. A **container OOM kill** happens *to* your process: the kernel decides, sends SIGKILL, and the interpreter is stopped between two bytecodes with no exception, no cleanup and no traceback. Which one you get is decided by *where the memory pressure becomes visible*, and that is a kernel question, not a Python question. ## Why the kill usually wins On Linux with the default **heuristic overcommit** policy, a large allocation request is mostly bookkeeping: the kernel reserves address space and returns success without committing physical memory. Pages are charged only on **first touch**, when the fault handler backs the page. A cgroup memory limit is enforced at exactly that charge point. So the sequence in a container is: 1. your allocation *succeeds*, 2. you begin filling the object, 3. and somewhere in the middle of an ordinary write the charge cannot be met, 4. reclaim fails, 5. and the cgroup OOM killer fires. There was never a NULL for CPython to turn into an exception, which is why wrapping the code in `try/except MemoryError` changes nothing at all. Worse, the kernel picks a victim inside the cgroup by score, so in a multi-process container the process that dies may not be the one that allocated -- a small supervisor can be killed for a large worker's growth. ## When you genuinely do get MemoryError Several situations still produce the catchable form. - A single request the kernel refuses on sight -- ask for a few exabytes and the heuristic says no immediately. - A per-process address-space ceiling configured for the process, which makes the allocator itself fail rather than the kernel act later. - A host with overcommit disabled entirely, where every mapping must be backed by memory plus swap at request time and `malloc` really can return NULL. - And a handful of cases CPython checks itself, such as a computed length that cannot possibly be allocated. There is also a real distinction between failing to allocate an object's payload and failing to allocate the small internal structure: both surface as `MemoryError`, but only the first is proportional to your data. ## Catching MemoryError is not a recovery strategy Even when the exception does arrive, the handler runs in a process that has just proved it cannot get memory. Formatting a log message, building a traceback, or constructing the exception's own message may need allocations that fail again. Other threads are hitting the same wall concurrently, so state can be inconsistent in ways your handler does not expect. The only handlers that reliably work are ones that **free something already reserved** -- dropping a cache you deliberately kept droppable -- and then either retry or exit deliberately. Treating `except MemoryError: continue` as backpressure produces a process that limps along corrupting throughput instead of restarting clean. ## Making failure catchable on purpose If you want a degradation path rather than a sudden kill, you have to move the ceiling *inside* the process, below the platform's limit, so your own code notices first. Concretely: - bound how much you buffer per work item, - cap the number of items in flight, - reject or spool payloads above a size you chose, - and stream rather than reading whole objects into memory. Those checks raise exceptions you wrote, at points you control, with headroom left to handle them. A process-level address-space ceiling can play the same role at a coarser grain, turning a would-be kill into a `MemoryError` at the allocation site -- at the cost that the exception can arrive anywhere, including inside library code that was never written to survive it. ## The interview point The distinction is diagnostic, not academic. - A traceback ending in `MemoryError` tells you a single allocation was too big for what remained and the process kept running -- look at that allocation. - An exit status of 137 with no traceback tells you the kernel enforced a limit against your resident footprint -- look at the peak, at concurrency, and at how much memory the process is holding that it never returns.

  • Why can catching MemoryError and retrying make things worse?
    The handler runs in a process that has just failed to get memory, and building the message, the traceback or any retry state may need allocations that fail again. Meanwhile other threads are hitting the same wall, so invariants can already be broken. A handler is only trustworthy when it releases memory you deliberately reserved as droppable, such as a cache, and then either retries once or exits so the supervisor can restart cleanly.
  • In a container running several processes, why might the wrong one die?
    The cgroup OOM killer chooses a victim by score across processes in that cgroup, not by which one requested memory. A modest supervisor or sidecar can be selected while the greedy worker keeps running, which makes the crash look unrelated to the growth. Running one process per container, or checking the kernel log line naming the killed process, avoids chasing the wrong component.
  • How would you give a memory-hungry worker a graceful degradation path instead of a kill?
    Put the ceiling inside the process, below the platform limit: cap payload sizes, bound how many items are in flight, stream instead of loading whole objects, and reject work when your own accounting says the budget is gone. Those raise exceptions you wrote, at points you control, with headroom to shed load, log the cause and exit deliberately rather than being cut off mid-instruction.

Asking for memory under overcommit is like being promised a huge warehouse: the paperwork is approved instantly, and you only discover the space does not exist when you try to walk into the far end of it.

saying these in an interview costs you the question

  • Says a container OOM kill can be caught with except MemoryError
  • Thinks a successful allocation means the memory is really available
  • Claims MemoryError means the machine is out of RAM
  • Treats except MemoryError plus retry as a backpressure mechanism
  • Assumes the process that allocated is always the one killed
  • Believes raising the container limit is the only possible fix

context

open as a page

Why does a containerized Python process exit with status 137 and print no traceback?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Status 137 is 128 + 9: the kernel delivered SIGKILL, most often when the process crossed its container memory limit. SIGKILL cannot be caught, so no except, finally or atexit code runs and nothing is printed.

open as a page

A webhook receiver's Python worker keeps dying at exit 137 after traffic bursts, though its average memory looks fine. How do you diagnose it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Limits are enforced against instantaneous resident size, so diagnose the peak, not the average: confirm a real OOM kill from the cgroup counters, sample resident size frequently enough to see the burst, then attribute the peak to per-request bytes multiplied by in-flight concurrency.

open as a page