skip to content

A webhook receiver's Python worker keeps dying at exit 137 after traffic bursts, though its average memory looks fine. How do you diagnose it?

level: seniorimportance: should knowfreq 44%

answer

  1. Averages hide the thing that kills you
  2. First prove it was an OOM event at all
  3. Peak equals per-item bytes times concurrency
  4. A plateau is not the same shape as a leak
  5. One slow dependency multiplies live payloads

basics

~20 s

Limits are enforced against instantaneous resident size, so diagnose the peak, not the average: confirm a real OOM kill from the cgroup counters, sample resident size frequently enough to see the burst, then attribute the peak to per-request bytes multiplied by in-flight concurrency.

solid answer

~50 s

Averages are the wrong instrument: the kernel charges resident pages continuously and kills on an instantaneous high-water mark, so a two-second spike between thirty-second metric scrapes is invisible. First confirm it really was an out-of-memory kill rather than a stop-timeout SIGKILL, using the cgroup's oom_kill counter and the kernel log. Then read the cgroup's peak counter, or sample resident size from a watchdog thread at sub-second resolution, to see the true high-water mark. Then attribute it: peak is roughly bytes held per in-flight request times the number in flight, and in a fan-out receiver every request's payload stays alive until its slowest dependency answers, so one slow dependency multiplies resident bytes. Finally, distinguish a leak, which grows across bursts, from a plateau, which is retained-but-unreturned memory. Fixes are bounding concurrency, streaming instead of buffering, and capping payload size.

code

console · 2 lines
console
cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/memory.current /sys/fs/cgroup/memory.peak
grep oom_kill /sys/fs/cgroup/memory.events

go deeper

for a junior

Recall that the limit is checked against a momentary peak rather than an average, so a spike can kill a process whose dashboard looks calm. Knowing to ask for peak memory rather than mean memory is the useful instinct here.

for a middle

Explain what resident size actually contains beyond your objects, and why deleting a large structure often does not shrink it. Be able to describe sampling the high-water mark rather than trusting a coarse scrape interval.

for a senior

Show a diagnosis order: prove the OOM event, capture the peak, attribute it to per-item bytes times in-flight concurrency, then separate a leak from a plateau. Then fix by bounding concurrency, streaming, capping payload size and tightening downstream timeouts.

for a principal

Own the sizing policy: headroom is set against peak under dependency degradation, not steady state, and worker counts multiply a per-process baseline against a shared ceiling. Decide what every service must emit -- peak memory and OOM counters -- so these incidents are diagnosable without a reproduction.

## Start by proving what killed it Exit 137 only says SIGKILL. Before optimising anything, check: - the cgroup's `memory.events` for an incrementing `oom_kill` count, - the kernel ring buffer for the line naming the killed process, - and the runtime's recorded exit reason. If none of them show an OOM event, you are chasing the wrong failure -- a shutdown escalation from SIGTERM produces the identical number. ## Measure the peak, not the average A memory limit is a ceiling on **instantaneous resident set size**, and the kill is instantaneous too. Dashboards built from a scrape every fifteen or thirty seconds smooth exactly the spike that killed you: a burst that lasts two seconds is very likely never sampled at all. Two better instruments exist. - The cgroup exposes a **peak counter** that records the high-water mark since it was reset, which survives the death of the process. - And inside the process, a small **watchdog thread** that samples resident size several times a second and records a running maximum -- emitting it periodically -- gives you the shape of the spike rather than a single number. The crucial property is that both are max-based, not sample-based. ## Resident size is not the Python heap The number the kernel enforces against includes far more than your objects: - the interpreter and its extension modules, - thread stacks, - the allocator's own pools and arenas, - and, in cgroup v2, the page cache your process generated by reading and writing files. Cache is reclaimable, so it usually causes reclaim rather than a kill, but it does inflate the counter and confuses naive comparisons against a heap measurement. So when the object graph you can account for is a few hundred megabytes and resident size is well above that, the gap is normal, not necessarily a bug. ## Freed memory is not necessarily returned Deleting a large structure drops it from Python's view immediately, but the pages may stay with the process. Small objects are freed into the allocator's pools and only handed back to the OS when an entire arena becomes empty, which fragmentation can prevent indefinitely; large allocations that were mapped separately usually do go back. The practical consequence is a **ratchet**: after each burst, resident size settles at a plateau higher than before, so the second and third bursts start closer to the ceiling than the first. That is why a service can survive its first spike and die on an identical one an hour later. Distinguishing this from a real leak is a matter of shape -- a **leak** keeps climbing across quiet periods too, whereas **retention** rises to a plateau and then stops. ## Attribute the peak to concurrency Peak footprint is approximately **bytes retained per in-flight unit of work times the number of units in flight**. In a receiver that accepts a webhook and fans out to a large dependency graph -- say seventeen downstream services per event -- the request payload, the parsed object graph and every partial response are all reachable until the slowest of those calls returns. The ordering assumption is the trap: code written as though responses come back in the order they were issued, releasing buffers as it goes, actually holds *all* of them until the last one lands. One dependency going from 50 ms to 5 s therefore multiplies the number of concurrently live payloads by a hundred without any change in request rate, which is exactly the profile of a service that is fine for weeks and then dies during someone else's incident. ## Fixes, in order of leverage 1. Bound in-flight work explicitly with a semaphore or a small fixed pool, so peak is a design parameter rather than an emergent property of dependency latency. 2. Stream instead of buffering: iterate over the body, spool oversized payloads to a temporary file, and never hold a whole response when a chunk will do. 3. Reject payloads above a size you chose, at the edge, before allocating. 4. Set aggressive timeouts on downstream calls so buffers cannot be held hostage by a slow dependency. 5. Right-size worker counts: each worker process pays its own resident baseline, and multiplying processes multiplies that baseline against a shared limit. ## A 3.14 note on workers The default `multiprocessing` start method on Unix other than macOS is now `forkserver` rather than `fork` (macOS and Windows use `spawn`, and `fork` must be requested explicitly). - With `fork`, children shared the parent's pages copy-on-write, so a warm parent made children look cheap until reference-count writes gradually unshared those pages. - With `forkserver` each child starts from a small clean server process and imports what it needs, so per-child memory is more predictable but no longer subsidised by the parent. If a service was tuned for worker counts under the old default, its resident footprint can change on upgrade.

  • How do you tell a leak from memory that is simply never returned to the OS?
    Compare shapes over several bursts and quiet periods. A leak keeps climbing, including while idle, because live objects accumulate. Retention rises during a burst and then flattens at a plateau that does not grow further, because the allocator is holding pages it could reuse but cannot hand back while any object still sits in the arena. Repeating the same burst and seeing the plateau stop rising is the cheapest discriminator.
  • Why can resident size be far larger than the total size of the objects you can account for?
    The enforced number covers the whole process: the interpreter and extension modules, thread stacks, allocator pools and arenas including fragmentation, and under cgroup v2 the page cache generated by file I/O. Only some of that corresponds to your object graph, and cache is reclaimable rather than fatal. A large gap between accounted objects and resident size is normal and not by itself evidence of a bug.
  • The dependency fan-out is the cause. What is the smallest change that stops the kills?
    Bound in-flight work, typically with a semaphore or a fixed-size pool sized so that concurrency times per-request retained bytes fits under the limit with headroom. That converts peak footprint from an emergent property of downstream latency into a number you chose. Tightening downstream timeouts helps too, since it caps how long each payload can stay reachable, but bounding concurrency is what makes the ceiling deterministic.
  • Why does the service survive the first burst and die on an identical one later?
    Resident size ratchets. After a burst, Python frees objects but the allocator often keeps the pages, since an arena returns to the OS only when it is entirely empty and fragmentation frequently prevents that. Each burst therefore starts from a higher floor, and eventually floor plus spike crosses the limit. The traffic did not change; the starting point did.

A memory limit is a doorway height, not an average: it does not matter that your load is usually low if, once an hour, everything is stacked up at once waiting for one slow porter.

saying these in an interview costs you the question

  • Reasons from average memory graphs and ignores the peak
  • Assumes exit 137 is an OOM kill without checking the counter
  • Treats resident size as equal to the size of live Python objects
  • Calls every plateau a memory leak
  • Proposes raising the limit before attributing the peak
  • Expects gc.collect to return the pages to the operating system

context