skip to content

A container exits with code 137 in production. How do you confirm it was a cgroup memory kill rather than something else, and how do you find the cause?

level: seniorimportance: must knowfreq 54%

answer

  1. 137 = 128+9 SIGKILL: OOM or external kill
  2. State.OOMKilled true + dmesg 'Memory cgroup out of memory'
  3. memory.events oom_kill counter; child killed = silent degradation
  4. memory.current includes page cache; check memory.stat anon vs file
  5. cap runtime heap below the container limit

basics

~20 s

137 means the process died from SIGKILL (128+9), which can be a cgroup OOM kill or an external kill such as a stop timeout. Confirm with docker inspect .State.OOMKilled, the kernel log line 'Memory cgroup out of memory', and the cgroup's memory.events oom_kill counter, then compare peak usage against memory.max.

solid answer

~60 s

Exit 137 = 128 + 9 = SIGKILL. Two common sources: the kernel's cgroup OOM killer when the container hits memory.max, or an external SIGKILL such as `docker stop` timing out or an operator running `docker kill`. Distinguish them: `docker inspect --format '{{.State.OOMKilled}}'` is true only for the memory kill, and the host log shows a line like `Memory cgroup out of memory: Killed process ...` naming the container's memcg. The cgroup's memory.events oom_kill counter increments too. Then find the cause. Compare the peak footprint against the limit; look at memory.stat to see whether the charge is anonymous memory (real heap growth or a leak) rather than reclaimable page cache. Note the kill targets the highest-scoring process *in that cgroup*: if it kills a worker child instead of PID 1, the container keeps running silently degraded. A frequent root cause is a runtime that never learned the limit - the JVM sizes its heap from the cgroup limit only with container support enabled, and native or off-heap growth still lives outside the heap cap. Fix by capping the runtime below the limit, correcting the leak, or raising the limit with headroom.

code

bash · 4 lines
bash
docker inspect --format '{{.State.ExitCode}} oom={{.State.OOMKilled}}' web
journalctl -k --since '10 min ago' | grep -i 'memory cgroup out of memory'
cat /sys/fs/cgroup/system.slice/docker-<id>.scope/memory.events   # oom / oom_kill
cat /sys/fs/cgroup/system.slice/docker-<id>.scope/memory.stat | head

go deeper

for a junior

Know that 137 means SIGKILL and that the usual cause is exceeding the memory limit; know where to check OOMKilled.

for a middle

Separate the possible SIGKILL senders, read memory.current versus memory.stat, and relate the runtime heap setting to the container limit.

for a senior

Run the full diagnosis: kernel log correlation, oom_kill counters, victim selection inside the cgroup, peak-versus-limit sizing, and the fix that makes the failure diagnosable.

for a principal

Set the policy: headroom standards, runtime caps below limits, alerting on oom_kill and pressure metrics, and how OOM risk trades against node density.

## Reading the exit code A container's exit code above 128 usually encodes a fatal signal: 128 + signal number. 137 is SIGKILL, 143 is SIGTERM. SIGKILL cannot be caught, so there are no application shutdown logs - which is exactly why people mistake an OOM kill for a crash. SIGKILL has several possible senders: 1. The kernel's cgroup OOM killer, when the container's memory cgroup cannot reclaim enough to satisfy an allocation under memory.max. 2. The container runtime, after `docker stop` waited out its grace period. 3. An operator or supervisor calling `docker kill`, or an out-of-band process manager. 4. The *host-level* OOM killer, when the machine as a whole is out of memory and picks a victim that happens to live in your container. ## Confirming a cgroup OOM `docker inspect --format '{{.State.OOMKilled}} {{.State.ExitCode}}' <c>` - OOMKilled true is decisive for cases 1 and 4. Kernel logs disambiguate further: `journalctl -k` or `dmesg -T` around the timestamp shows either `Memory cgroup out of memory: Killed process 1234 (java) total-vm:... anon-rss:...` with a memcg pointing at the container's cgroup, or a system-wide OOM report with no memcg. If the container is still alive, `cat /sys/fs/cgroup/memory.events` shows `oom` (allocation failures that triggered OOM handling) and `oom_kill` (actual kills) counters; `oom_kill` incrementing while the container keeps running is the smoking gun for the next point. ## The kill target is not always PID 1 The cgroup OOM killer chooses a victim inside the cgroup by badness score, adjusted by oom_score_adj. In a multi-process container - a supervisor with workers, a web server with children, a shell wrapper - it may kill a child. The container then keeps running with reduced or broken capacity, no restart, no alert from the orchestrator, and only degraded behavior visible. Any container with more than one process needs the oom_kill counter monitored, not just the exit code. ## Interpreting the memory numbers memory.current is the total charge, and it includes page cache from files the container read or wrote. That is why `docker stats` can show a container apparently pinned at its limit while it is perfectly healthy: cache is reclaimable, so the kernel evicts it under pressure instead of killing anything. Break the number down with memory.stat: `anon` is anonymous memory (heaps, stacks - not reclaimable, the usual killer), `file` is page cache, plus slab and sock entries for kernel-side charges. A steadily rising `anon` across restarts is a leak; a high `file` with stable `anon` is usually benign I/O. memory.peak (or memory.max_usage_in_bytes on v1) gives the high-water mark, which is what you must size against. ## Common root causes - **Runtime unaware of the limit.** /proc/meminfo is not namespaced, so a process that sizes itself from "total system memory" will pick a heap far larger than the cgroup allows. Modern JVMs read cgroup limits by default and honor MaxRAMPercentage; older ones and many other runtimes need explicit flags (Node's max-old-space-size, Go's GOMEMLIMIT). - **Off-heap growth.** A JVM heap capped at 400 MB in a 512 MB container still allocates metaspace, thread stacks, code cache, direct byte buffers and native library memory. The container limit must cover the whole process, not just the heap. - **A genuine leak or unbounded buffering** - caching everything, reading whole files into memory, an unbounded queue. - **Limit set too tight**, sized against average rather than peak, so a batch job or a traffic spike tips it over. ## The fix and the guardrail Size the limit from observed peak plus headroom, and cap the runtime's own allocation *below* the container limit so it fails with a diagnosable application-level error (an OutOfMemoryError with a heap dump) rather than an opaque SIGKILL. Add alerting on the oom_kill counter and on memory.current approaching memory.max, so you see pressure before the kill. If usage genuinely grows without bound, no limit is the answer - fix the leak.

  • docker stats shows the container sitting at its memory limit, but nothing has been killed. Is that a problem?
    Usually not by itself. The reported usage includes page cache from file I/O, which is reclaimable: when a new allocation needs room, the kernel evicts cache instead of invoking the OOM killer. Look at memory.stat and compare anon (non-reclaimable) with file (cache). A high file component with stable anon is normal for I/O-heavy workloads; steadily growing anon is the real warning sign.
  • The container's oom_kill counter is increasing, but the container never restarts. What is happening?
    The cgroup OOM killer picked a victim inside the container that is not PID 1 - a worker or child process - so the main process survives and the container never exits. Capacity silently degrades and no orchestrator notices. Monitor the memory.events oom_kill counter directly, and prefer a design where the container has one meaningful process so a kill is visible as a restart.
  • How would you distinguish an OOM kill from a stop-timeout kill when both produce exit 137?
    Check State.OOMKilled, which is true only for a memory kill, and correlate with the kernel log: a cgroup OOM writes an explicit 'Memory cgroup out of memory' record naming the process and the memcg. A stop-timeout kill instead correlates with a deployment or stop event and typically takes exactly the stop timeout, with no kernel OOM record at all.

saying these in an interview costs you the question

  • Treating exit 137 as automatically an application crash or bug in the image
  • Assuming 137 always means OOM, ignoring stop timeouts and docker kill
  • Reading docker stats memory usage as if it were all non-reclaimable
  • Believing an OOM kill always terminates the container
  • Sizing a container limit to the JVM heap size and forgetting metaspace, stacks and native memory

context