skip to content

A container stopped with exit code 137. How do you determine whether the kernel's out-of-memory killer did it or something else sent SIGKILL, and what evidence do you collect?

level: seniorimportance: must knowfreq 58%

answer

  1. 137 = SIGKILL — sender unknown from the code alone
  2. `.State.OOMKilled` = the cgroup-OOM flag
  3. False + 137 → stop timeout, docker kill, or host OOM
  4. `dmesg -T | grep -i oom` names the cgroup and victim
  5. Child-process kill can set the flag without stopping the container

basics

~20 s

137 only says "killed by SIGKILL". Check docker inspect --format '{{.State.OOMKilled}} {{.State.ExitCode}}' <container>: true means the kernel out-of-memory killer fired. If it is false, look for a stop that timed out into SIGKILL, a docker kill, or a host-level kill — using timestamps, daemon logs, and dmesg.

solid answer

~1 min

Exit 137 = 128 + 9 = SIGKILL, and SIGKILL has several possible senders, so the code alone is not a diagnosis. My sequence: 1. `docker inspect --format '{{.State.OOMKilled}} {{.State.ExitCode}} {{.State.StartedAt}} {{.State.FinishedAt}}' <ctr>`. `OOMKilled=true` is the direct answer: the container hit its memory limit and the kernel killed the process in its cgroup. 2. If false, ask **who else could have sent SIGKILL**: a `docker stop` whose grace period expired (the container would have received SIGTERM ~10s earlier — correlate `FinishedAt` with a deploy or a restart), an explicit `docker kill`, or an operator/supervisor on the host. 3. Corroborate outside Docker: `dmesg -T | grep -i -E 'oom|killed process'` and the journal show kernel OOM events including the cgroup and RSS at kill time — this also catches a **host-level** OOM (the machine ran out of memory) where `OOMKilled` may not be set on the container. 4. Read `docker logs --tail 100` for whether shutdown began (a SIGTERM handler logging "shutting down" before the kill points at a stop timeout, not memory) and check memory metrics leading up to `FinishedAt`. Watch the trap: `OOMKilled` reflects a kill inside the container's cgroup, so a killed child process can set it while the container survives — and a host OOM can kill your process with it false.

code

bash · 5 lines
bash
docker inspect --format \
  '{{.State.ExitCode}} oom={{.State.OOMKilled}} start={{.State.StartedAt}} end={{.State.FinishedAt}} err={{.State.Error}}' \
  web

dmesg -T | grep -iE 'out of memory|killed process|oom-kill' | tail -20

go deeper

for a junior

Know that 137 means SIGKILL and that docker inspect has an OOMKilled field which tells you if memory was the reason.

for a middle

List the alternative senders — stop-timeout escalation, docker kill, host OOM — and use timestamps plus logs to tell them apart.

for a senior

Run the full evidence chain: inspect fields, kernel OOM lines with the cgroup and RSS, log correlation, memory trend up to FinishedAt; state the root cause with the evidence that supports it.

for a principal

Turn the diagnosis into policy: distinguish container-limit OOM from node oversubscription in alerting, treat repeated stop-timeout 137s as a graceful-shutdown defect class, and require kernel evidence before any limit is raised.

## What 137 does and does not tell you 137 is `128 + 9`: the container's main process was terminated by **SIGKILL**. SIGKILL cannot be caught, blocked, or handled, so the process ran no cleanup — no flushing buffers, no draining connections, no releasing locks. Everything you need to know beyond that is *who sent it*, and the exit code is silent on that. Candidates: - The **kernel OOM killer**, because the container exceeded its memory cgroup limit. - The **kernel OOM killer at host scope**, because the machine as a whole ran out of memory and your process was the fattest target. - **Docker itself**, escalating a `docker stop` whose grace period expired. - **An explicit kill** — `docker kill`, or a `kill -9` from a human or a supervisor on the host. Each of those has a different fix, so misattributing 137 to "OOM" by reflex wastes an incident. ## Step 1: the container's own record Docker records an OOM flag alongside the exit code: ``` docker inspect --format '{{.State.ExitCode}} oom={{.State.OOMKilled}} start={{.State.StartedAt}} end={{.State.FinishedAt}}' web ``` `OOMKilled=true` with 137 is as close to a confirmed diagnosis as you get from Docker alone: the kernel signalled an out-of-memory condition for that container's cgroup and killed a process in it. `.State.Error` occasionally carries additional text. The `StartedAt`/`FinishedAt` pair matters as much as the flag. A container that died 90 seconds after start, repeatedly, at the same point in its warm-up, looks very different from one that died 6 days in — the first suggests a startup allocation that exceeds the limit, the second a leak or a traffic spike. ## Step 2: rule the alternatives in or out With `OOMKilled=false`, work through the other senders: **Stop-timeout escalation.** `docker stop` sends SIGTERM (or `STOPSIGNAL`), waits — 10 seconds by default — and then sends SIGKILL. A container that is slow to shut down therefore ends at 137 even though nobody intended a force kill. Evidence: the container logs show shutdown starting ("draining", "closing listeners") roughly the grace period before `FinishedAt`, and the kill coincides with a deploy, a `docker compose down`, or a host reboot. The fix is application-side shutdown speed or a longer `-t`/`stop_grace_period`, not memory. **Explicit kill.** `docker kill` sends SIGKILL immediately, with no preceding SIGTERM in the logs. Check the daemon journal (`journalctl -u docker`) and your own automation/cron for anything that force-kills containers. **Host-level OOM.** When the *machine* runs out of memory, the kernel picks a victim across all processes. This kills your container's process with SIGKILL — 137 — but because the kill did not come from the container's own memory cgroup limit, the container's `OOMKilled` flag may be false. This is the case most often misdiagnosed, and it is why you check the kernel, not just Docker. ## Step 3: kernel evidence The authoritative record of any OOM kill lives in the kernel ring buffer and the journal: ``` dmesg -T | grep -iE 'out of memory|killed process|oom-kill' journalctl -k --since '2026-08-12 09:00' | grep -i oom ``` A cgroup OOM line names the cgroup path (which contains the container ID), the victim's PID and comm, and its RSS at the moment of death. That single line resolves both "was it OOM?" and "was it *my* container's limit or the host's?" — and it also identifies cases where a *child* process inside the container was the victim. ## Step 4: the child-process subtlety The OOM killer chooses a process, not a container. If it kills a worker child rather than PID 1, the container may keep running in a degraded state while `OOMKilled` gets set — so the flag being true does not always come with a 137, and a healthy-looking container can be quietly losing workers. Conversely, on modern cgroup v2 setups configured to kill the whole group, everything dies together. When symptoms are "the app is weird but still up", check the flag and the kernel log even without an exit at all. ## Step 5: turning evidence into action - **Confirmed cgroup OOM**: the container's memory ceiling and the workload disagree. Either the workload genuinely needs more headroom, or something inside is growing without bound (heap, cache, connection pool, native buffers). Historical memory metrics up to `FinishedAt` separate "steady sawtooth then spike" (burst) from "monotonic climb" (leak). - **Host OOM**: the node is oversubscribed — the fix lives at scheduling/placement level, not in this container. - **Stop timeout**: shutdown is slower than the grace period; measure it and either speed it up or widen the window. - **External kill**: find the caller; unexplained SIGKILLs are an operational problem in their own right. Always write the conclusion down with its evidence (`OOMKilled` value, the dmesg line, the timestamps). "It exited 137" is not a root cause, and a 137 that gets attributed to memory without kernel evidence tends to come back a week later as the same incident.

  • `OOMKilled` is false but `dmesg` shows an out-of-memory kill naming your process. How is that possible?
    The container flag reflects an OOM event attributed to that container's own memory cgroup. When the *host* runs out of memory, the kernel picks a victim globally and kills it with SIGKILL; the container still reports 137 but was never over its own limit, so the flag stays false. The conclusion is that the node is oversubscribed rather than that this workload is misconfigured, and the remedy is placement or node sizing rather than raising this container's ceiling.
  • The logs show your app printing "shutting down" about ten seconds before it exits 137. What does that tell you?
    It received SIGTERM, began a graceful shutdown, and did not finish within `docker stop`'s grace period — so Docker escalated to SIGKILL. This is a shutdown-duration problem, not a memory problem: `OOMKilled` will be false and `dmesg` will be clean. Either make shutdown faster (bound the drain, close listeners first) or extend the timeout, and confirm the exit becomes 0 afterwards.
  • Can a container be OOM-affected without exiting at all?
    Yes. The kernel OOM killer selects a process, so if it kills a worker child instead of PID 1, the container keeps running while losing capacity, and the OOM flag can be set with no exit code recorded. Symptoms are silent worker loss and degraded throughput rather than a restart, which is why you check `dmesg` and the flag when an app behaves oddly even though it appears up.

saying these in an interview costs you the question

  • Concluding "OOM" from the 137 alone without checking `.State.OOMKilled` or the kernel log
  • Not knowing a `docker stop` that exceeds its grace period also produces 137
  • Assuming `OOMKilled=false` rules out memory entirely, missing host-level OOM kills
  • Believing the application could have caught SIGKILL and logged its own last words
  • Raising the memory limit as a reflex without checking whether usage was climbing monotonically (a leak) or spiking (a burst)

context