skip to content

kubectl describe pod shows a container with Last State: Terminated, Reason: OOMKilled, Exit Code: 137, and a rising restart count. Explain exactly what happened and how you would fix it.

level: middleimportance: must knowfreq 68%

answer

  1. 137 = 128 + 9 (SIGKILL), no graceful shutdown
  2. OOMKilled = container cgroup limit; Evicted = node pressure
  3. 137 + Reason: Error = ignored SIGTERM, not memory
  4. JVM/Node need explicit max-heap below the limit
  5. requests=limits → Guaranteed QoS, but the ceiling is still hard

basics

~20 s

The container's processes exceeded the memory limit in its cgroup, so the kernel's OOM killer sent SIGKILL — exit 137 is 128+9. Fix by measuring real memory use, then either raising the limit or making the runtime respect it (for example JVM heap sizing).

solid answer

~60 s

A container's `resources.limits.memory` becomes a hard cgroup memory ceiling on the node. When the processes inside exceed it, the kernel OOM killer kills them immediately — no SIGTERM, no graceful shutdown. The kubelet records `Reason: OOMKilled` and exit code 137 (128 + SIGKILL), then restarts the container, which is why you end up in CrashLoopBackOff. Diagnosis: confirm it is the *container* limit and not node-level pressure — a container OOM kill shows OOMKilled on a Running pod's container status, while node memory pressure evicts the whole pod with status `Evicted`. Then look at actual usage over time (`kubectl top pod`, or the container working-set metric in your monitoring) and at whether it is a steady climb (leak) or a spike (large request, batch job, cache warm). Fixes, in order of preference: make the app cgroup-aware — a JVM needs `-XX:MaxRAMPercentage`, Node.js needs `--max-old-space-size` — bound the caches or batch sizes causing spikes, and only then raise the limit based on measured peaks plus headroom. Setting requests equal to limits gives the pod Guaranteed QoS and predictable behaviour.

code

bash · 3 lines
bash
kubectl get pod api-0 -o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}{" "}{.status.containerStatuses[0].lastState.terminated.exitCode}{"\n"}'
kubectl top pod api-0 --containers
kubectl exec -it api-0 -- sh -c 'cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/memory.current'

go deeper

for a junior

Know that 137 with OOMKilled means the container exceeded its memory limit and was killed by the kernel, and that the fix starts with looking at real usage.

for a middle

Explain cgroup enforcement, requests versus limits and QoS classes, and why runtimes like the JVM need explicit heap sizing under the limit.

for a senior

Separate container OOM from node eviction and from grace-period force-kills, read the working-set time series to tell a leak from a spike, and justify limit changes with measured peaks plus headroom.

for a principal

Discuss cluster-wide policy: request/limit ratios and oversubscription, defaults enforced by LimitRange or admission policy, the capacity cost of padded limits, and alerting on OOM rate as a leading indicator.

## The mechanism Every container you give `resources.limits.memory` runs in a cgroup with that value as a hard ceiling. Memory here means the container's working set: anonymous memory (heap, stacks), plus page cache attributable to it, minus what the kernel can reclaim. When an allocation cannot be satisfied within the ceiling and nothing more can be reclaimed, the kernel's OOM killer picks a process in that cgroup and sends SIGKILL. SIGKILL cannot be caught or handled, so the application gets no chance to flush, log or shut down cleanly. The kubelet observes the termination and records it in the container status: `lastState.terminated.reason = OOMKilled`, `exitCode = 137`. The 137 is the shell convention of 128 plus the signal number, and 9 is SIGKILL. With the default `restartPolicy: Always` the container is restarted, and if it OOMs again quickly you land in CrashLoopBackOff. ## Distinguishing the three memory failures This matters because the fixes differ: 1. **Container limit OOM** — what the question describes. One container is killed; the pod object survives; other containers and the node are unaffected. Cause: the workload wants more memory than its own limit. 2. **Node memory pressure eviction** — the kubelet, not the kernel, decides the node is running low and evicts pods (status `Evicted`, and the pod is not restarted in place). Cause: the node is oversubscribed, usually because requests were set far below real usage. 3. **Force-kill after grace period** — exit 137 again, but `Reason: Error`, not OOMKilled. The process ignored SIGTERM and the kubelet escalated to SIGKILL after `terminationGracePeriodSeconds`. Nothing to do with memory. So exit code 137 alone is not enough; always read the Reason next to it. ## Requests, limits and QoS `requests.memory` is what the scheduler reserves; `limits.memory` is the enforcement ceiling. Their relationship sets the pod's QoS class: equal requests and limits on every container gives **Guaranteed**; a request lower than the limit gives **Burstable**; neither set gives **BestEffort**. QoS influences which pods the kubelet evicts first under node pressure, but it does **not** soften the hard limit — a Guaranteed pod that exceeds its limit is killed just the same. If you set no memory limit at all, the container can never be OOMKilled by its own cgroup, but it can starve the node and get the pod evicted, which is usually worse because it takes neighbours down with it. ## Why applications exceed limits - **Runtime not cgroup-aware.** Older JVMs sized the heap from the host's total RAM, happily choosing a 16 GB heap inside a 1 GB container. Modern JVMs honour container limits by default, but you still need `-XX:MaxRAMPercentage` because the *total* process footprint is heap plus metaspace, thread stacks, code cache, direct byte buffers and native allocations. A heap set to 100% of the limit will always be OOMKilled. Node.js needs `--max-old-space-size` below the limit for the same reason. - **Unbounded growth.** An in-memory cache with no eviction, an unbounded queue, an accumulating list — a leak that shows as a slow climb toward the ceiling with a sawtooth of restarts. - **Spikes.** Reading a whole file or result set into memory, a large export, a batch that scales with input size. Steady-state usage looks fine and the pod dies only occasionally. - **Limit set by copy-paste.** Someone pasted `512Mi` from another service without measuring. ## Investigating properly Look at the memory time series for the container — the working-set metric — over hours or days, across all replicas. A straight upward line ending at the limit is a leak; a flat line with occasional vertical spikes is a workload-dependent allocation. Correlate the kill timestamps with traffic, deploys and batch schedules. Inside the container, `/sys/fs/cgroup/memory.max` and `memory.current` (cgroup v2) show the limit and current usage as the kernel sees them, which is what the runtime should be reading. For JVMs, capture a heap histogram or dump before the kill if you can, or enable native memory tracking when the heap looks innocent and the total footprint does not. ## Fixing Order matters. First make the runtime respect the limit — that is free and prevents a whole class of failure. Second, bound the thing that grows: cache size, page size, batch size, streaming instead of buffering. Third, raise the limit to measured peak plus headroom (a rule of thumb is 20–30%, more for GC-heavy runtimes), and raise the request with it so the scheduler actually places the pod where the memory exists. Consider setting request equal to limit for latency-sensitive services so they are not competing for memory they may not get. What not to do: doubling the limit repeatedly without measuring, which just moves the failure further out and wastes cluster capacity; or removing the limit entirely, which converts a contained container failure into a node-level incident.

  • How do you tell an OOMKilled container apart from a pod evicted because of node memory pressure?
    An OOMKilled container is killed by the kernel inside its own cgroup: the pod object stays, the container status shows Reason OOMKilled with exit 137, and the restart count increases. A node-pressure eviction is a kubelet decision: the pod's phase becomes Failed with reason Evicted, the pod is not restarted in place, and the node reports a MemoryPressure condition. The first means the container's own limit is too low or its usage too high; the second means the node is oversubscribed, usually because requests understate real usage.
  • Setting a JVM's -Xmx equal to the container memory limit is a common mistake. Why does it still get OOMKilled?
    The cgroup limit covers the entire process footprint, not just the Java heap. Metaspace, thread stacks, the JIT code cache, GC structures, direct/mapped byte buffers and any native libraries all allocate outside the heap. If the heap alone is allowed to reach the ceiling, the first non-heap allocation pushes the container over it. Leave headroom — a MaxRAMPercentage around 65–75% is a typical starting point, tuned against measured native usage.

The memory limit is a lift's weight sensor, not a suggestion. Once you step over the rating the doors do not politely ask anyone to leave — the lift simply stops, mid-journey, with everything you were carrying still inside.

saying these in an interview costs you the question

  • Reading exit code 137 as automatically meaning out-of-memory without checking whether the Reason is OOMKilled or Error
  • Believing Guaranteed QoS or a matching request protects a container from being killed at its limit
  • Setting the JVM max heap equal to the container memory limit
  • Removing the memory limit to 'fix' the OOM, converting a contained failure into node-wide pressure
  • Expecting the application to catch the kill and shut down gracefully — SIGKILL cannot be handled

context