skip to content

How do you tell a Go 'fatal error: out of memory' from the kernel OOM-killing the process?

level: seniorimportance: nice to knowfreq 30%

answer

  1. one writes to your logs, one does not
  2. check whether stderr has anything at all
  3. the kernel uses a signal, not a message
  4. exit 137 versus a full goroutine dump

basics

~20 s

Look for output. Go's abort writes 'fatal error: out of memory' plus goroutine stacks to stderr and exits with status 2. A kernel OOM kill sends SIGKILL, so there is no Go output and the container reports exit code 137.

solid answer

~50 s

The two look identical on a dashboard and completely different in the logs. If the Go runtime asks the operating system for memory and is refused, it throws `fatal error: out of memory`, prints a diagnostic naming the size it could not allocate and how much was already in use, dumps goroutine stacks, and exits with status 2. You get a snapshot of every goroutine at the instant of death, so you can see which call path was allocating. If the kernel's out-of-memory killer picks the process instead, it delivers SIGKILL: nothing Go-authored is written anywhere, the process simply disappears, the exit status surfaces as 137 under a container runtime, and the only record is in the kernel log. The postmortem rule is short - no Go output means the kernel decided, so look at the memory limit and the allocation trend rather than hunting for a guilty goroutine.

code

text · 10 lines
text
# Go's own abort: written by the runtime to stderr, exit status 2
runtime: out of memory: cannot allocate 4294967296-byte block (2147483648 in use)
fatal error: out of memory

goroutine 51 [running]:
...

# Kernel OOM kill: nothing from Go at all, only the status
$ echo $?
137

go deeper

for a junior

Know that a Go program can die of memory in two different ways, and that the fast way to tell them apart is whether the process wrote anything before it disappeared.

for a middle

Explain what the runtime prints when an allocation fails, why a SIGKILL leaves no output, and what exit status each produces so you can read a crash report correctly.

for a senior

An interviewer expects you to run the triage: classify the death, read the goroutine dump if there is one, distinguish a steady climb from an input-driven spike, and turn that into a concrete fix rather than a memory-limit increase.

for a principal

Own the evidence policy: what diagnostic settings services run with by default, what memory signals are retained long enough to establish a trend, and how postmortems are required to state which of the two failures actually occurred.

## Two ways a Go process dies of memory **The runtime gives up.** The Go runtime needs more memory from the operating system, asks for it, and is refused. It cannot continue, so it throws a fatal error. You see something in the shape of `runtime: out of memory: cannot allocate <n>-byte block (<m> in use)` followed by `fatal error: out of memory`, then goroutine stacks on standard error, then exit status 2. **The kernel gives up.** The process crosses a memory limit - a cgroup limit in a container, or overall system pressure - and the kernel's out-of-memory killer selects it and sends SIGKILL. SIGKILL cannot be caught, blocked or handled. The Go runtime never gets a chance to write anything. The process is simply gone. ## Telling them apart in a postmortem | Evidence | Runtime abort | Kernel OOM kill | |---|---|---| | Application stderr | `fatal error: out of memory` plus goroutine stacks | nothing at all | | Exit status | 2 | reported as 137 by container runtimes (128 plus signal 9) | | Where the record lives | your logs | the kernel log, and your orchestrator's event or termination reason | | Evidence quality | a stack for every goroutine at the moment of failure | none from the process itself | The single question that resolves it: **did the process write anything on the way out?** If yes, the runtime decided and you have a dump to read. If the last log line is an ordinary request completing and then silence, the kernel decided. ## Which one you actually get In a container with a memory limit, the kernel usually wins. The cgroup limit sits well below the point at which the host would refuse an allocation outright, so the process crosses its limit while the runtime's allocations are still succeeding perfectly well. That is why containerised Go services almost always die the silent way, and why engineers coming from a bare-metal background are surprised there is no dump. The runtime's own abort is more common when there is no container limit and the machine itself runs out, or when a single enormous allocation is requested - a length or capacity computed from untrusted input, a file read whole into memory, a decoder that trusts a declared size. A huge allocation with a plainly invalid length gives a panic instead, so a fatal out-of-memory usually means the size was legal but unavailable. ## Reading the dump when you do get one The goroutine dump is the payoff. Because nothing unwinds, every stack is exactly where it was: - The goroutine at the top, the one that made the failing allocation, names the call path that asked for the memory. - The size reported in the message tells you whether the failure was one absurd request or ordinary demand hitting a wall. - If several hundred goroutines are sitting in the same handler, each holding a buffer, the shape of the problem is concurrency without a bound rather than a single leak. How much of that you get is governed by `GOTRACEBACK`. `none` suppresses the goroutine stacks entirely, `all` includes every user goroutine, `system` adds runtime frames and goroutines the runtime created itself, and `crash` makes the process terminate in an operating-system-specific way so a core dump can be collected. It is read by the process that crashes, so setting it never improves the dump already in your hands - it improves the next one. Services that are expected to abort under memory pressure are worth running with it raised so the first crash is already readable. ## What neither of them does Neither path runs deferred functions. No buffered output is flushed, no temporary file is removed, no shutdown hook executes. Anything you needed to persist had to have been persisted before the crash. ## Turning the finding into an action - **Kernel OOM kill.** Establish the trend before the kill: was memory climbing steadily over hours, or did it spike in seconds? A steady climb points at retained state; a spike points at a request or a batch whose size is driven by input. Correlate the kill time with traffic and with deploys, and check whether the limit was recently lowered. - **Runtime abort.** Read the dump first. A single failing allocation of an implausible size is an input-validation bug and is fixed by bounding the size, not by adding memory. - **Either way.** Record which of the two happened in the postmortem explicitly. Writing 'the service OOMed' without saying whether Go or the kernel decided throws away the most useful fact you had.

  • What does the Go abort give you that a SIGKILL does not?
    Every goroutine's stack at the instant the allocation failed, plus the size of the block that could not be obtained and how much was already in use. That usually names the allocating call path directly. A SIGKILL leaves nothing from the process, so you reconstruct the cause from memory metrics sampled before the kill.
  • Why is the runtime's own out-of-memory abort rarer in a container than an OOM kill?
    Because a cgroup memory limit normally sits far below the point at which the host would refuse an allocation. The process crosses its limit while the runtime's requests are still being satisfied, so the kernel kills it before Go ever sees a failed allocation.
  • Does raising GOTRACEBACK help with a dump you already have?
    No. It is read by the process that crashes, so it only shapes future crashes. If you expect a service to abort, run it with the level raised in advance; otherwise the first crash gives you whatever the default produced and you have to reproduce to get more.
  • Was any deferred cleanup performed in either case?
    None. A fatal runtime error terminates without unwinding, and SIGKILL cannot be handled at all, so nothing in the program runs. Buffered writes are lost, temporary files remain, and any state that existed only in memory is gone - which is why crash-time durability cannot live in a deferred call.

saying these in an interview costs you the question

  • Treats every container OOM as a Go runtime out-of-memory error
  • Expects a goroutine dump after a SIGKILL
  • Reads exit code 137 as an application-level panic
  • Assumes buffered logs were flushed before the process died
  • Declares a leak from one crash with no memory trend