skip to content

Why can a Go service be OOM-killed with GOMEMLIMIT set to the container's memory limit?

level: seniorimportance: must knowfreq 52%

answer

  1. the kernel counts more than the runtime does
  2. soft means it may be exceeded
  3. cgo, mapped files and the binary sit outside
  4. concurrent marking allows a brief overshoot
  5. reserve the non-Go footprint you measured

basics

~10 s

GOMEMLIMIT governs only memory the Go runtime manages, and it is a target rather than a cap. The binary, cgo allocations and mapped files sit outside it, so an equal setting leaves no headroom.

solid answer

~50 s

Two things are wrong with setting the ceiling equal to the allowance. First, `GOMEMLIMIT` only accounts for **runtime-managed** memory — the Go heap, goroutine stacks and runtime metadata. Everything else the process consumes is invisible to it: the mapped binary, memory allocated by C through cgo, memory-mapped files. The kernel counts all of that, so the process is always larger than the number you set. Second, the ceiling is soft. If the live working set legitimately grows past it, the collector cannot reclaim reachable data, so the runtime keeps allocating and simply exceeds the target — and it can overshoot transiently anyway, because allocation can outrun a concurrent mark phase. The fix is to set the ceiling *below* the allowance by the measured non-Go footprint plus a margin for that overshoot, and to treat a live set that genuinely exceeds the remainder as a capacity problem, not a tuning one.

code

go · 10 lines
go
s := []metrics.Sample{
	{Name: "/memory/classes/total:bytes"},
	{Name: "/memory/classes/heap/released:bytes"},
}
metrics.Read(s)

// Runtime-managed memory: what GOMEMLIMIT governs.
// Anything the platform reports beyond this is outside its accounting.
managed := s[0].Value.Uint64() - s[1].Value.Uint64()
log.Printf("runtime-managed bytes: %d", managed)

go deeper

for a junior

Take away the headline: the ceiling only budgets the Go runtime's own memory, so it must be set below the container's allowance rather than equal to it. You are not expected to size the gap yourself.

for a middle

Be able to list what falls outside the ceiling's accounting — the binary's mappings, cgo allocations, mapped files — and explain why a concurrent collector can briefly overshoot its target even when everything is configured correctly.

for a senior

This is the production judgment call: distinguish an accounting problem, which a lower ceiling fixes, from a live-data problem, which no ceiling fixes. Describe the experiment that tells them apart rather than asserting a cause.

for a principal

Frame it as capacity rather than configuration. When the working set genuinely exceeds the remainder, the decision to buy a larger allowance, shard the work, or retain less state belongs to whoever owns the memory budget, and tuning cannot substitute for it.

## The scenario An aggregation service runs in a container with a fixed memory allowance. It holds its working set in one large map that grows with the size of the window it is aggregating. Someone reads about `GOMEMLIMIT`, sets it to exactly the container's allowance, and ships. The service is still killed — abruptly, with no panic and no Go stack trace, because the process is terminated from outside rather than failing from inside. There are three distinct reasons this happens, and a good answer separates them. ## Reason 1: the ceiling accounts for less than the kernel does The limit covers memory **the Go runtime manages**: the heap, goroutine stacks, and the runtime's own bookkeeping. It excludes mappings of the binary itself, memory allocated by C code called through cgo, memory-mapped files, and memory the operating system holds on the program's behalf. Whatever is doing the killing counts the whole process. So the process footprint is structurally larger than the quantity the ceiling governs, and setting the two equal guarantees you are over. How much larger depends entirely on the program: a pure-Go service with a small binary may sit only tens of megabytes above it; a service that maps a large index file or calls into a C library may sit hundreds of megabytes above it, or worse. The practical consequence: **the ceiling is a budget for the Go part only**, and it must be set to the allowance minus everything else, minus a margin. That number is measured, not guessed — you can read the quantity the runtime is comparing against from `runtime/metrics` and compare it with the footprint your platform reports, and the difference is the non-Go part you have to reserve. ## Reason 2: soft means the runtime is allowed to lose The collector reclaims unreachable objects. It cannot reclaim a live map. If the aggregation window grows so that live data alone approaches the ceiling, the pacer responds by collecting more and more often, finds progressively less garbage each time, and the total keeps climbing anyway. Nothing in the runtime refuses the allocation. The program crosses the ceiling, then crosses the allowance, and is killed. This is the case people most often misdiagnose as "the limit didn't work". It worked exactly as specified. The mistake was expecting a pacing target to enforce something the program's own data makes impossible. ## Reason 3: overshoot is normal, briefly The collector is concurrent: the program keeps allocating while marking is in progress. If the allocation rate is high enough, the total can rise past the target before the cycle finishes reclaiming anything. The pacer's job is to start early enough that this is rare and small, but a burst — a sudden spike in request size, a batch that arrives all at once — can produce a real transient above the ceiling. If the ceiling is flush with the allowance, that transient is fatal; if there is headroom, it is invisible. ## Confirming it rather than guessing The useful experiment is a **soak test held near the ceiling**: run the service with a working set deliberately sized close to the number, hold it there, and watch what happens over time with `GODEBUG=gctrace=1` printing a line per cycle. Two shapes appear, and they mean different things: - Cycles at a healthy spacing, the heap goal well under the ceiling, and yet the process is killed — the extra footprint is outside the ceiling's accounting. Reason 1. Lower the ceiling; the Go part was never the problem. - Cycles collapsing into each other with the goal pinned at the ceiling — live data has reached the ceiling. Reason 2. No ceiling value fixes this; the working set has to shrink or the allowance has to grow. Moving the ceiling between runs, and setting it from code with `debug.SetMemoryLimit` so it can be changed without a rebuild, turns this from an argument into a measurement. ## Setting the number A workable procedure: 1. measure the process footprint with the Go ceiling set very low, so almost all of it is non-Go memory — that is your fixed cost; 2. subtract that fixed cost, plus a margin for allocation overshoot, from the container's allowance; 3. set the ceiling to what is left, ideally derived at startup from the allowance rather than hardcoded; 4. verify with the soak test above that the service's real working set fits comfortably inside it. If step 4 fails, you have discovered a capacity fact, not a tuning failure. Aggregating a larger window needs more memory, and the honest options are less retained state, a larger allowance, or partitioning the work — never a lower ceiling, which only exchanges the kill for something slower and no less broken.

  • How would you measure the non-Go part of the footprint to decide the headroom?
    Run the service with the ceiling set deliberately low so the Go part is small and bounded, then compare the footprint the platform reports against the runtime-managed total exposed by `runtime/metrics`. The difference is the fixed cost — binary mappings, cgo allocations, mapped files — that has to be reserved out of the allowance before you decide the ceiling.
  • The soak test shows the heap goal pinned at GOMEMLIMIT and cycles back to back. What does that tell you?
    That the live working set itself has reached the ceiling, so the collector has no garbage left to buy room with. No value of the ceiling fixes it: raising it postpones the problem, lowering it converts the kill into permanent thrashing. The real options are retaining less state, partitioning the work, or getting a larger allowance.
  • Why is there no Go stack trace when this happens?
    Because the process was terminated from outside rather than failing from within. A Go program that cannot get memory from the operating system does report a runtime failure, but a process ended by an external memory enforcement mechanism is simply stopped, so the last thing in the log is ordinary output. Absence of a Go-level error is itself the diagnostic signal.

saying these in an interview costs you the question

  • Expects the ceiling to prevent the process being killed
  • Sets GOMEMLIMIT equal to the container's memory allowance
  • Forgets cgo and mapped memory sit outside the ceiling's accounting
  • Lowers the ceiling further when live data already exceeds it
  • Looks for a Go panic or stack trace that will never exist