skip to content

After lowering GOMEMLIMIT, a Go service stops crashing but pins CPU and stalls — what is happening?

level: seniorimportance: nice to knowfreq 33%

answer

  1. alive but useless
  2. the collector never gets a break
  3. heap goal pinned at the ceiling
  4. a limiter prevents the total spiral
  5. less live data, or more room

basics

~10 s

Live data now sits just under the ceiling, so the collector restarts almost as soon as a cycle ends. Because GOMEMLIMIT is soft the process survives, but collection takes the CPU and throughput collapses.

solid answer

~50 s

Lowering the ceiling did not shrink the working set; it moved the target below what the working set needs. The pacer now aims at a number the collector can barely reach, so cycles run back to back, each one reclaiming very little, and the program's own goroutines get a shrinking share of the CPU. Latency explodes and throughput collapses, but the process stays up, because the ceiling is a target and not an enforced cap — the runtime is allowed to exceed it rather than stop. The runtime's GC CPU limiter caps collection at roughly half the CPU so the program keeps making some progress instead of freezing entirely, which is precisely why you get a catatonic service rather than a clean failure. You have traded an abrupt death for a slow one. The real fix is less live data, or more room — not a lower ceiling.

go deeper

for a junior

The takeaway is that a lower ceiling does not make a program keep less data. It only leaves the collector less room, and when that room runs out the program gets slow rather than safe.

for a middle

Explain the mechanics: cycle spacing is the gap between live data and the ceiling, so closing that gap means back-to-back cycles, and the softness of the ceiling is why the process survives instead of failing.

for a senior

Demonstrate that you would recognise this in production from cycle spacing and GC CPU share rather than from memory alone, and that you would treat a degraded surviving instance as an unresolved incident, not a fix that worked.

for a principal

Own the point that this mode can be worse than the crash it replaced, because a degraded instance keeps taking traffic. Whether health checks should fail on degradation is a decision with fleet-wide consequences.

## What you actually changed Lowering the memory ceiling changes one thing: the number the pacer aims at. It does not change how much data the program keeps reachable. If an aggregation service holds a window's worth of state in one large map, that map is the same size before and after the change — the collector simply has less room in which to work around it. As long as the live set is comfortably below the ceiling, the difference between them is the space that garbage can fill before a cycle must run. Halve the ceiling and you halve that space. Bring the ceiling close to the live set and the space nearly vanishes: a cycle finishes, the program allocates a small amount, the total is at the target again, and another cycle starts. That is the thrash. ## Why the process does not simply die Because the ceiling is soft. Nothing in the runtime refuses an allocation or terminates the program for being over target. The pacer's only lever is *how hard the collector works*, and when that lever runs out, the runtime does the pragmatic thing: it lets the program exceed the ceiling. Specifically, a **GC CPU limiter** caps collection at roughly half the available CPU. Without it, a collector chasing an unreachable target could take essentially all the CPU and the program would make no progress at all — a true death spiral. With it, the program keeps a share of the machine and slowly proceeds, while memory drifts above the target. So the observable state is a process that is alive, responds to health checks slowly or not at all, burns CPU constantly, and serves almost nothing. From an availability standpoint this is often *worse* than the crash it replaced, because a dead process gets replaced and a catatonic one keeps taking traffic. ## Recognising it The signature is unmistakable once you look for it: - `GODEBUG=gctrace=1` prints one line per cycle: the lines arrive continuously, with almost no gap, and the heap goal is pinned at or just under the ceiling rather than tracking live data; - the fraction of CPU attributed to garbage collection climbs and then plateaus near the limiter's cap; - the runtime exposes, through `runtime/metrics`, both the cumulative count of completed cycles and the last cycle in which the CPU limiter was engaged — a limiter that has been enabled recently is a direct confirmation rather than an inference; - application latency degrades in proportion to the CPU the collector took, not in proportion to load. The cleanest way to see all of it is the experiment this failure deserves: a **soak test held near the ceiling**. Size the workload so the live set approaches the number, hold it there for minutes rather than seconds, and move the ceiling between runs — from code with `debug.SetMemoryLimit`, so a run needs no rebuild. You will see a clear knee: above some value, cycles are spaced and throughput is flat; below it, cycles merge and throughput falls off a cliff. That knee is a property of the service's live set, and it tells you exactly how much room the working set actually needs. ## The fixes, in order of honesty **Reduce the live set.** This is the only fix that changes the underlying fact. Bound the aggregation window, evict or spill completed groups, use a more compact representation for the accumulated state, or partition the work across more instances so each holds less. Everything else moves the pain around. **Give it more room.** Raise the ceiling — and, if that puts the process over its allowance, raise the allowance. That is a capacity decision with a cost attached, but it is a legitimate answer: some workloads simply need the memory. **Choose the other failure.** Remove the ceiling, accept that the process will occasionally be killed, and let it be restarted. This is a real option, not a defeat, when a restarted instance is more useful than a degraded one — but it is a policy decision about what the service owes its callers, not a tuning trick. **Not a fix:** forcing extra collections, or lowering the ceiling further. The collector is already running as often as it can, and the problem is that it cannot find anything to reclaim. ## The lesson to carry A memory ceiling is a genuinely useful tool, but it converts an out-of-memory failure into a performance failure. That conversion is only an improvement while there is real garbage for the collector to find. Once the live set approaches the ceiling, the ceiling has nothing left to give, and continuing to tune it is choosing between two shapes of the same underlying shortage.

  • What does Go's GC CPU limiter do here, and why does it exist?
    It caps garbage collection at roughly half the available CPU. Without it, a collector chasing a target it can never reach would consume nearly the whole machine and the program would stop making progress entirely. With it, the runtime prefers exceeding the soft ceiling over starving the program, which is why the outcome is a slow service rather than a frozen one.
  • Why is a thrashing instance often worse operationally than one that is killed?
    A dead process is replaced quickly and stops taking traffic. A thrashing one stays registered, keeps accepting requests it cannot serve in time, and may drag callers into timeouts and retries that spread the damage. Unless health checks are tuned to fail on degradation rather than only on death, the degraded instance is the more expensive of the two.
  • How do you find the right ceiling experimentally rather than by argument?
    Hold a soak test near the ceiling and move the value between runs with `debug.SetMemoryLimit`, watching cycle spacing under `GODEBUG=gctrace=1` alongside throughput. There is a clear knee: above it cycles are spaced and throughput is flat, below it cycles merge and throughput collapses. That knee measures how much room the live set genuinely needs.

It is the difference between a car that stalls and one held at the redline in first gear — still running, burning everything, going almost nowhere.

saying these in an interview costs you the question

  • Lowers the ceiling again because the process stopped being killed
  • Blames the collector for a bug rather than an impossible target
  • Forces extra collections with runtime.GC to relieve the pressure
  • Treats a surviving process as a resolved incident
  • Assumes a lower ceiling reduces how much data the service retains