skip to content

A Go controller with GOMEMLIMIT set runs GC cycles back to back as its live heap grows - why?

level: seniorimportance: nice to knowfreq 26%

answer

  1. headroom is goal minus live
  2. the ceiling does not grow with the live set
  3. each cycle reclaims almost nothing
  4. cost rises on two axes at once
  5. something has to cap collection CPU

basics

~20 s

The live heap has grown until it nearly fills GOMEMLIMIT, leaving almost no headroom between live memory and the limit-derived goal. Each cycle reclaims little and the next triggers at once, so GC CPU climbs while throughput collapses.

solid answer

~50 s

Collection frequency is roughly allocation rate divided by headroom, and headroom is the gap between the live heap and the goal. Once the live set grows large enough that the limit-derived goal sits just above it, that gap collapses: a cycle finishes, frees almost nothing because the memory is genuinely reachable, and the heap is immediately back at the goal. GC CPU share climbs while useful work falls - the pacer thrashing, each cycle starting before the previous one paid for itself. The runtime's GC CPU limiter bounds it by capping collection at 50% of CPU, after which the soft limit is exceeded rather than defended. Confirm it by running the same load at two `GOGC` settings: if cycle count and GC CPU share barely move, the limit is setting the goal. The finding is a grown working set, not a misconfigured collector.

go deeper

for a junior

Know that a Go program collects when its heap reaches a goal, and that if the goal sits barely above the memory that is genuinely still in use, collecting cannot free anything worth the effort.

for a middle

Be able to derive it: headroom is goal minus live heap, cycle frequency is allocation rate over headroom, and a fixed ceiling means headroom shrinks as the live set grows.

for a senior

Demonstrate the separation - the collector is behaving as specified, the live working set grew, and the cost rose on two axes at once. Name the experiment that confirms which goal is binding before proposing anything.

for a principal

Own the framing that a soft memory limit converts a memory problem into a CPU problem and then, at the CPU cap, back into a memory one, and be able to say which of those your service can absorb.

## The shape of the failure A controller reconciling desired state against observed state keeps its whole watched set in memory, so its live heap grows with the size of the estate it is watching. It runs under `GOMEMLIMIT`. For months nothing is remarkable. Then, past some estate size, garbage-collection CPU share climbs steeply, cycles start following one another with no gap, and request latency degrades across the board without any single operation being slow. Nothing has leaked, no code has changed, and no knob has moved. ## The arithmetic behind it Two relationships explain the whole thing. **Headroom is what the program allocates into.** Headroom is `goal - live heap`. With `GOGC` alone the goal is a multiple of the live heap, so headroom grows in step with the live set - a bigger program gets bigger, rarer cycles. With `GOMEMLIMIT` in play, the pacer follows whichever goal is smaller, and the limit-derived goal is a fixed ceiling minus everything else the runtime holds. It does not grow with the live set. It shrinks against it. **Frequency is allocation rate over headroom.** Hold the allocation rate roughly constant and shrink headroom towards zero, and cycle frequency rises without bound. This is the pacer thrashing: a cycle completes, reclaims a small amount of genuine garbage, and the heap is straight back at the goal, so the next cycle triggers before the previous one has come close to paying for its own scan. Each cycle is also getting *more expensive* at the same time, because scan work is proportional to the live set that just grew. More frequent cycles over a larger live set is a multiplicative cost, which is why the degradation looks like a cliff rather than a slope. ## What stops it running away Left alone, that loop would consume every processor. The runtime includes a **GC CPU limiter** that caps total garbage-collection CPU at 50%. When it binds, the collector stops trying to hold the heap under the soft limit and lets it be exceeded instead. That is the designed behaviour of a *soft* limit, and it produces a diagnostic signature worth recognising: memory quietly drifts above the configured limit at the same moment CPU stops climbing. Someone reading that as "the limit is not working" has it backwards - the limit worked until collecting harder cost more than it was worth, then yielded. ## Confirming it rather than guessing The cheap experiment is to run the same load test twice at two different `GOGC` settings and compare cycle count and GC CPU share between the runs. - If both figures move roughly as `GOGC` predicts, `GOGC` is setting the goal and pacing is ordinary. - If both figures barely move, the limit-derived goal is the smaller of the two in both runs and `GOGC` is not participating in the decision at all. That is the confirmation. The second thing to establish is whether the live heap growth is legitimate. A working set that scales with the estate being reconciled and a slow leak look identical on a memory graph; they differ in whether the live set plateaus when the estate stops growing. Holding the estate fixed and watching whether the live heap settles separates them in one run. ## The judgment the diagnosis leads to Once it is established that the live set genuinely grew and the collector is behaving exactly as specified, the space of real answers is about the working set and the capacity around it, not about pacing. Reducing what is retained, or giving the process room proportional to what it must retain, addresses the cause. Reaching for a pacing knob does not: lowering `GOGC` cannot help, because `GOGC` is not the binding constraint, and where headroom is already near zero there is no ratio that creates more of it. The senior move in the interview is exactly that separation - showing that the collector is not the fault, naming the mechanism that made the cost non-linear, and knowing which observation distinguishes a growing working set from a leak.

  • What stops this becoming an unbounded death spiral?
    The runtime's GC CPU limiter caps total collection CPU at 50%. Once it binds, the collector stops defending the soft limit and lets memory rise above it instead. Memory drifting over the configured limit exactly as CPU stops climbing is the signature of that valve opening.
  • How would a load test show that the memory limit, not GOGC, is setting the goal?
    Run the same load twice at two different GOGC settings and compare cycle count and GC CPU share. If they move as GOGC predicts, GOGC is binding. If they barely differ between the two runs, the limit-derived goal was the smaller one in both, so GOGC never entered the decision.
  • Would the same pattern appear with no memory limit configured?
    No. With GOGC alone the goal is a multiple of the live heap, so headroom grows as the live set grows and cycles become larger and rarer rather than tighter. You would see rising peak memory and longer scans instead of back-to-back cycles - a different failure, on a different resource.
  • How do you separate a legitimately growing working set from a leak here?
    Stop the growth in the input - hold the reconciled estate fixed - and watch whether the live heap plateaus. A working set proportional to the estate settles; a leak keeps climbing against a constant input. On a memory graph alone the two are indistinguishable.

It is a removal van that must be emptied before more can be loaded, parked outside a house that is nearly full. Every trip out carries a single box, and the van is back at the door immediately.

saying these in an interview costs you the question

  • Calls it a leak without checking whether the live set plateaus
  • Says a soft limit means the runtime will kill the process
  • Expects lowering GOGC to reduce GC CPU when the limit is binding
  • Assumes garbage-collection CPU can grow without bound
  • Reads memory drifting above the limit as the limit not working