skip to content

A Go daemon's RSS stays at 900 MB while gctrace reports a 60 MB live heap after a scrape spike — how do you tell a leak from retained memory?

level: seniorimportance: must knowfreq 55%

answer

  1. flat versus climbing, not high versus low
  2. watch the third number over many cycles
  3. empty spans are still charged to the process
  4. idle minus released is the sentence you write

basics

~20 s

Track the live-heap figure in gctrace across many cycles, not one sample. Flat while resident memory stays high means the runtime is holding freed spans it has not given back. A real leak makes that figure climb steadily.

solid answer

~50 s

This is the classic leak that is not one. The decisive evidence is the **third number** of the gctrace heap triple — the live heap after marking — read across a long series of cycles. Flat means the reachable heap is not growing, no matter what resident memory says. Cross-check with `runtime.MemStats`: a small `HeapAlloc` with a large `HeapIdle` and a `HeapReleased` well below it says the memory is free inside the runtime and simply still charged to the process, because the runtime keeps a spike's worth of spans rather than returning them and re-faulting them on the next spike. Resident memory behaves like a high-water mark that decays slowly. Before you close the postmortem, rule out non-heap growth too: `StackInuse` against the goroutine count for a goroutine leak, and `Sys` minus `HeapSys` for runtime structures. If the live heap *is* climbing, the diagnosis flips and the next step is finding the allocation sites.

code

text · 9 lines
text
# during the spike
gc 118 @610.4s 2%: 0.09+41+0.07 ms clock, 0.7+12/38/0+0.5 ms cpu, 840->871->433 MB, 866 MB goal, 8 P
gc 119 @612.9s 2%: 0.11+44+0.08 ms clock, 0.8+14/41/0+0.6 ms cpu, 861->889->441 MB, 883 MB goal, 8 P

# an hour later, load back to normal
gc 402 @4210.7s 2%: 0.05+6.1+0.04 ms clock, 0.4+1.1/5.2/0+0.3 ms cpu, 121->124->60 MB, 123 MB goal, 8 P
gc 403 @4225.1s 2%: 0.06+5.9+0.04 ms clock, 0.4+0.9/5.0/0+0.3 ms cpu, 120->123->60 MB, 122 MB goal, 8 P

# live heap: 433 MB -> 60 MB and flat. RSS: still 900 MB.

go deeper

for a junior

Take away one fact: resident memory sitting above the live heap is normal for a Go process. The runtime keeps memory it has freed internally, so a gap between the two is not by itself evidence of anything wrong.

for a middle

Be able to name the specific evidence — the live-heap figure a completed cycle reports, and HeapIdle against HeapReleased — and explain why empty spans are still charged to the process even though the objects are gone.

for a senior

Run the investigation: series not sample, live heap trend first, MemStats to confirm, then goroutine stacks and non-Go memory to rule out. Say out loud what would have changed your conclusion.

for a principal

Turn the finding into a decision. This is a capacity question — the footprint tracks the worst spike — so own whether the answer is a bigger limit, a bounded intake path, or an accepted cost, and stop the team from tuning the collector to fix an input problem.

## The shape of the incident A daemon that scrapes and re-exports measurements takes a bad spike — a target returns far more series than usual — and its resident memory jumps to 900 MB. An hour later, load is normal, throughput is normal, latency is normal, and resident memory is still 900 MB. Someone files it as a memory leak. The job is to establish whether it is one, and the whole answer is available from the collector's own output plus `runtime.MemStats`. ## Step one: read the live heap across cycles, not once With `GODEBUG=gctrace=1`, every completed cycle prints a heap triple: heap at cycle start, heap at cycle end, and the live heap that survived marking. Only the third figure is a retention measure — the first two mostly reflect where the collector was triggered and how much the program allocated while marking ran concurrently. The distinction that settles the question is a trend, not a value: - **Live heap flat across hundreds of cycles.** The reachable heap is not growing. Whatever the resident set is doing, the program is not accumulating objects it cannot release. Not a leak. - **Live heap climbing cycle after cycle.** Real retention. The trigger point and the collection frequency will be rising behind it, and you now have an allocation-site hunt on your hands rather than a memory-accounting question. A single sample cannot distinguish these, which is the single most common mistake in this postmortem. Retention is a slope. ## Step two: confirm with MemStats `runtime.ReadMemStats` gives the accounting behind the trace. In the not-a-leak case the pattern is unmistakable: - `HeapAlloc` small — the objects really are gone. - `HeapInuse` small — few spans still hold anything. - `HeapIdle` large — many spans are completely empty. - `HeapReleased` well below `HeapIdle` — most of that empty memory has not been given back to the kernel. - `Sys` close to the resident set — the process did take this memory from the operating system and still has it. `HeapIdle - HeapReleased` is the number to quote in the writeup: memory that is free inside the runtime and still charged to the process. That is not a leak in any useful sense of the word. It is the runtime keeping the ground it won during the spike so that the next spike costs no syscalls and no page faults. A leak looks completely different: `HeapAlloc` large and rising, `HeapInuse` tracking it, `HeapIdle` small. The memory is in live objects, not in empty spans. ## Step three: rule out the non-heap explanations If the live heap is flat but resident memory keeps *climbing* rather than merely staying high, the growth is not in heap objects and you should stop looking there. - **Goroutine leak.** Every goroutine has a stack. `StackInuse` and `StackSys` grow with the goroutine count, and a count that only climbs is the tell. This shows up in stack accounting, not in any heap field. - **Runtime structures.** `Sys` minus `HeapSys` covers collector metadata and other runtime allocations. - **Memory outside the Go heap.** Anything mapped by C code or by direct mapping is invisible to every field discussed here, and to gctrace entirely. If the Go accounting says one number and the operating system says a much larger one that keeps growing, that gap is where to look next. ## Why resident memory behaves this way Resident memory is a high-water mark that decays lazily. The runtime frees objects promptly and returns pages to the kernel on its own schedule, deliberately unhurried, because handing memory back and immediately re-acquiring it is pure waste for a workload that spikes repeatedly. For a daemon whose load is bursty by nature — scrape intervals, batch windows — that is the correct behaviour, not a defect. It is also why steady-state resident memory converges toward the peak rather than the average, and why capacity planning for such a service should be done against the spike. ## Writing it up The postmortem should say three things and no more: the live heap was flat at roughly 60 MB across the whole window, so no retention occurred; the resident set reflects idle spans the runtime is holding, quantified as `HeapIdle - HeapReleased`; and the service's memory footprint is therefore governed by its worst spike, which is a capacity question rather than a bug. If the team wants the footprint lower, that is a conversation about bounding the spike itself, and the number to put in front of them is the peak live heap, not the resident set. ## The trap in one line Resident memory above the live heap is the normal, expected state of every Go process. Calling that a leak, or acting on a single sample of it, is what turns a capacity observation into a week of wasted investigation.

  • Which gctrace number would actually be climbing if this were a real leak?
    The third one in the heap triple, the live heap after marking, cycle after cycle. The trigger point and the collection frequency follow it upward, so you also see consecutive cycle timestamps crowding closer together. A first number that is large while the third stays flat means the collector is firing at a high heap target, not that the program is retaining anything.
  • The live heap is flat but resident memory keeps climbing anyway. Where do you look next?
    Outside the heap. Check the goroutine count against `StackInuse` and `StackSys` for goroutines that never finish — every one carries a stack. Then look at `Sys` minus `HeapSys` for runtime structures. If the operating system reports far more than `Sys` and the gap grows, the memory is mapped outside the Go heap entirely and none of these fields will ever show it.
  • Why is a single ReadMemStats sample a weak basis for this postmortem?
    Because it is one instant, and every question here is about a slope. HeapAlloc taken just before a collection looks alarming and just after looks reassuring, for the same healthy program. You need a series across many cycles to distinguish a heap that is growing from one that is oscillating around a stable live set, and the gctrace stream gives you that series for free.
  • What does this incident imply for how the service should be sized?
    Its steady-state footprint converges toward its worst spike rather than its average, because the runtime holds the ground it won. So capacity is planned against peak live heap, and the memory limit has to sit above the spike, not above the quiet-hours reading. Lowering the footprint means bounding the spike itself — capping what a single scrape can pull in — rather than tuning the collector.

A warehouse that has shipped out its stock has not gotten smaller. The runtime keeps renting the empty floor space because the next delivery is coming, and the rent bill is what your resident-memory graph shows.

saying these in an interview costs you the question

  • Calls any resident memory above the live heap a leak
  • Assumes Go returns freed pages to the OS immediately
  • Compares RSS against HeapAlloc and ignores HeapIdle
  • Draws the conclusion from one sample instead of a series
  • Never checks goroutine stacks before blaming the heap