skip to content

An index service rebuilds a snapshot hourly; its heap profile halves each time but container memory never drops. Leak or retention?

level: seniorimportance: should knowfreq 42%

answer

  1. the dashboard and the profile watch different things
  2. confirm the live heap actually fell first
  3. peak during rebuild, not steady state
  4. idle minus released is the retained part
  5. one forced release on a canary settles it

basics

~20 s

Read runtime.MemStats beside the profile. If HeapAlloc really fell and HeapIdle minus HeapReleased grew, nothing leaked: the runtime kept the pages from the rebuild peak. If HeapAlloc stayed high, the old snapshot is still reachable and it is a leak.

solid answer

~50 s

Work the numbers in order rather than arguing about the graph. First confirm the live heap actually fell, with `runtime.MemStats.HeapAlloc` — not the profile, which is only current as of the last completed GC. If it did, compute `HeapIdle - HeapReleased`; a large value says the runtime is retaining free pages, which is exactly what you expect from a rebuild that holds the old and new snapshots live at the same time and therefore peaks at roughly twice the snapshot size. Next check whether `Sys` accounts for the RSS at all; if it does not, the memory is off-heap. Confirm on one canary with a single `debug.FreeOSMemory` call: if RSS drops, it was retained free memory and there is no leak. The fix is then to lower the peak — build into a reused buffer, or release the old snapshot sooner — because peak, not steady state, sizes what the runtime holds.

code

go · 10 lines
go
type Index struct {
	mu   sync.RWMutex
	data []byte // the packed snapshot
}

func (ix *Index) Swap(next []byte) {
	ix.mu.Lock()
	ix.data = next // only here does the old snapshot become unreachable
	ix.mu.Unlock()
}

go deeper

for a junior

Know that a container's memory graph and a Go heap profile can disagree without anything being broken, and that runtime.MemStats is the next place to look rather than the incident channel.

for a middle

Walk the numbers in order: HeapAlloc to confirm the live heap really fell, then HeapIdle minus HeapReleased to see what the runtime kept. Explain why the rebuild's peak matters more than its steady state.

for a senior

Demonstrate a decision tree that ends in an action: confirm on a canary with one FreeOSMemory call, then attack the peak by reusing the buffer instead of holding two full snapshots at once.

for a principal

Own the tradeoff between a simple rebuild that briefly doubles memory and a streaming rebuild that is harder to review and maintain. Decide which one the team carries, and size the instances to match that decision.

## The symptom A service keeps a large packed snapshot in memory as a single `[]byte` and answers lookups against it. Every hour it rebuilds. After each rebuild the heap profile shows the live heap back where it started, roughly half of what it was mid-rebuild. The container's memory graph, meanwhile, plateaus after the first rebuild and never comes down. Someone asks whether the service leaks. The honest answer is that the graph and the profile are measuring different things, and the question is which of three distinct situations you are in. There is a decision procedure, and it takes about five minutes. ## Step 1 — confirm the live heap really fell Do not take this from the profile. A heap profile is a snapshot as of the most recently completed GC cycle, so it can be stale by a whole cycle, and around a rebuild the cycles are exactly where the interesting transitions happen. Call `runtime.ReadMemStats` and read `HeapAlloc`, which is current as of the call. - `HeapAlloc` fell and stays at the pre-rebuild level → the old snapshot really was collected. Go to step 2. - `HeapAlloc` did not fall, or falls less each hour → **this is a leak**, and step 2 is irrelevant. Something still references the old snapshot. ## Step 2 — measure what the runtime is retaining ``` retained := HeapIdle - HeapReleased ``` Idle spans hold no objects. Idle spans that have not been released are mapped, resident, and empty. If `retained` is large — comparable to the snapshot size — you have your answer: **nothing leaked; the runtime is holding the arena it needed at the peak.** And the peak is the crux. A rebuild that must keep serving lookups cannot free the old snapshot until the new one is ready, so both are live simultaneously and the peak live heap is roughly twice the snapshot. The runtime sized its heap for that peak. Afterwards, the live set halves, but the pages do not evaporate — they become idle spans the runtime keeps so it can do the same thing again next hour without asking the kernel for anything. That is the runtime behaving exactly as designed. ## Step 3 — decide whether the memory is even in the Go heap If `HeapAlloc` is small and `retained` is small but RSS is still high, the memory is not in the object heap. Read `Sys`: it bounds the address space the runtime obtained across heap, stacks and internal structures. If `Sys` is far below RSS, look outside the runtime entirely — cgo allocations, explicit mappings, files mapped into the address space. If `Sys` is high but the heap fields are low, look at goroutine stacks, which are counted in `Sys` and never in a heap profile. ## Step 4 — the confirming experiment On one canary instance, call `debug.FreeOSMemory` once and watch RSS. - RSS drops sharply → it was retained free heap memory. Confirmed: no leak. - RSS does not move → the memory is live or off-heap, and step 1 or step 3 was answered wrong. Go back. Do this once, on a canary, as an experiment. Leaving the call in on a schedule converts a diagnosis into a permanent tax: the arena you release is the arena the next rebuild will fault straight back in. ## Step 5 — fix the right thing If the verdict is retention, the lever is the **peak**, because the peak is what the runtime sizes for: - **Reuse the buffer.** Build the next snapshot into a `[]byte` you already own rather than allocating a fresh one each hour. The peak stops doubling and the retained arena stops being twice what the service actually needs. - **Shorten the overlap.** If the old snapshot only has to remain servable while the new one is being finalised rather than while it is being built, the window where both are fully live shrinks. - **Stream the build.** If the packed form can be produced incrementally into the destination buffer, the intermediate structures never coexist with two full snapshots. And if the verdict from step 1 was a leak, the usual culprits in this shape of service are: a package-level variable or cache still pointing at the old snapshot; a goroutine from the previous rebuild still running and holding it; and, characteristically for a packed `[]byte`, a small sub-slice handed out to a caller — a slice keeps its entire backing array alive, so one 40-byte record retained from a 500 MB snapshot pins all 500 MB. ## What to tell the person asking The useful sentence is not "it is fine" but "the live heap halves as expected, the runtime retains about N megabytes of free pages sized to the rebuild peak, and the way to make the graph match the profile is to stop doubling the peak." That names the mechanism, quantifies it, and points at a change someone can make.

  • Why does an hourly rebuild that keeps serving lookups roughly double the peak heap?
    Because the old snapshot has to stay reachable and servable until the new one is complete, so for the duration of the build both are live at once. The runtime sizes its heap for that moment. Everything after the swap is a shrink from a peak the process genuinely reached, which is why the steady-state live heap tells you nothing about the memory the runtime is holding.
  • runtime.MemStats shows HeapAlloc did not fall after the rebuild. Where do you look?
    At what still references the old snapshot. Typical causes: a package-level variable or cache never overwritten, a goroutine spawned by the previous rebuild still alive and holding it, or a small sub-slice of the packed `[]byte` handed to a caller — a slice pins its entire backing array, so one retained record keeps the whole snapshot alive. A heap profile grouped by allocation site plus the reference path usually names it in minutes.
  • How would you change the rebuild so the memory graph tracks the live heap more closely?
    Lower the peak rather than chase the steady state. Build into a buffer the service already owns instead of allocating a fresh snapshot each hour, shorten the window in which both snapshots are fully live, and produce the packed form incrementally where possible. If the peak stops doubling, the arena the runtime retains stops being twice the working set.
  • Why not just call debug.FreeOSMemory after every rebuild and close the ticket?
    Because you would release the exact pages the next rebuild needs an hour later, paying page faults and an unscheduled collection to reacquire them, in exchange for a nicer graph. It is reasonable as a one-shot experiment on a canary to prove the memory was free rather than leaked; it is a poor standing fix for a peak that repeats on a schedule.

saying these in an interview costs you the question

  • Declares a leak from the memory graph alone
  • Trusts the heap profile as a live view of the heap
  • Never reads HeapIdle or HeapReleased
  • Adds a periodic FreeOSMemory call as the fix
  • Optimises steady-state memory instead of the peak
  • Forgets a sub-slice keeps the whole backing array alive