skip to content

A Go queue worker's memory climbs for days while message throughput is flat — how do you tell a goroutine leak from a heap leak?

level: seniorimportance: should knowfreq 45%

answer

  1. two trends, not two numbers
  2. does it come back down overnight?
  3. the slope has a rate attached to it
  4. growth per hour divided by leaks per hour
  5. only then open the goroutine profile

basics

~10 s

Compare two trends: runtime.NumGoroutine and in-use heap bytes. A goroutine count that ratchets up and never falls back when the queue drains means goroutines are leaking, dragging their captured data with them.

solid answer

~50 s

Start with the cheapest discriminator: export `runtime.NumGoroutine()` as a gauge and look at it over days, not at one sample. A goroutine leak has a signature — the count rises in step with messages processed or with the error rate, and never comes back down during a quiet period. A heap leak shows flat goroutines and rising in-use bytes. If the count is climbing, take the goroutine profile from `runtime/pprof`, which groups live goroutines by where they are parked; a leak is one line with thousands of goroutines on it, and that line names the bug. Correlate the slope with a rate you already measure, because a leak rate tracking the error rate says the missing exit is on an error path. Go 1.27 also ships a `goroutineleak` profile reporting only goroutines the runtime can prove will never resume.

go deeper

for a junior

Know that Go exposes runtime.NumGoroutine and that a count which only ever rises is suspicious. Be able to say why leaked goroutines cost memory rather than CPU.

for a middle

Explain what to compare — goroutine count against in-use heap bytes — and why the trend matters more than any single sample. Know that the goroutine profile groups live goroutines by where they are parked.

for a senior

Show a triage order that rules things out cheaply before profiling, and use arithmetic to test the theory: growth per hour divided by leaks per hour should equal one in-flight message. Say what you would tie the slope to and what that implies about which code path is at fault.

for a principal

Decide what this service exports by default so the next incident starts with data: a goroutine gauge, in-use bytes, and a rate to correlate against. Own the call on when a climbing count should page someone versus open a ticket.

## Why this triage matters before any profiling "Memory grows, traffic is flat" is compatible with several very different bugs, and they have different fixes: an unbounded cache, a slice that keeps a large backing array alive, a connection pool that never trims, or leaked goroutines each holding the message they were decoding. Guessing wastes the on-call hour. One comparison separates them almost immediately. ## The discriminator: two trends, not two numbers Export both as gauges and look at a multi-day chart: - `runtime.NumGoroutine()` — the count of goroutines that currently exist, parked ones included. - heap in-use bytes — from `runtime.ReadMemStats` (`HeapAlloc`) or, preferably, `runtime/metrics`, which is the cheaper, non-stop-the-world source. A **single reading of `NumGoroutine` proves nothing**: legitimate in-flight work moves it, and a worker that runs one goroutine per message is *supposed* to spike. The signature of a leak is in the shape: - it **ratchets** — rises under load and does not return to its baseline when the queue drains overnight; - its slope is proportional to some rate you can name, usually messages consumed or errors returned; - restarting the process resets it to the baseline and the climb resumes at the same slope. Goroutines flat and memory rising points the other way — at a cache, a growing map, a retained buffer — and you go to a heap profile instead. Both rising together is the common case for a goroutine leak, because the goroutines are what is holding the memory. ## What the goroutine count is actually telling you about memory Each leaked goroutine contributes far more than its stack. The stack is small; the pinned data is not. A worker that pulls a JSON message off a queue, decodes it into a map, and then strands the goroutine holding that map is leaking the decoded document, not a few kilobytes of stack. So the arithmetic that confirms the theory is: (bytes of growth per hour) ÷ (goroutines leaked per hour) ≈ the size of one in-flight message plus its decoded form. If that number is plausible, you have your explanation without opening a heap profile at all. ## Then, and only then, the goroutine profile Once the trend says goroutines, the goroutine profile from `runtime/pprof` groups every live goroutine by its stack, so leaks announce themselves as one call site with an implausible count next to it. The distinction to keep in mind while reading it is between goroutines that are *supposed* to be parked (an accept loop, a ticker consumer, an idle pool worker) and goroutines parked at a place that should have completed in milliseconds. Go 1.27 added a `goroutineleak` profile in `runtime/pprof` that narrows this for you by reporting only goroutines the runtime can prove can never resume — the ones blocked on channels nothing can ever touch again. ## The traps in this triage - **Reading RSS alone.** Resident memory includes stacks, heap, and memory the runtime has not returned to the OS. It rises for all of these leaks and discriminates between none of them. - **Sampling `NumGoroutine` once during an incident.** The number is big; so what? Without the baseline and the slope you cannot tell a leak from a burst. - **Blaming the collector.** Go's collector is not generational and does not compact, but it does not skip garbage; if in-use bytes are rising, something is genuinely reachable, and a parked goroutine is one of the ways things stay reachable. - **Waiting for a deadlock error.** The runtime only aborts when *every* goroutine is asleep, which never happens while the consumer loop is still running. ## Closing the loop The fix is upstream of all this: every goroutine the worker starts needs an exit that does not depend on the happy path — a buffered result channel, a closed input channel on every producer return path, a cancellation case in the `select`. The value of the trend metric is that it turns "we think that is fixed" into evidence: after the deploy, the count should return to baseline during the quiet window and stay there. Keep the gauge on the dashboard permanently; it is one integer and it is the earliest warning this class of bug gives you.

  • The goroutine count is flat but in-use heap bytes climb steadily. What are you looking for instead?
    Retention rather than blocked goroutines: an unbounded cache or map that is never evicted from, a slice header keeping a large backing array alive, accumulated entries in a long-lived struct, or a growing set of open resources. A heap profile comparing in-use bytes between two snapshots hours apart points at the allocation site holding them.
  • Why is one sample of runtime.NumGoroutine a poor leak signal?
    Because legitimate concurrency moves it. A worker running a goroutine per in-flight message is expected to sit in the hundreds under load. What identifies a leak is that the count does not return to baseline when the work drains, and that its floor rises monotonically across days. Only the trend carries that information.
  • The leak's slope tracks your downstream error rate rather than your message rate. What does that tell you?
    That the missing exit path is on an error branch, not the common one. Something returns early — a producer that skips its close, a caller that abandons a result it started — only when the operation fails. It also explains why the leak appeared the day a dependency got flaky rather than the day the code shipped.
  • How do you verify the fix rather than assume it?
    Keep the goroutine gauge on the dashboard and watch a full quiet period after the deploy: the count must fall back to its baseline and stay there while the previous build's floor kept rising. In tests, compare the goroutine count before and after the exercised code, or use the goroutine-leak profile to assert nothing outlived the run.

saying these in an interview costs you the question

  • Diagnoses from resident memory alone, which rises for every leak
  • Takes one goroutine count reading and calls it high
  • Blames the garbage collector for reachable memory
  • Never checks whether the count falls back when traffic drains
  • Jumps straight to a heap profile without checking goroutine count