skip to content

A Go streaming service's goroutine profile total climbs for days while its heap stays flat. How do you confirm a leak and find the call site?

level: seniorimportance: should knowfreq 52%

answer

  1. one sample is not a trend
  2. two captures, minutes apart
  3. count the stacks, not the goroutines
  4. the growing stack traces to one call site
  5. a flat heap proves nothing here

basics

~20 s

Take two goroutine profiles minutes apart from the same process and diff them: a leak is the one stack whose count keeps growing. Then read the full dump and find the call site those goroutines are parked at. A flat heap proves nothing.

solid answer

~50 s

One snapshot only tells you how many goroutines exist; the difference between two tells you which stack is growing. Capture `/debug/pprof/goroutine` twice, minutes apart, and compare them with `go tool pprof -base first second` — a leak is a single stack whose count rises monotonically while the rest stay steady. Then pull `?debug=2` and read the blocks grouped by parked call site: thousands of identical stacks, all in `chan send` on the per-subscriber push, all carrying a long wait, all attributed by one `created by` line to the function that starts a writer per stream. That names the lifecycle to fix. The flat heap is not counter-evidence: a parked goroutine costs roughly its stack plus what it holds, so ten thousand of them are invisible next to a real heap. And capture both profiles **before** anyone restarts the process — a restart destroys exactly the evidence you need.

code

text · 7 lines
text
# two snapshots, five minutes apart, from the same process
curl -s http://localhost:6060/debug/pprof/goroutine > g1.pb.gz
sleep 300
curl -s http://localhost:6060/debug/pprof/goroutine > g2.pb.gz

# what grew in that window
go tool pprof -base g1.pb.gz g2.pb.gz

go deeper

for a junior

Know that a flat memory graph does not rule out a leak, and that a goroutine leak shows up in the goroutine profile rather than the heap profile. Being able to say which instrument to reach for is enough here.

for a middle

Explain why two captures beat one: a single dump reports how many goroutines exist, while the difference between two reports which stack is growing. Be ready to describe the diff mechanically.

for a senior

Walk the whole path aloud: capture twice, diff, group by parked call site, separate the legitimately long-lived, name the spawning function. Say explicitly what you capture before anyone is allowed to restart the process.

for a principal

Own the mitigation call. Decide when a scheduled restart is an honest stopgap with a deadline and a captured profile behind it, and when it has quietly become the reason nobody ever fixes the leak.

## Why this leak hides Picture a service that holds one long-lived server-streaming call open per subscriber: a goroutine per subscriber pushes protocol frames down its stream. When a subscriber drops, the publisher keeps handing frames to that subscriber's channel, and the writer goroutine — which will never run again because the connection is gone — sits parked forever on the send. One goroutine per dropped subscriber, accumulating for days. Nothing in a memory dashboard reacts. A parked goroutine's live cost is roughly its stack — starting at a couple of kilobytes and grown to whatever depth it reached — plus whatever state it still references. Ten thousand of them might be tens of megabytes, which disappears into the noise of a service with a multi-gigabyte heap. The damage arrives later and sideways: the per-subscriber channels and buffers those goroutines pin, the sockets and file descriptors they hold, scheduler and GC work proportional to goroutine count, and finally a slow slide into unresponsiveness that looks like nothing in particular. This is the failure mode that teaches people to look at the goroutine census and not just at memory. ## Step one: two snapshots, not one The single most common mistake is reading one dump, seeing twelve thousand goroutines, and declaring a leak. Plenty of healthy services legitimately hold tens of thousands — one per connection is a normal design in Go. A count is not evidence; **growth** is. So capture the profile twice, several minutes apart, from the same process: ``` curl -s http://localhost:6060/debug/pprof/goroutine > g1.pb.gz sleep 300 curl -s http://localhost:6060/debug/pprof/goroutine > g2.pb.gz go tool pprof -base g1.pb.gz g2.pb.gz ``` The `-base` flag subtracts the first from the second, so what you are looking at is *what appeared in those five minutes and did not go away*. In a healthy process the diff is close to noise: handlers come and go and roughly balance. In a leaking one, a single stack dominates the difference. If you prefer text, the same story is visible in two `?debug=1` captures — that view already groups goroutines by identical stack with a count in front, so you can subtract by eye. Do the capture more than twice if you can. A leak has a slope; knowing the slope tells you how much runway you have and, later, whether your fix actually worked. ## Step two: read the population at the parked call site Once the diff names a stack, pull the full dump with `?debug=2` and look at the blocks that share it. What you want out of them: - **The wait reason.** `chan send` says a value has nowhere to go: the receiver is gone or is not draining. `chan receive` says nothing is arriving and nobody closed the channel. `select` says every case is dead at once — often a `case <-ctx.Done()` on a context that is never cancelled. - **The wait duration.** Blocks parked longer than about a minute carry a minutes figure. In a service where requests finish in milliseconds, that is your filter. - **The innermost frame.** It names the exact line that parked — which channel, which lock, which read. - **The `created by` line.** It names the function that ran the `go` statement. Thousands of identical stacks all attributed to `Subscribe` turns "goroutines are leaking" into "the writer we start per subscriber is never being shut down", which is a statement someone can fix. ## Step three: separate the leaked from the merely long-lived A dump is full of goroutines that have been parked for hours and are perfectly healthy: accept loops, background tickers, idle workers waiting on a job channel, connection handlers in `IO wait`. None of these are leaks. Three things distinguish a leak from a long life: 1. **It grows.** That is what the diff established, and it is the strongest signal you have. 2. **It concentrates.** A leak piles up on one stack; structural goroutines exist in small fixed numbers. 3. **It cannot be woken.** Ask what event would resume this goroutine and whether that event can still happen. A send to a channel whose only reader has returned will never complete, no matter how long you wait. Duration alone convicts nobody, and that is the trap this scenario sets. ## The mitigation call Someone will propose a nightly restart until the fix ships. It works — the count resets and the service survives another day — and it is a legitimate stopgap, but on two conditions. First, **capture the evidence before the restart**: both profiles and a full dump, saved somewhere that outlives the process, because a restart destroys the only copy of the state you need. Second, attach a deadline and a metric. Track the profile total's slope so you know how much runway a restart actually buys and whether it is shrinking; alert on the slope rather than on the restart. A restart with a captured profile and a dated ticket behind it is honest mitigation. A restart that quietly becomes the operating procedure is how a leak lives for a year.

  • Why does a per-subscriber goroutine leak stay invisible on memory dashboards for days?
    Because a parked goroutine costs little: its stack, which starts at a couple of kilobytes, plus whatever state it still references. Ten thousand leaked goroutines might amount to tens of megabytes, which is noise beside a multi-gigabyte heap. The visible damage arrives through the resources they pin — channels, buffers, sockets — and through scheduler and GC work that scales with goroutine count, long after the leak began.
  • How do you tell a leaked goroutine from a legitimately long-lived one in the same dump?
    Three tests. It grows across two captures taken minutes apart, where a structural goroutine's count is fixed. It concentrates on a single stack, where healthy long-lived goroutines exist in small numbers spread across several. And nothing can wake it: ask what event would resume it and whether that event can still occur. A long wait on its own convicts nobody — an idle worker looks identical.
  • The on-call engineer wants a nightly restart until the fix ships. What do you insist on first?
    Capture both profiles and a full dump and store them outside the process before the restart, because the restart destroys the only evidence. Then attach a deadline and a measurement: track how fast the goroutine total climbs so you know what runway a restart buys and whether it is shrinking. A restart with captured evidence and a dated ticket is mitigation; one without either becomes the permanent workaround.
  • The diff shows a growing stack, but the innermost frame is inside the standard library. What now?
    Read further down the same block. The frames below the standard-library call are yours, and the `created by` line at the bottom names the function that spawned the goroutine. A park inside a channel or network primitive is expected; the interesting question is which of your call sites got it there and what was supposed to release it.

A single goroutine dump is a photograph of a crowd: you cannot tell who is waiting for a train and who has been standing there for a week. Two photographs minutes apart tell you instantly, because only one group never moves and keeps getting bigger.

saying these in an interview costs you the question

  • Concludes there is no leak because heap and RSS look flat
  • Reads one profile snapshot and calls it a trend
  • Assumes any goroutine with a long wait is leaked
  • Restarts the process before capturing any profile
  • Hunts a goroutine leak in the heap profile
  • Treats a high goroutine count as a defect by itself