A Go streaming service's goroutine profile total climbs for days while its heap stays flat. How do you confirm a leak and find the call site?
answer
- one sample is not a trend
- two captures, minutes apart
- count the stacks, not the goroutines
- the growing stack traces to one call site
- a flat heap proves nothing here
basics
~20 sTake two goroutine profiles minutes apart from the same process and diff them: a leak is the one stack whose count keeps growing. Then read the full dump and find the call site those goroutines are parked at. A flat heap proves nothing.
solid answer
~50 sOne snapshot only tells you how many goroutines exist; the difference between two tells you which stack is growing. Capture `/debug/pprof/goroutine` twice, minutes apart, and compare them with `go tool pprof -base first second` — a leak is a single stack whose count rises monotonically while the rest stay steady. Then pull `?debug=2` and read the blocks grouped by parked call site: thousands of identical stacks, all in `chan send` on the per-subscriber push, all carrying a long wait, all attributed by one `created by` line to the function that starts a writer per stream. That names the lifecycle to fix. The flat heap is not counter-evidence: a parked goroutine costs roughly its stack plus what it holds, so ten thousand of them are invisible next to a real heap. And capture both profiles **before** anyone restarts the process — a restart destroys exactly the evidence you need.
code
text · 7 lines# two snapshots, five minutes apart, from the same process
curl -s http://localhost:6060/debug/pprof/goroutine > g1.pb.gz
sleep 300
curl -s http://localhost:6060/debug/pprof/goroutine > g2.pb.gz
# what grew in that window
go tool pprof -base g1.pb.gz g2.pb.gzgo deeper
Know that a flat memory graph does not rule out a leak, and that a goroutine leak shows up in the goroutine profile rather than the heap profile. Being able to say which instrument to reach for is enough here.
Explain why two captures beat one: a single dump reports how many goroutines exist, while the difference between two reports which stack is growing. Be ready to describe the diff mechanically.
Walk the whole path aloud: capture twice, diff, group by parked call site, separate the legitimately long-lived, name the spawning function. Say explicitly what you capture before anyone is allowed to restart the process.
Own the mitigation call. Decide when a scheduled restart is an honest stopgap with a deadline and a captured profile behind it, and when it has quietly become the reason nobody ever fixes the leak.
## Why this leak hides Picture a service that holds one long-lived server-streaming call open per subscriber: a goroutine per subscriber pushes protocol frames down its stream. When a subscriber drops, the publisher keeps handing frames to that subscriber's channel, and the writer goroutine — which will never run again because the connection is gone — sits parked forever on the send. One goroutine per dropped subscriber, accumulating for days. Nothing in a memory dashboard reacts. A parked goroutine's live cost is roughly its stack — starting at a couple of kilobytes and grown to whatever depth it reached — plus whatever state it still references. Ten thousand of them might be tens of megabytes, which disappears into the noise of a service with a multi-gigabyte heap. The damage arrives later and sideways: the per-subscriber channels and buffers those goroutines pin, the sockets and file descriptors they hold, scheduler and GC work proportional to goroutine count, and finally a slow slide into unresponsiveness that looks like nothing in particular. This is the failure mode that teaches people to look at the goroutine census and not just at memory. ## Step one: two snapshots, not one The single most common mistake is reading one dump, seeing twelve thousand goroutines, and declaring a leak. Plenty of healthy services legitimately hold tens of thousands — one per connection is a normal design in Go. A count is not evidence; **growth** is. So capture the profile twice, several minutes apart, from the same process: ``` curl -s http://localhost:6060/debug/pprof/goroutine > g1.pb.gz sleep 300 curl -s http://localhost:6060/debug/pprof/goroutine > g2.pb.gz go tool pprof -base g1.pb.gz g2.pb.gz ``` The `-base` flag subtracts the first from the second, so what you are looking at is *what appeared in those five minutes and did not go away*. In a healthy process the diff is close to noise: handlers come and go and roughly balance. In a leaking one, a single stack dominates the difference. If you prefer text, the same story is visible in two `?debug=1` captures — that view already groups goroutines by identical stack with a count in front, so you can subtract by eye. Do the capture more than twice if you can. A leak has a slope; knowing the slope tells you how much runway you have and, later, whether your fix actually worked. ## Step two: read the population at the parked call site Once the diff names a stack, pull the full dump with `?debug=2` and look at the blocks that share it. What you want out of them: - **The wait reason.** `chan send` says a value has nowhere to go: the receiver is gone or is not draining. `chan receive` says nothing is arriving and nobody closed the channel. `select` says every case is dead at once — often a `case <-ctx.Done()` on a context that is never cancelled. - **The wait duration.** Blocks parked longer than about a minute carry a minutes figure. In a service where requests finish in milliseconds, that is your filter. - **The innermost frame.** It names the exact line that parked — which channel, which lock, which read. - **The `created by` line.** It names the function that ran the `go` statement. Thousands of identical stacks all attributed to `Subscribe` turns "goroutines are leaking" into "the writer we start per subscriber is never being shut down", which is a statement someone can fix. ## Step three: separate the leaked from the merely long-lived A dump is full of goroutines that have been parked for hours and are perfectly healthy: accept loops, background tickers, idle workers waiting on a job channel, connection handlers in `IO wait`. None of these are leaks. Three things distinguish a leak from a long life: 1. **It grows.** That is what the diff established, and it is the strongest signal you have. 2. **It concentrates.** A leak piles up on one stack; structural goroutines exist in small fixed numbers. 3. **It cannot be woken.** Ask what event would resume this goroutine and whether that event can still happen. A send to a channel whose only reader has returned will never complete, no matter how long you wait. Duration alone convicts nobody, and that is the trap this scenario sets. ## The mitigation call Someone will propose a nightly restart until the fix ships. It works — the count resets and the service survives another day — and it is a legitimate stopgap, but on two conditions. First, **capture the evidence before the restart**: both profiles and a full dump, saved somewhere that outlives the process, because a restart destroys the only copy of the state you need. Second, attach a deadline and a metric. Track the profile total's slope so you know how much runway a restart actually buys and whether it is shrinking; alert on the slope rather than on the restart. A restart with a captured profile and a dated ticket behind it is honest mitigation. A restart that quietly becomes the operating procedure is how a leak lives for a year.
- Why does a per-subscriber goroutine leak stay invisible on memory dashboards for days?Because a parked goroutine costs little: its stack, which starts at a couple of kilobytes, plus whatever state it still references. Ten thousand leaked goroutines might amount to tens of megabytes, which is noise beside a multi-gigabyte heap. The visible damage arrives through the resources they pin — channels, buffers, sockets — and through scheduler and GC work that scales with goroutine count, long after the leak began.
- How do you tell a leaked goroutine from a legitimately long-lived one in the same dump?Three tests. It grows across two captures taken minutes apart, where a structural goroutine's count is fixed. It concentrates on a single stack, where healthy long-lived goroutines exist in small numbers spread across several. And nothing can wake it: ask what event would resume it and whether that event can still occur. A long wait on its own convicts nobody — an idle worker looks identical.
- The on-call engineer wants a nightly restart until the fix ships. What do you insist on first?Capture both profiles and a full dump and store them outside the process before the restart, because the restart destroys the only evidence. Then attach a deadline and a measurement: track how fast the goroutine total climbs so you know what runway a restart buys and whether it is shrinking. A restart with captured evidence and a dated ticket is mitigation; one without either becomes the permanent workaround.
- The diff shows a growing stack, but the innermost frame is inside the standard library. What now?Read further down the same block. The frames below the standard-library call are yours, and the `created by` line at the bottom names the function that spawned the goroutine. A park inside a channel or network primitive is expected; the interesting question is which of your call sites got it there and what was supposed to release it.
A single goroutine dump is a photograph of a crowd: you cannot tell who is waiting for a train and who has been standing there for a week. Two photographs minutes apart tell you instantly, because only one group never moves and keeps getting bigger.
saying these in an interview costs you the question
- Concludes there is no leak because heap and RSS look flat
- Reads one profile snapshot and calls it a trend
- Assumes any goroutine with a long wait is leaked
- Restarts the process before capturing any profile
- Hunts a goroutine leak in the heap profile
- Treats a high goroutine count as a defect by itself