What is a goroutine leak in Go, and why does the runtime never reclaim a goroutine that is blocked forever?
answer
- started, never returns
- nothing outside can stop it
- the collector traces from it, not to it
- its stack pins whatever it captured
- count only ever goes up
basics
~20 sA goroutine leak is a goroutine that never returns because it is blocked forever on an operation nobody will ever complete. The runtime cannot reclaim it, so its stack and everything it references stay alive until the process exits.
solid answer
~50 sA goroutine ends exactly one way: its function returns (or it panics, or it calls `runtime.Goexit`). Nothing outside can stop it. A leak is a goroutine parked forever on an operation that will never complete — a send with no receiver, a receive with no sender and no `close`, a `range` over a channel nobody closes, a `WaitGroup.Wait` whose counter never hits zero. Goroutines are garbage-collection roots, so the collector traces *from* a blocked goroutine's stack and never decides the goroutine itself is garbage; even a channel that is unreachable from the rest of the program will not free it. The cost is its stack plus everything reachable from its frames — the captured request, the decoded message, the buffer. So leaks show up as memory that climbs forever and a `runtime.NumGoroutine()` that only goes up.
code
go · 8 linesfunc leaks(msgs chan []byte) {
go func() { <-make(chan struct{}) }() // nobody sends, nobody closes
go func() { msgs <- []byte("x") }() // nobody receives
go func() {
for range msgs { // ends only when msgs is closed
}
}()
}go deeper
Be able to say what a goroutine leak is in one sentence and name two shapes that cause it — a send with no receiver, a receive on a channel nobody closes. Know that nothing outside a goroutine can stop it.
Explain why the collector cannot reclaim a parked goroutine: goroutines are GC roots, so reachability is traced from them. Be ready to say what a leak actually costs and why it looks like a memory leak rather than a CPU problem.
Show that you treat an exit path as part of the contract of starting a goroutine, and that you can spot the failure branch where the exit disappears. Expect to be asked how you would notice the leak in production before it pages you.
Frame this as a codebase-wide invariant rather than a bug class: every goroutine start needs a stated termination condition, and reviews should reject the ones that only terminate in the happy case. Decide what leak signal the platform exports by default.
## What a goroutine leak is A goroutine is a function running concurrently, started with `go f()`. It terminates in exactly three ways: the function returns, it panics and the process dies, or it calls `runtime.Goexit`. There is no external stop and no timeout. A **goroutine leak** is a goroutine that can never reach any of those endings, because it is parked forever on an operation no other goroutine will ever complete. The usual shapes: - a **send** on a channel whose receiver has gone away (`ch <- v` with nobody left to receive); - a **receive** on a channel that nobody will send to and nobody will `close`; - a `for range` over a channel that is never closed — `range` over a channel ends *only* on `close`; - `sync.WaitGroup.Wait` whose counter never returns to zero because one worker returned early without its `Done`; - a `select` where every case is a channel that is now dead and there is no cancellation case and no `default`; - a lock nobody unlocks. ## Why the runtime cannot clean it up Every live goroutine is a **root** for the garbage collector. The GC starts from each goroutine's stack and traces outwards to decide what is reachable; it never reasons in the other direction and concludes that a goroutine is unreachable. So a goroutine blocked on a channel is kept alive even when that channel is referenced by nothing else in the program. The Go runtime deliberately does not try to prove "this goroutine can never run again" during collection — that is a whole-program liveness question, not a reachability one. There is one check that looks like a safety net and is not. If **every** goroutine in the process is asleep, the runtime aborts with `fatal error: all goroutines are asleep - deadlock!`. A leak inside a running service never triggers it: the accept loop, timer goroutines and background GC workers are all runnable, so the process is not deadlocked — it just has one more parked goroutine than it had a minute ago. Leaks are silent by construction. ## What each leaked goroutine costs Three things: 1. **Its stack.** Goroutine stacks are small and grow on demand, so this alone is cheap — thousands of leaked goroutines are only megabytes of stack. 2. **Everything reachable from its frames.** This is usually the expensive part: the message it decoded, the HTTP request it captured, the `[]byte` buffer it was about to write. Those objects are pinned as long as the goroutine is parked, which is forever. This is why a goroutine leak is normally *observed* as a memory leak. 3. **GC work.** The collector must scan every goroutine's stack each cycle, so mark time grows with the goroutine count even though parked goroutines are never scheduled. They cost essentially no CPU otherwise: a parked goroutine is off the run queues entirely, which is exactly why the service still looks healthy on a latency dashboard while it slowly eats the box. ## How you see one `runtime.NumGoroutine()` returns the current count. Exported as a gauge and watched over hours, a leak is unmistakable: the count ratchets up with load and never comes back down when load drops. A single reading tells you almost nothing, because in-flight work legitimately moves the number. The goroutine profile from `runtime/pprof` then groups live goroutines by where they are parked, so a thousand goroutines sitting on the same channel send line is the answer. ## The discipline that prevents them Every goroutine you start needs a guaranteed exit path that does not depend on the happy case happening. Concretely: give a result channel enough buffer that the sender never blocks even if the reader walked away; make sure whoever owns the sending side closes the channel on every return path, including the error path; and give any goroutine that can block a cancellation case to select on. If you cannot say out loud what makes a goroutine return in the failure case, it leaks in the failure case.
- If the channel a goroutine is blocked on becomes unreachable from the rest of the program, will the garbage collector free that goroutine?No. The goroutine is itself a GC root, so the collector traces outward from its stack and keeps the channel alive rather than the other way round. Go does not attempt to prove that a parked goroutine can never resume during collection. It stays parked, holding its stack and everything it captured, until the process exits.
- Why doesn't the runtime's "all goroutines are asleep - deadlock!" error catch leaks in a real server?That check fires only when the runtime sees that *every* goroutine is blocked, so no progress is possible anywhere. A live server always has runnable goroutines — the listener, timers, background collection work — so the condition is never met. The leaked goroutines are simply parked while the rest of the process keeps serving.
- Roughly what does one leaked goroutine cost?Its stack is small — kilobytes, growing only if it needed a deep call chain. The real cost is everything reachable from its frames: the captured request, the decoded payload, the buffer it was about to send. That can be orders of magnitude larger than the stack, which is why goroutine leaks surface as memory growth rather than CPU growth.
It is a worker who was sent to wait at a loading dock for a truck that got cancelled. Nobody tells him, he never clocks out, and he keeps holding the paperwork he was given.
saying these in an interview costs you the question
- Says the garbage collector eventually cleans up blocked goroutines
- Believes a goroutine can be killed from outside once it is running
- Thinks the runtime's deadlock detector catches leaks in a live server
- Assumes a leaked goroutine burns CPU while it waits
- Claims a leaked goroutine only wastes its own small stack