skip to content

Why do a Go fan-in merge's forwarding goroutines leak when the consumer abandons the merged channel?

level: seniorimportance: should knowfreq 52%

answer

  1. who receives after the consumer leaves
  2. a parked sender never wakes
  3. the collector does not reclaim goroutines
  4. count goroutines against live feeds
  5. put the send inside a select

basics

~20 s

Each forwarder is blocked sending into the merged channel. Nothing will ever receive again, so those goroutines park forever, keep their last value and their source alive, and the merged channel never closes. Give every forwarder a cancellation case.

solid answer

~50 s

A forwarder's loop is `for v := range src { out <- v }`, and that send only completes when someone receives. If the consumer returns early — an error, a timeout, a `break` out of the range loop — nobody ever receives again, so every forwarder parks on its send permanently. Parked goroutines are roots: they are never collected, they pin their stack, the value in flight and the source channel, and because they never return, the closer goroutine stays in `Wait` too. One abandoned merge over five feeds leaks six goroutines, and a service that re-creates it on every reconnect leaks steadily. The fix is to make the send abandonable: pass a `context.Context` into the merge, have each forwarder `select` on `out <- v` and `<-ctx.Done()`, and have the consumer `defer cancel()` so every exit path releases the merge.

code

go · 20 lines
go
func merge[T any](ctx context.Context, srcs ...<-chan T) <-chan T {
	out := make(chan T)
	var wg sync.WaitGroup
	for _, src := range srcs {
		wg.Go(func() {
			for v := range src {
				select {
				case out <- v:
				case <-ctx.Done():
					return
				}
			}
		})
	}
	go func() {
		wg.Wait()
		close(out)
	}()
	return out
}

go deeper

for a junior

Remember that a goroutine blocked on a channel send stays alive forever if nobody ever receives. Abandoning a channel does not clean up the goroutines feeding it.

for a middle

Explain the mechanics: the send is a rendezvous with no receiver, the goroutine is a GC root that pins its stack and its source, and the closer goroutine stuck in Wait leaks too.

for a senior

Demonstrate the operational loop: a goroutine gauge climbing in steps, the goroutine profile grouped by stack showing chan send at one line, then a Context threaded through the merge with defer cancel at the consumer.

for a principal

Make it a boundary rule others follow: any goroutine sending into a channel it does not own needs a stop that does not depend on the receiver, and the exported merge signature should make the cancellation obligation visible at every call site.

## The scenario A market-data service merges quote streams from several providers into one channel and a single consumer loop reads it. One day the consumer stops early — a downstream write fails, a deadline elapses, or someone adds a `break` when a bad quote arrives — and returns without draining the merged channel. Nothing crashes. Nothing logs. The service keeps serving. But the goroutine count starts climbing, and a week later the process is holding gigabytes of goroutine stacks and blocked feeds. ## Why a blocked forwarder is permanent The forwarder body is `for v := range src { out <- v }`. That send is a rendezvous: it completes only when a receive executes on `out`. Once the consumer is gone there will never be another receive, so the goroutine is parked in a channel send with no possible wakeup. Three consequences follow, and the second is the one interviewers listen for: 1. **The goroutine is never collected.** Go's garbage collector reclaims unreachable *memory*; it has no notion of an unreachable goroutine. A goroutine blocked forever is a live goroutine forever, and it is itself a GC root — its stack keeps the value in flight, the source channel and anything those reference alive. 2. **Blocking propagates upstream.** The provider goroutine feeding that source now has no reader either, so it parks on its own send in turn. One abandoned consumer can stall an entire chain. 3. **The merged channel is never closed.** The forwarders never return, so the `WaitGroup` counter never reaches zero and the closer goroutine sits in `Wait` forever. That is a seventh leaked goroutine on a five-source merge, and it also means any other code waiting for the merged channel to close waits forever. ## Diagnosing it The symptom is a monotonically rising goroutine count that never falls back after load subsides. `runtime.NumGoroutine()` exported as a gauge is enough to see it; a step of exactly *sources + 1* per incident is a strong hint that a merge is the culprit. Compare that count against the number of live provider connections — if goroutines grow while the feed count stays flat, something per-merge is not exiting. The goroutine profile then names it precisely, because it groups goroutines by identical stack and shows the blocking state: ``` goroutine 217 [chan send]: main.merge[...].func1() /app/feed/merge.go:14 +0x8c ``` A large count of goroutines in the `chan send` state, all at the same line inside the merge, is the whole diagnosis. This is not something the race detector will show you: nothing here is a data race, it is a liveness bug, and `-race` reports neither. ## The fix: make the send abandonable A forwarder must be able to give up, which means the send must appear in a `select` alongside a cancellation signal: ```go for v := range src { select { case out <- v: case <-ctx.Done(): return } } ``` The merge takes a `context.Context` as its first parameter, and the consumer owns cancelling it — `ctx, cancel := context.WithCancel(parent)` followed by `defer cancel()`, so that every exit path from the consumer, including a panic unwinding through it, releases the merge. When `cancel` runs, `ctx.Done()` closes, every parked forwarder wakes on that case and returns, the `WaitGroup` drains, and the closer goroutine closes the merged channel and exits. The whole merge unwinds within one scheduling round. Note what this does **not** promise: the value the forwarder was carrying is dropped, so if losing an in-flight quote matters you need acknowledgement at a higher level. Cancellation is abandonment, and that has to be an explicit product decision, not an accident. ## The alternative fix, and its limit Instead of cancelling you can *always drain*: `go func() { for range merged {} }()` on the abandonment path. Every forwarder eventually completes its send, its source runs dry, and the merge closes itself. This is only correct if every source is guaranteed to close on its own. For a live feed that never ends it just moves the leak — now the drainer goroutine runs forever. Draining is right for a bounded batch; cancellation is right for a stream. ## Making it a habit The rule worth stating in review is: **any goroutine that sends into a channel it does not own must have a way to stop that does not depend on the receiver.** In practice that is a `ctx.Done()` case on every send in a forwarder, a `defer cancel()` at every consumer, and a documented statement on the merge's exported signature about who cancels. Once the merge takes a `Context` as its first parameter, the obligation is visible at every call site instead of living in a comment.

  • Once the forwarders return on ctx.Done(), is everything released?
    The merge's own goroutines are, but not necessarily the producers above it. A provider goroutine still sending into a source channel now has no reader and parks in turn. Cancellation has to reach the producers as well — usually the same context — and the merge should document that it stops reading its sources but does not close them.
  • Is draining the merged channel instead of cancelling a valid fix?
    Only when every source is guaranteed to close. Running `for range merged {}` on a goroutine lets each forwarder finish, the WaitGroup drain and the closer run, which is fine for a bounded batch. For an endless feed the drainer itself never returns, so you have swapped several leaked goroutines for one — cancellation is the right tool for a stream.
  • What tells you at 3am that this is what is happening?
    A goroutine count that climbs and never falls, stepping up by roughly the number of sources per incident while the number of live provider connections stays flat. The goroutine profile confirms it: a large group of goroutines in the `chan send` state, all at the same line in the merge. The race detector shows nothing, because this is liveness, not a data race.
  • Why does the process not simply crash or report the deadlock?
    Go's deadlock detector only fires when every goroutine in the process is blocked, and a real service always has a listener or a ticker runnable. The leaked forwarders are a small parked minority, so the runtime sees progress and stays silent. Nothing surfaces it except memory growth and your own goroutine metric.

saying these in an interview costs you the question

  • Says the garbage collector will clean up a blocked goroutine
  • Assumes an unread channel just drops the values sent to it
  • Buffers the merged channel and calls the leak fixed
  • Expects the runtime to report a deadlock panic
  • Reaches for the race detector to find a blocked sender