A Go fan-in merge calls wg.Wait() and close(out) inline before returning out — why does it hang?
answer
- who receives while the function waits
- an unbuffered send is a rendezvous
- the caller has no channel yet
- constructor, not a drain
- empty sources make it pass
basics
~20 sBecause the forwarders are blocked sending. With an unbuffered output channel nothing can receive until the caller holds the channel, so every forwarder parks on its send, the WaitGroup counter never reaches zero, and Wait never returns.
solid answer
~50 sA merge function is a constructor, not a drain. Its forwarders and its consumer have to be alive at the same time: on an unbuffered `out`, a forwarder's first send blocks until somebody receives, and nobody can receive before the merge function has returned the channel. Waiting inline therefore deadlocks — the forwarders wait for a consumer, and the merge waits for the forwarders. That is why the wait and the `close(out)` are moved into their own goroutine and the channel is returned immediately. Buffering `out` does not really fix it either: it only works if the buffer is at least as large as everything the sources will ever send, which you cannot know. The bug is also data-dependent, so a test where every source is already closed and empty passes happily and hides it.
code
go · 14 linesfunc mergeBroken[T any](srcs ...<-chan T) <-chan T {
out := make(chan T)
var wg sync.WaitGroup
for _, src := range srcs {
wg.Go(func() {
for v := range src {
out <- v // parks: the caller has no channel yet
}
})
}
wg.Wait() // never returns while any value is pending
close(out)
return out
}go deeper
Remember the rule of thumb: a function that starts senders must hand the channel back before anything can be received. Waiting for your own producers inside that function is a deadlock.
Trace the cycle out loud — forwarder parks on an unbuffered send, so Wait never returns, so the channel is never returned, so nothing ever receives — and give the fix: wait and close in a separate goroutine.
Point out that the bug is data-dependent and slips past a test over empty sources, and that a real service will not panic on it. Say how you would test it: real values plus a bounded deadline.
Frame it as an API rule the team can apply everywhere: functions that return a channel return it immediately, and any waiting they do happens off the caller's path. Encode that in review guidance rather than rediscovering it.
## The broken code ```go func mergeBroken[T any](srcs ...<-chan T) <-chan T { out := make(chan T) var wg sync.WaitGroup for _, src := range srcs { wg.Go(func() { for v := range src { out <- v } }) } wg.Wait() close(out) return out } ``` It reads as if it is being careful — wait for everyone, close cleanly, hand back a finished channel. It hangs the moment any source actually carries a value. ## Why it hangs An unbuffered channel send does not complete until a receive is executing on the other side; the two operations are a rendezvous. Here is the cycle: 1. A forwarder receives a quote from its source and executes `out <- v`. There is no receiver, so the goroutine parks. 2. Because it is parked, it never returns, so the `WaitGroup` entry for it is never released. 3. `wg.Wait()` in the merge function therefore never returns. 4. Because `mergeBroken` never returns, the caller never gets `out`, and so no receive is ever executed. Each party is waiting on the other. Note that `go vet` will not flag this and the compiler is perfectly happy: it is a legal program with a runtime dependency cycle. It will not even produce the friendly `all goroutines are asleep - deadlock!` panic in a real service, because that check only fires when **every** goroutine in the process is blocked. In a server with a listener and a metrics ticker still running, the merge just silently wedges the caller. ## Producer and consumer must overlap in time The underlying principle is broader than merging: with unbuffered channels you cannot sequence "produce everything, then consume everything" inside one goroutine. Producing and consuming are concurrent by construction. A function that both starts producers and waits for them must not be the same function that the consumer is waiting on. So the merge function's job is only to *set things up*: create the output, start one forwarder per source, start a closer, return. Everything that waits happens in a goroutine the caller is not blocked on. ```go go func() { wg.Wait() close(out) }() return out ``` ## Why buffering is not the fix Candidates often propose `make(chan T, 16)`. That changes when the deadlock happens, not whether it can. Sends complete without a receiver only while the buffer has room; the seventeenth value blocks and you are back in the same cycle. The buffered version is arguably worse, because it now deadlocks depending on how much data the sources happen to carry — a bug that appears in production and not in the small test. The only buffer that removes the hang is one at least as large as the total number of values that will ever flow, which for a live feed is unbounded. And even if you could size it, you would have destroyed backpressure: the merge would absorb everything into memory instead of letting a slow consumer slow the producers down. ## Why the test suite misses it Run `mergeBroken` over sources that are already closed and empty and it behaves perfectly: each forwarder's `range` ends immediately without ever sending, the counter drops to zero, `Wait` returns, `close(out)` runs, and the caller gets a closed empty channel. A unit test that merges two empty channels and asserts "no values, channel closed" goes green. The failure needs a value in flight, which is why it tends to be found by the engineer wiring up the first real feed rather than by CI. It is worth writing the test with actual values in every source and a bounded overall deadline, so that a hang is reported as a failure instead of a stuck run. ## The diagnosis in the moment If you are staring at a wedged process, a goroutine dump is decisive: you will see N goroutines parked in a channel send inside the forwarder, plus one goroutine parked in `sync.WaitGroup.Wait` in the merge function itself, and the caller blocked on the merge call. That triangle — senders waiting, a `Wait` waiting, and a caller waiting on the constructor — is the signature of this exact mistake, and it is worth being able to name on sight.
- Would giving the merged channel a buffer fix it?Only until the buffer fills. A buffer of sixteen deadlocks on the seventeenth value, so the bug becomes data-dependent rather than fixed. The only size that always works is one that holds everything the sources will ever send, which is unbounded for a live feed — and it would remove backpressure onto the producers. Close asynchronously instead.
- A unit test over two already-closed empty channels passes against this broken merge. Why?With no values to forward, each forwarder's range loop ends immediately without executing a single send, so the WaitGroup counter reaches zero, `Wait` returns, the close runs and the caller receives a closed empty channel. The deadlock needs a value in flight, so the test must send real values and enforce a deadline to catch it.
- Why does the process not report Go's 'all goroutines are asleep - deadlock!' panic?That detection fires only when every goroutine in the process is blocked. A real service still has an HTTP listener, a ticker or a signal handler runnable, so the runtime sees progress is possible and stays quiet. The merge just wedges its caller, and you find it in a goroutine dump rather than as a crash.
saying these in an interview costs you the question
- Blames the WaitGroup counter rather than the blocked sends
- Says a small buffer on the output channel fixes it
- Thinks close(out) would release the parked senders
- Assumes the forwarders complete before the function returns
- Expects the runtime to panic with a deadlock message