In a first-result-wins fan-out over N mirrors, why must the result channel have one buffer slot per attempt?
answer
- one receive, many senders
- count who is left holding a value
- an unbuffered send needs a partner
- the loser has nowhere to put it
- capacity equals the number of attempts
basics
~20 sBecause only one value is ever received. With an unbuffered channel the losing goroutines block forever on their send, leaking themselves and everything they hold. One slot per attempt lets every sender complete and exit even though nobody is listening.
solid answer
~50 sThe caller receives once and returns, so after the winner is taken there is no receiver left. A send on an unbuffered channel only completes when a receiver is ready to take it, so each losing goroutine parks in `ch <- result{...}` for the life of the process — a permanent goroutine leak, and it pins the goroutine's stack plus the downloaded bytes and the connection it still references. Sizing the channel `make(chan result, len(mirrors))` gives every attempt somewhere to drop its value unconditionally, so the send always completes and the goroutine returns. The alternative is a non-blocking send — a `select` with the send in one case and a `default`, or a `ctx.Done()` case — but the buffer is simpler and cannot drop a value. Confirmation is easy: the goroutine profile shows N-1 goroutines parked in `chan send` per call.
code
go · 12 lines// leaks N-1 goroutines per call: after the single receive,
// no receiver is left and each remaining send blocks forever
bad := make(chan result)
// every attempt has a reserved slot, so its send always
// completes and the goroutine returns even with no receiver
good := make(chan result, len(mirrors))
for _, m := range mirrors {
go func() { good <- fetchFrom(ctx, m) }()
}
r := <-good // one receive; the other values are simply abandonedgo deeper
Remember the rule and the reason: a send on an unbuffered channel needs someone receiving at that moment, and after a race only one receive ever happens. Give the channel one slot per attempt.
Explain the rendezvous semantics of an unbuffered send, why the abandoned senders park forever, and what capacity guarantees instead. Be able to say what the leaked goroutine keeps alive besides itself.
Diagnose it from a live process: goroutine count growing with request count, then a goroutine profile full of identical chan send stacks. Know the alternatives — non-blocking send, cancellation-aware send — and why plain capacity is the default.
Treat it as a review invariant rather than a bug you catch twice: for every fan-out, senders must be able to complete without a receiver, and that property should be visible at the make call, not inferred from the control flow around it.
## The failure this rule exists to prevent Start three goroutines to fetch the same archive from three mirrors, have each send its outcome to a shared channel, and receive once: ```go ch := make(chan result) // unbuffered — the bug for _, m := range mirrors { go func() { ch <- fetchFrom(ctx, m) }() } r := <-ch return r.body, r.err ``` This returns the right answer and leaks two goroutines on **every call**. It is the classic first-result-wins bug because the program is correct in its output and wrong in its resource behaviour, so tests pass and a long-running process dies days later. ## Why the losers block A send on an unbuffered channel is a rendezvous: it completes only when a receiver is simultaneously ready to take the value. The caller performed exactly one receive and has returned; nobody will ever receive again. So the second and third goroutines reach `ch <- ...` and park there, permanently. Nothing rescues them — a blocked goroutine is not a leaked object the collector can clean up. It is a live goroutine, it is a GC root, and it keeps alive: - its own stack, - the `[]byte` archive it just finished downloading (potentially megabytes), - the HTTP response and the underlying connection it still references, - whatever the closure captured. Call the download path once per request in a service and you are leaking two goroutines and two archives per request. The process grows until it is killed. ## Why capacity fixes it `make(chan result, len(mirrors))` gives the channel as many slots as there are senders. A send on a buffered channel with a free slot completes immediately without any receiver. Since each attempt sends exactly once, there is a slot reserved for it no matter what the caller did — so every goroutine's send returns, every goroutine returns, and the abandoned values sit in the channel until the channel itself becomes unreachable and is collected along with them. The sizing rule is *one slot per sender that may still send after the last receive*, and the safe way to say it is: capacity equals the number of attempts. Some people argue N-1 is enough because the caller takes one. It is, in this exact shape, but the reasoning is fragile — add a `select` with a timeout, an early return before the receive, or a retry, and N-1 silently becomes a leak again. Size it to N and stop thinking about it. ## The alternatives, and when to prefer them - **Non-blocking send.** `select { case ch <- r: default: }` — the send is attempted and dropped if nobody can take it. This bounds the goroutine correctly, but it *discards* results, so it only makes sense when a late result is genuinely useless. With a buffer of N you never discard anything. - **Cancellation-aware send.** `select { case ch <- r: case <-ctx.Done(): }` — the attempt gives up sending when the shared context is cancelled. Useful when the channel is deliberately unbuffered for backpressure reasons; unnecessary when it is sized to N. - **Just size the buffer.** Simplest, allocates N slots of one small struct, and has no correctness edge cases. This is the idiomatic answer. Note that the buffer is not a performance optimisation here and not an ordering device. Its only job is to guarantee that a send can always complete, which is what lets an abandoned goroutine exit. ## How you find it in a running program The goroutine profile is the direct evidence: `runtime/pprof`'s goroutine profile, or `/debug/pprof/goroutine?debug=2` if the process exposes the pprof handlers, prints every goroutine with its stack and its block reason. A leak of this shape is unmistakable — a large and *growing* count of goroutines all sitting at the same source line with the reason `chan send`, N-1 of them per call the process has made. `runtime.NumGoroutine()` climbing monotonically with request count is the cheap smoke test that tells you to go look. Go also ships a dedicated goroutine-leak profile — `goroutineleak` in `runtime/pprof`, served at `/debug/pprof/goroutineleak` — which reports goroutines that are permanently blocked rather than merely blocked right now. That distinction is exactly this bug: a loser parked on a send that no receiver can ever satisfy. ## The reviewer's rule of thumb Whenever you see `go func() { ch <- ... }()` fanned out N ways, find the receive. If the receiving side does not consume exactly as many values as will be sent, the channel must be buffered to cover the difference or the sends must be non-blocking. In first-result-wins the receiving side consumes one and the senders are N, so the difference is the whole point.
- Would a capacity of N-1 also work, since the caller consumes one value?In this exact shape it does, because one value is handed to the receiver and the remaining N-1 fit in the buffer. But the reasoning depends on the receive definitely happening; add an early return, a timeout case, or a second attempt round and N-1 silently becomes a leak. Size it to the number of attempts so the property is local and obvious.
- How would you confirm this leak in a running service rather than by reading the code?Watch `runtime.NumGoroutine()` against request count — monotonic growth is the smoke test. Then take the goroutine profile from `runtime/pprof` and look for many goroutines stopped at the same source line with the reason `chan send`. Go's goroutine-leak profile, at `/debug/pprof/goroutineleak`, reports specifically those that are permanently blocked rather than momentarily blocked.
- When is a non-blocking send preferable to buffering the channel?When a late value is genuinely worthless and you would rather drop it than allocate for it — `select` with the send in one case and `default` in the other lets the attempt discard its result and exit. It bounds the goroutine just as well, but it loses information, so buffering is the better default when N is small and known.
- What exactly does a leaked losing goroutine hold onto?Its own stack, everything its closure captured, and the value it is trying to send — which in a download race is the fully materialised archive, plus the response and connection it still references. That is why the leak shows up as growing memory and exhausted connections, not just a rising goroutine count.
saying these in an interview costs you the question
- Thinks a blocked goroutine is eventually garbage collected
- Says buffering is only a throughput optimisation
- Believes the runtime reclaims goroutines whose result is unwanted
- Sizes the channel to 1 because only one result is used
- Claims the leak is harmless because the values are small