In a worker pool, why must close(results) happen only after wg.Wait() returns?
answer
- two channels, two different closers
- one close is a signal, one is a promise
- a send on a closed channel panics
- what proves every sender has gone
- wait and close in their own goroutine
basics
~20 sClosing a channel that a worker might still send on panics. wg.Wait() returning is the only proof that every worker has left its loop, so close(results) must follow it. close(jobs) merely tells the workers to finish.
solid answer
~50 sThe two closes mean different things and have different owners. `close(jobs)` is the producer saying "no more work", and it is what ends each worker's `for j := range jobs` loop. `close(results)` is a promise to the collector that no value will ever appear again — and the senders on `results` are the workers, so only something that observes all of them can make that promise. `wg.Wait()` is that observation: every worker runs `defer wg.Done()`, so when the counter hits zero no goroutine can send on `results` again. Close it earlier and a still-running worker panics with `send on closed channel`, or, if the timing happens to work, the collector's range ends early and results are silently dropped. With an unbuffered results channel the wait-and-close pair lives in its own goroutine so the owner can be receiving while workers are still sending.
code
go · 30 linesjobs := make(chan int)
results := make(chan int) // unbuffered: the owner must be receiving
var wg sync.WaitGroup
for i := 0; i < workers; i++ {
wg.Add(1) // before the go statement, never inside it
go func() {
defer wg.Done()
for j := range jobs {
results <- j * 2
}
}()
}
go func() {
for _, in := range inputs {
jobs <- in
}
close(jobs) // the sender closes: ends every worker loop
}()
go func() {
wg.Wait() // proof that no worker can send again
close(results) // only now
}()
total := 0
for r := range results {
total += r
}go deeper
Remember the sequence and be able to recite it: close the jobs channel, wait for the workers, then close the results channel. Know that sending on an already closed channel panics.
Explain why the WaitGroup is the proof rather than a formality, and why the wait-and-close pair sits in its own goroutine when the results channel is unbuffered.
Be ready to describe both failure modes from a real incident: the crash from a send on a closed channel, and the quieter one where the collector stops early and a batch reports success with missing rows.
Own the shape as a reviewable convention: which goroutine owns each close, how it is documented on an exported pool API, and how you would keep the whole team from re-deriving this ordering by hand.
## Two channels, two different closes, two different owners A pool with a `jobs` channel and a `results` channel has two closes in it, and they mean opposite things. - `close(jobs)` is a **signal to the workers**: "no more work is coming." It is called by the producer — the goroutine that sends on `jobs` — because the convention in Go is that the sender closes, and a receiver that closes leaves other senders to panic. - `close(results)` is a **promise to the collector**: "no further value will ever appear here." Its only legal caller is somebody who can prove that every sender on `results` has finished. The workers are the senders on `results`. The pool owner cannot see inside them. The `WaitGroup` is what turns "I think they are done" into proof. ## The retirement order 1. The producer sends its last job and calls `close(jobs)`. 2. Each worker's `for j := range jobs` drains whatever is buffered and then ends. 3. Each worker returns, running its `defer wg.Done()`. 4. `wg.Wait()` returns. This is the exact moment at which no goroutine can ever send on `results` again. 5. `close(results)` — now safe. 6. The collector's `for r := range results` drains the last values and ends. Skip step 4 and put `close(results)` next to `close(jobs)` and you get one of two outcomes, both bad. If a worker is still in its loop, its next `results <- r` panics with `panic: send on closed channel`, which takes the process down — a panic in a worker goroutine cannot be recovered by the goroutine that started it. If, by luck of timing, every worker happens to have finished already, the program works — and will fail in production when the machine is busier. Even in the lucky case the collector stops early: `range` over a closed channel ends as soon as the buffer is empty, so results that were still in flight are silently dropped. ## Why the wait usually lives in its own goroutine With a **buffered** `results` channel large enough for every result, the owner can call `wg.Wait()` inline, then `close(results)`, then drain. With an **unbuffered** or small-buffered channel it cannot: the workers block on their sends until somebody receives, so the owner must already be inside the receive loop while they are still running. If it calls `wg.Wait()` first, it never reaches the loop and the whole pool wedges. Hence the standard shape: ```go go func() { wg.Wait() close(results) }() for r := range results { // running while the workers are still sending collect(r) } ``` The closer goroutine does nothing but wait and close. The owner is free to stream results out as they arrive, so memory does not have to hold the whole result set. ## Where `wg.Add` goes `wg.Add(1)` must run in the goroutine that starts the worker, **before** the `go` statement — not as the first line inside the worker. If it runs inside, `wg.Wait()` can observe a counter of zero before the goroutine has been scheduled, return immediately, and close `results` under a worker that is about to send. That is the same panic, arrived at from a different direction. Recent Go ships a `waitgroup` analyzer in `go vet` that flags a misplaced `Add` inside the goroutine. Recent Go also offers `wg.Go(func() { ... })`, which does the `Add(1)` and the `defer Done()` for you and removes that whole class of mistake: ```go for i := 0; i < workers; i++ { wg.Go(func() { for j := range jobs { results <- do(j) } }) } ``` The ordering rule is unchanged: `wg.Wait()` still has to return before `close(results)`. ## Related traps - **Never have a worker close `results`.** With N workers, the first to exit closes it and the rest panic on their next send, and the second one to exit panics on the double close. - **A `WaitGroup` must not be copied** after use. Capture it by reference (a closure over a `var wg sync.WaitGroup`, or a `*sync.WaitGroup` field), never pass it by value. - **Closing is not cancelling.** `close(jobs)` does not interrupt a worker mid-job; it only ends the loop once the channel is drained. Stopping early is a cancellation concern, not a close concern. ## What an interviewer is checking That you can say *why* `wg.Wait()` is the proof, not just that the order is memorised: it is the only construct in the pool that observes every worker's exit.
- What actually happens if close(results) is called right next to close(jobs)?Any worker still in its loop panics on its next send with `send on closed channel`, and a panic in a worker goroutine takes the process down. If the timing happens to spare you, the collector's `for r := range results` ends as soon as the buffer empties, so results still in flight are silently lost — the worse outcome, because it looks like success.
- Why is the wg.Wait()/close(results) pair usually put in its own goroutine?Because with an unbuffered results channel the workers block on their sends until someone receives. If the owner calls `wg.Wait()` inline it never reaches its receive loop, so the workers never finish, so the counter never drops — the pool wedges. Putting the wait in a separate goroutine lets the owner stream results out while the workers run.
- Where must wg.Add(1) be called, and why does it matter?In the goroutine that starts the worker, before the `go` statement. Inside the worker it races: `wg.Wait()` can see a zero counter before the goroutine is scheduled, return immediately, and close the results channel under a worker about to send. Recent Go's `go vet` ships a `waitgroup` analyzer that reports exactly this misplacement.
- Should a worker ever close the results channel itself?No. With N workers the first one to exit closes it, the others panic on their next send, and the second one to exit panics on the double close. Closing belongs to a single goroutine that can observe all senders — which in a pool means the one holding the WaitGroup.
close(jobs) is telling the sorters no more post is coming. close(results) is removing the outbound mailbox, and you only do that once you have watched every sorter walk out.
saying these in an interview costs you the question
- Closes the results channel next to the jobs channel
- Has each worker close the results channel on exit
- Thinks closing a channel unblocks pending sends
- Calls wg.Wait() before entering the collector loop
- Puts wg.Add(1) inside the worker goroutine
- Passes the sync.WaitGroup by value into the worker