A batch runner sends all 500 jobs before reading from an unbuffered results channel and stops making progress. Why?
answer
- count who is left to receive
- an unbuffered send needs a waiting receiver
- workers parked in send, submitter parked in send
- the collector never got to run
- submit from its own goroutine
basics
~20 sEvery worker is blocked sending its first result, because the only goroutine that would receive from the results channel is the one still sending jobs. With no worker receiving, the jobs send blocks too, and nothing can move.
solid answer
~50 sIt is a cycle of waits. A send on an unbuffered channel completes only when some goroutine is at the matching receive, and the only goroutine that will ever receive from `results` is the owner — which is still in its submission loop. So each worker finishes a job, parks in `results <- r`, and stops receiving from `jobs`; once all N workers are parked, the owner's next `jobs <- j` parks too. It is size-dependent, which is why tests pass: with a buffer of sixteen it works for any batch of sixteen. The fix that scales is to move submission into its own goroutine so the owner goes straight into `for r := range results` and something is always receiving. Buffering `results` to the job count also works but holds the whole result set in memory, so it only survives while the batch stays small. Adding workers or a sleep fixes nothing.
code
go · 24 linesjobs := make(chan int)
results := make(chan int) // unbuffered
var wg sync.WaitGroup
for i := 0; i < 4; i++ {
wg.Add(1)
go func() {
defer wg.Done()
for j := range jobs {
results <- j * 2 // all four park here
}
}()
}
for j := 0; j < 500; j++ {
jobs <- j // parks once every worker is stuck in the send above
}
close(jobs)
wg.Wait()
close(results)
for r := range results { // never reached
_ = r
}go deeper
Remember that a send on an unbuffered channel waits for a receiver, so someone must be reading the results channel while the workers are producing.
Trace the cycle out loud: workers parked in the results send, so nobody receives jobs, so the submitter parks. Then name the fix that does not depend on batch size.
An interviewer expects the diagnosis from evidence - a goroutine dump showing N goroutines in chan send - plus the judgment that buffering to the job count trades a hang for unbounded memory.
Own the guardrail rather than the fix: what the pool's contract says about streaming versus collecting, and the completeness test that stops this shape being reintroduced across teams.
## The topology of the wedge Draw the goroutines and the edges between them. - One goroutine (call it the owner) is running `for j := range inputs { jobs <- j }`. It intends to read `results` afterwards. - Four workers are running `for j := range jobs { results <- do(j) }`. - `results` is unbuffered, so a send on it completes only when some goroutine is *at* a receive. Nobody is at a receive on `results`, because the only goroutine that will ever be there is the owner, and the owner is still in its submission loop. So the first worker to finish a job parks in `results <- r`. So does the second, third and fourth. Now no goroutine is receiving from `jobs` either — all four are parked on the results send — so the owner's next `jobs <- j` parks too. Every goroutine is blocked, and each is waiting for a goroutine that is waiting for it. That is a deadlock in the textbook sense: a cycle of waits. ## Why "it worked in the test" The wedge is size-dependent, which is what makes it a production bug rather than a compile error. - If `results` had a buffer of, say, 16, the first sixteen results would be absorbed and the program would run to completion for any batch of sixteen jobs or fewer. Seventeen jobs and it hangs. - If `jobs` is buffered, the owner gets further before it parks, but it still parks. - With one worker and one job it always works. So the classic story is: unit test with three jobs passes, staging with a hundred passes because somebody buffered `results` to 128, production with fifty thousand hangs on a Tuesday. ## Confirming it in a running process The Go runtime prints `fatal error: all goroutines are asleep - deadlock!` and a full stack dump — but **only when every goroutine in the process is blocked**. A batch job does hit that. A long-lived service almost never does: an HTTP server goroutine, a metrics ticker, or the runtime's own timer work is still runnable, so the process just sits there consuming no CPU and finishing no requests. To confirm it in that case, get goroutine stacks: - send `SIGQUIT` (Ctrl-\ on a terminal), which dumps every goroutine's stack and exits; `GOTRACEBACK=all` makes the dump complete; - or, if `net/http/pprof` is wired up, fetch `/debug/pprof/goroutine?debug=2`, which prints the same stacks without killing anything. You are looking for N goroutines parked in `chan send` at the same source line — the worker's send on `results` — and one goroutine parked in `chan send` at the submission line. The count matching the worker count is the tell. ## The three fixes, and what each costs **Submit from its own goroutine.** The owner starts the producer as `go func() { ...; close(jobs) }()` and goes straight into `for r := range results`. Now something is always receiving, the workers never park indefinitely, and memory holds only the results you have not processed yet. This is the default answer: it streams, and it works for any job count. **Buffer `results` to the number of jobs.** Legitimate *only* when the job count is known, bounded and small, because you are now holding the entire result set in memory. It converts a hang into memory growth — which is a fine trade for a 500-row batch and a terrible one for a 50-million-row one. An interviewer will ask "and when the batch is ten times bigger?"; the honest answer is that this fix does not survive it. **Collect in a separate goroutine.** Same effect as the first fix from the other side: the owner submits, and a collector goroutine drains `results` into a slice or writes them out. You then need one more synchronisation to know when the collector is done. What is *not* a fix: adding workers (they all park in the same place), adding a `time.Sleep`, closing `jobs` earlier, or reaching for a mutex. None of them puts a receiver on `results`. ## Catching it before production The test that finds this is a **completeness assertion**: submit K jobs, count the results, assert that K went in and K came out — with K chosen larger than any buffer in the pool, so the buffered path cannot mask it. Run it under a bounded timeout (`go test -timeout 30s`) so a wedge fails the run instead of hanging; when the timeout fires, `go test` panics and prints every goroutine's stack, which hands you the same `chan send` evidence for free. The assertion is also the right regression test for the whole retirement sequence: a pool that drops results because `results` was closed too early fails the same check, with K in and fewer than K out.
- You give the results channel a buffer of 500 and the hang goes away. What have you accepted?That the whole result set is resident in memory, and that the fix is sized to today's batch. When the batch becomes 50 million rows the hang returns as memory growth or an OOM. Buffering is legitimate when the job count is known, bounded and small; otherwise stream by submitting in a goroutine and collecting as results arrive.
- In a long-running service the pool wedges but the runtime never prints a deadlock error. Why, and how do you confirm the diagnosis?The runtime reports `fatal error: all goroutines are asleep - deadlock!` only when every goroutine in the process is blocked; a live HTTP server or timer goroutine keeps it out. Get goroutine stacks instead — SIGQUIT with GOTRACEBACK=all, or the goroutine profile at /debug/pprof/goroutine?debug=2 — and look for N goroutines parked in chan send at the worker's send line.
- How would you write a test that catches this before production?A completeness assertion: submit K jobs, count the results, assert K in equals K out — with K larger than any buffer in the pool so buffering cannot mask the wedge. Run it under a bounded `go test -timeout`, so a wedge fails the run instead of hanging; when that timeout fires, go test panics and prints every goroutine's stack, handing you the evidence.
- Would adding more workers help?No. Every additional worker parks in the same send on the results channel, so you just have more blocked goroutines and the submitter still stalls one job later. Nothing changes until some goroutine is receiving from the results channel; that is what the fix has to introduce.
The dispatcher will not read the out-tray until he has finished handing out every job, and the workers cannot take another job until someone empties the out-tray. Both sides wait forever.
saying these in an interview costs you the question
- Blames the Go scheduler or GOMAXPROCS for the hang
- Enlarges the buffer without asking how many jobs there can be
- Says closing the jobs channel would unblock the workers
- Claims a deadlock always prints a fatal runtime error
- Adds a time.Sleep before collecting results
- Adds more workers to get past the stall