Why does sync.WaitGroup.Wait hang forever when a worker goroutine returns early on an error?
answer
- it is only a counter
- which exit paths reach your decrement?
- an error return skips the last line
- Wait has no timeout at all
- defer it on the goroutine's first line
basics
~20 sA WaitGroup is only a counter: Add raises it, Done lowers it, Wait blocks until it reaches zero. A worker that returns before reaching its Done leaves the counter above zero, so Wait never unblocks. Put Done in a defer.
solid answer
~50 s`sync.WaitGroup` counts outstanding work and nothing more. `wg.Add(1)` before each `go` statement raises the counter, each `wg.Done()` lowers it, and `wg.Wait()` parks the calling goroutine until the counter hits zero. If `Done` sits at the bottom of the worker function, every early exit — an error return, a `continue`-style guard clause, a recovered panic — skips it, the counter never reaches zero, and `Wait` blocks for the life of the process. There is no timeout on `Wait`; it waits forever. The fix is mechanical: make `defer wg.Done()` the first statement of the goroutine so every exit path decrements exactly once. Call `wg.Add` in the parent before `go`, never inside the goroutine where it can race with `Wait`, and pass the group as a `*sync.WaitGroup` — copying one splits the counter so the parent's copy never moves.
code
go · 10 lineswg.Add(1)
go func() {
f, err := os.Open(name)
if err != nil {
return // skips the wg.Done below
}
defer f.Close()
process(f)
wg.Done()
}()go deeper
Know the three calls and the one rule: Add before go, defer wg.Done() as the goroutine's first line, Wait in the parent. Be able to point at an early return and say why it hangs.
Explain the counter mechanics and both directions of failure — a missed decrement hangs Wait, an Add inside the goroutine makes Wait return too early — and why copying a WaitGroup silently splits the count.
Show how you would find it in a running process: a goroutine parked in Wait with no workers left is a stale count, not slow work. Talk about getting errors out on a buffered channel so the error path cannot itself park before the decrement.
Own the convention rather than the bug: decide whether teams start goroutines through a helper that owns Add/Done and error collection, so the hand-rolled version cannot appear in review at all, and what the vet configuration enforces in CI.
## What a WaitGroup actually is A `sync.WaitGroup` is a counter with a queue of waiters attached. Three operations touch it: - `Add(n)` adds `n` to the counter (usually 1, before starting a goroutine); - `Done()` is exactly `Add(-1)`; - `Wait()` parks the calling goroutine until the counter is zero, then wakes every waiter. That is the whole model. It knows nothing about goroutines, errors, results or time. It cannot tell that a worker died; it only knows whether the number it is holding is zero. ## The failure ``` wg.Add(1) go func() { f, err := os.Open(name) if err != nil { return // the Done below never runs } process(f) wg.Done() }() wg.Wait() // parks forever the first time a file is missing ``` The happy path decrements and everything looks fine in tests. The first production input that triggers the error branch leaves the counter at 1 permanently. `Wait` has no deadline and no cancellation, so the parent goroutine is parked for the rest of the process's life — and whatever that parent was doing (a request handler, a shutdown sequence, a batch stage) stops there. If the parent holds a lock or a connection, that is held too. This is the same defect with a different mask: a worker that panics and is recovered by a deferred function that then returns, or a worker whose `Done` is inside an `if` branch, or a loop that calls `Add(1)` per item but returns from the loop early on the first failure without matching decrements. ## The rule that removes the whole class Make the decrement the goroutine's first deferred statement: ``` wg.Add(1) go func() { defer wg.Done() ... }() ``` Deferred calls run when the function returns *however* it returns, including while a panic unwinds. Put it first and it is impossible for a later edit to add an exit path that skips it — which is the real value, because these bugs are almost always introduced by a later edit to a function that was correct when written. ## Where Add belongs, and the opposite bug `Add` must be called by the goroutine that starts the work, **before** the `go` statement. Moving `wg.Add(1)` inside the new goroutine creates the mirror-image failure: the parent may reach `Wait` before the new goroutine gets scheduled, see a counter of zero, and return immediately while the workers are still running. That is not a hang — it is a silent early return, results discarded, sometimes a write to a closed resource. It is harder to spot than the hang. Modern toolchains catch it: `go vet` gained a `waitgroup` check in Go 1.25 that reports an `Add` call misplaced inside the new goroutine. ## The other ways to break a WaitGroup - **Copying it.** A `WaitGroup` must not be copied after first use. Passing `wg sync.WaitGroup` by value into a worker gives that worker its own counter; the parent's counter never drops. Always pass `*sync.WaitGroup`, or close over it. - **Over-decrementing.** More `Done` calls than `Add` calls drives the counter negative and panics with a negative-counter error. Two `defer wg.Done()` lines, or a `Done` plus a deferred `Done`, do this. - **Reusing while waiting.** Adding to a group whose counter has hit zero while another goroutine is still inside `Wait` is unsupported. Use a fresh group per round. ## Since Go 1.25 `wg.Go(f)` starts `f` in a new goroutine and performs the `Add` and the deferred `Done` itself, so the whole family of mistakes above is unreachable for work started that way. On older toolchains, or where you need the goroutine's body inline, write `Add` before `go` and `defer wg.Done()` first. ## Getting errors out Because the group carries no results, an error has to travel some other way. A buffered channel with capacity equal to the number of workers is the simple form: every worker can send its error without blocking, so the send cannot itself become a second hang, and the parent drains it after `Wait` returns. A channel that is too small turns the error path into a parked send, which stops the worker before its deferred `Done` — the original bug, one level down. ## Diagnosing it A process hung this way shows a goroutine parked in `sync.(*WaitGroup).Wait` in a full goroutine dump, with the workers gone. That combination — a waiter and no workers — is the signature: the count is stale, not the work.
- Where exactly should wg.Add go, and what breaks if it is inside the goroutine?In the parent, immediately before the `go` statement. Inside the goroutine it races with `Wait`: the parent can reach `Wait` before the goroutine is scheduled, see zero, and return while workers are still running — an early return with discarded results rather than a hang. Go 1.25's `go vet` `waitgroup` check reports that placement.
- If a worker panics before its Done, does Wait still hang?With `defer wg.Done()` the decrement runs while the panic unwinds, so the group is consistent. But an unrecovered panic in any goroutine terminates the whole program, so `Wait` never gets to return anyway. Only if that goroutine recovers in its own deferred function does the group's state matter — and then the deferred `Done` has already run.
- What happens if Done is called more times than Add?The counter goes negative and the call panics with a negative WaitGroup counter. The usual cause is a `Done` left at the end of the function after `defer wg.Done()` was added at the top, so both fire. The panic is loud and immediate, which makes it far less dangerous than the silent hang.
saying these in an interview costs you the question
- Puts wg.Done() at the end of the function instead of in a defer
- Calls wg.Add(1) inside the goroutine it is counting
- Passes sync.WaitGroup by value to the worker function
- Believes Wait gives up after some timeout
- Adds a Done in the error branch as well, driving the counter negative