sync.WaitGroup.Wait never returns and a Go test hangs to its timeout — how do you diagnose it?
answer
- the counter never reached zero
- let the timeout dump every stack
- are the workers still in the dump?
- an Add separated from its go statement
- -race and a CPU profile show nothing here
basics
~20 sLet the go test timeout fire: it dumps every goroutine's stack. Find the one parked in sync.(*WaitGroup).Wait, then read the other stacks — either a Done was never reached, or a counted goroutine is itself blocked and never finishes.
solid answer
~50 sA stuck `Wait` means the counter never reached zero, and the counter is the only state involved, so there are exactly two families of cause: an `Add` with no matching `Done`, or a counted goroutine that never returns. The dump does the sorting for you — `go test` panics at its timeout (10 minutes by default, so lower it with `-timeout 30s`) and prints every goroutine's stack; a running binary gives you the same thing via `SIGQUIT` with `GOTRACEBACK=all`, or `/debug/pprof/goroutine?debug=2`. Find the frame in `sync.(*WaitGroup).Wait`, note which function is below it, then count the counted goroutines still alive. If they are all gone, the bug is bookkeeping: a loop that `Add`s then `continue`s past the `go` statement, a `Done` that is not deferred and gets skipped by an early return, or a group copied by value. If they are still there, the WaitGroup is innocent — read where *they* are parked and fix that.
code
go · 11 linesfor _, m := range migrations {
wg.Add(1)
if !m.Enabled {
continue // BUG: no goroutine will Done for this unit
}
go func() {
defer wg.Done()
apply(m)
}()
}
wg.Wait() // one Add is unmatched, so this never returnsgo deeper
Know that a hang here means the counter never hit zero, and that the go test timeout prints every goroutine's stack for you to read.
Explain the two families — an unmatched Add versus a worker that never finishes — and name the code shapes that cause the first: an Add separated from its go statement, and a Done that is not deferred.
Show the working method: lower -timeout, read the dump, decide from the surviving goroutines which family it is, and say clearly why -race and a CPU profile contribute nothing to a hang.
Turn one incident into prevention: enforce Add adjacent to go (or the API that pairs them), require every worker to take a cancellable context so liveness hangs become bounded failures, and keep a goroutine dump reachable in production.
## First, narrow it to two families `sync.WaitGroup` has one piece of state: a counter. `Wait` returns when it is zero. So a `Wait` that never returns means the counter never reached zero, and there are only two ways for that to be true: 1. **Bookkeeping** — more `Add` than `Done` was ever going to run. The counted goroutines have all exited (or were never started) and the counter is still positive. 2. **Liveness** — the arithmetic is right, but one or more counted goroutines are still alive and blocked, so their `Done` has not run yet and never will. The difference is decided by one question: **are the counted goroutines still in the dump?** Everything below is about getting that dump and reading it. ## Getting the evidence **In a test.** `go test` kills a hung test at its timeout — 10 minutes by default — and, crucially, prints a panic followed by the stack of *every* goroutine in the process. Set `-timeout 30s` while you are iterating so you are not waiting ten minutes for evidence you already know is coming. **In a running binary.** Send it `SIGQUIT` (Ctrl-\ on a terminal) and the runtime prints all goroutine stacks and exits. `GOTRACEBACK=all` widens the dump to include runtime-internal goroutines when the default output is not enough. **In a service you must not kill.** If `net/http/pprof` is registered, `/debug/pprof/goroutine?debug=2` gives the same full stacks, live, with no downtime. Go 1.27 also ships a goroutine-leak profile (`goroutineleak` in `runtime/pprof`, served at `/debug/pprof/goroutineleak`) which reports goroutines that can no longer make progress — useful when the raw dump has thousands of entries and you need the parked ones singled out. ## Reading it Find the goroutine whose stack contains `sync.(*WaitGroup).Wait`. The frame immediately below it names the function that is joining — your migration runner, say. That confirms the symptom but tells you nothing about the cause. Now scan the rest of the dump for the goroutines that were supposed to call `Done` — they will have the worker function in their stacks. - **None of them exist.** It is a bookkeeping bug. Go read the launching loop. - **They are there, parked.** The WaitGroup is a bystander. Whatever they are blocked on — a channel send with no receiver, a lock held elsewhere, a network call with no deadline — is the real bug, and fixing `Wait` is not the fix. ## The bookkeeping bugs, in the order they occur in practice **An `Add` whose `go` statement never runs.** A filter or an early error inside the loop body after the `Add`: ```go for _, m := range migrations { wg.Add(1) if !m.Enabled { continue // BUG: nothing will ever Done for this unit } go func() { defer wg.Done() apply(m) }() } wg.Wait() // one Add is unmatched, so this blocks forever ``` The fix is to move `wg.Add(1)` to the line immediately before `go`, so the increment and the launch cannot be separated by any control flow. This is why `Add` directly above `go` is a convention and not just a style. **A `Done` that is not deferred.** `wg.Done()` at the bottom of the worker is skipped by every early `return` and by any panic that is recovered further out. `defer wg.Done()` as the first statement is the only form that survives every exit path. **A group copied by value.** A helper taking `sync.WaitGroup` instead of `*sync.WaitGroup` decrements a copy, so the caller's counter never moves and the symptom is identical. `go vet`'s copylocks check reports this, and `go test` runs vet by default — so check the vet output before you go stack-diving. **Bulk `Add(len(units))` over a loop that skips units.** Same shape as the first case, one line higher up. ## Prevention, once you have found it Go 1.25 added a `waitgroup` analyzer to `go vet` that reports the mirror-image mistake — `wg.Add(1)` written inside the goroutine it counts, which makes `Wait` return *too early* rather than never. Running vet catches that class for free. The structural fix for most of this is Go 1.25's `wg.Go(f)`: it increments, runs `f`, and decrements when `f` returns, so there is no separate `Add` to strand and no `Done` to skip. It does not help with the liveness family — a worker blocked forever still blocks `Wait` — which remains an argument for giving every worker a `context.Context` and a deadline rather than trusting it to finish. ## A note on what does not help The race detector will not find this. `-race` reports unsynchronised concurrent access to memory that actually happened; a counter that never reaches zero is correctly synchronised and perfectly deterministic. Likewise a CPU profile shows nothing, because a hung program is not burning CPU. The goroutine dump is the tool for this failure, and it is almost always sufficient on its own.
- The dump shows the counted goroutines still alive and blocked. Why is that not a WaitGroup bug?Because the arithmetic is correct: each live goroutine still owes exactly one `Done`, and it will run when the goroutine finishes. `Wait` is faithfully reporting that the work is not done. The bug is wherever those goroutines are parked — a send on a channel nobody reads, a lock held elsewhere, a request with no deadline. Fix that, and the join resolves itself.
- How do you get the same evidence from a production service you cannot restart?Register `net/http/pprof` and fetch `/debug/pprof/goroutine?debug=2`, which returns full stacks for every goroutine with no downtime. Go 1.27 additionally serves a goroutine-leak profile at `/debug/pprof/goroutineleak` that isolates goroutines which can no longer make progress, which is far more readable when the process holds thousands of them.
- Why is placing wg.Add(1) on the line immediately before the go statement a rule rather than a style preference?Because any statement between them can skip the launch — a `continue`, an early `return`, an error check, a filter — leaving an increment that nothing will ever match, and the symptom is a hang far from the cause. Keeping them adjacent makes the pairing locally verifiable by eye. Go 1.25's `wg.Go` enforces it structurally.
- Would running the tests with -race help find this?No. The race detector reports unsynchronised concurrent memory access that actually occurred during the run; it says nothing about a counter that is correctly synchronised and simply never reaches zero. A hung program also burns no CPU, so a CPU profile is empty. The goroutine stack dump is the right instrument, and usually the only one needed.
saying these in an interview costs you the question
- Reaches for the race detector to explain a hang
- Adds a sleep or a shorter timeout instead of reading the dump
- Never lowers -timeout and waits ten minutes for evidence
- Removes the Wait call to make the test pass
- Blames the scheduler rather than counting Add against Done
- Ignores that the counted goroutines are still alive in the dump