skip to content

sync.WaitGroup.Wait never returns and a Go test hangs to its timeout — how do you diagnose it?

level: seniorimportance: should knowfreq 50%

answer

  1. the counter never reached zero
  2. let the timeout dump every stack
  3. are the workers still in the dump?
  4. an Add separated from its go statement
  5. -race and a CPU profile show nothing here

basics

~20 s

Let the go test timeout fire: it dumps every goroutine's stack. Find the one parked in sync.(*WaitGroup).Wait, then read the other stacks — either a Done was never reached, or a counted goroutine is itself blocked and never finishes.

solid answer

~50 s

A stuck `Wait` means the counter never reached zero, and the counter is the only state involved, so there are exactly two families of cause: an `Add` with no matching `Done`, or a counted goroutine that never returns. The dump does the sorting for you — `go test` panics at its timeout (10 minutes by default, so lower it with `-timeout 30s`) and prints every goroutine's stack; a running binary gives you the same thing via `SIGQUIT` with `GOTRACEBACK=all`, or `/debug/pprof/goroutine?debug=2`. Find the frame in `sync.(*WaitGroup).Wait`, note which function is below it, then count the counted goroutines still alive. If they are all gone, the bug is bookkeeping: a loop that `Add`s then `continue`s past the `go` statement, a `Done` that is not deferred and gets skipped by an early return, or a group copied by value. If they are still there, the WaitGroup is innocent — read where *they* are parked and fix that.

code

go · 11 lines
go
for _, m := range migrations {
	wg.Add(1)
	if !m.Enabled {
		continue // BUG: no goroutine will Done for this unit
	}
	go func() {
		defer wg.Done()
		apply(m)
	}()
}
wg.Wait() // one Add is unmatched, so this never returns

go deeper

for a junior

Know that a hang here means the counter never hit zero, and that the go test timeout prints every goroutine's stack for you to read.

for a middle

Explain the two families — an unmatched Add versus a worker that never finishes — and name the code shapes that cause the first: an Add separated from its go statement, and a Done that is not deferred.

for a senior

Show the working method: lower -timeout, read the dump, decide from the surviving goroutines which family it is, and say clearly why -race and a CPU profile contribute nothing to a hang.

for a principal

Turn one incident into prevention: enforce Add adjacent to go (or the API that pairs them), require every worker to take a cancellable context so liveness hangs become bounded failures, and keep a goroutine dump reachable in production.

## First, narrow it to two families `sync.WaitGroup` has one piece of state: a counter. `Wait` returns when it is zero. So a `Wait` that never returns means the counter never reached zero, and there are only two ways for that to be true: 1. **Bookkeeping** — more `Add` than `Done` was ever going to run. The counted goroutines have all exited (or were never started) and the counter is still positive. 2. **Liveness** — the arithmetic is right, but one or more counted goroutines are still alive and blocked, so their `Done` has not run yet and never will. The difference is decided by one question: **are the counted goroutines still in the dump?** Everything below is about getting that dump and reading it. ## Getting the evidence **In a test.** `go test` kills a hung test at its timeout — 10 minutes by default — and, crucially, prints a panic followed by the stack of *every* goroutine in the process. Set `-timeout 30s` while you are iterating so you are not waiting ten minutes for evidence you already know is coming. **In a running binary.** Send it `SIGQUIT` (Ctrl-\ on a terminal) and the runtime prints all goroutine stacks and exits. `GOTRACEBACK=all` widens the dump to include runtime-internal goroutines when the default output is not enough. **In a service you must not kill.** If `net/http/pprof` is registered, `/debug/pprof/goroutine?debug=2` gives the same full stacks, live, with no downtime. Go 1.27 also ships a goroutine-leak profile (`goroutineleak` in `runtime/pprof`, served at `/debug/pprof/goroutineleak`) which reports goroutines that can no longer make progress — useful when the raw dump has thousands of entries and you need the parked ones singled out. ## Reading it Find the goroutine whose stack contains `sync.(*WaitGroup).Wait`. The frame immediately below it names the function that is joining — your migration runner, say. That confirms the symptom but tells you nothing about the cause. Now scan the rest of the dump for the goroutines that were supposed to call `Done` — they will have the worker function in their stacks. - **None of them exist.** It is a bookkeeping bug. Go read the launching loop. - **They are there, parked.** The WaitGroup is a bystander. Whatever they are blocked on — a channel send with no receiver, a lock held elsewhere, a network call with no deadline — is the real bug, and fixing `Wait` is not the fix. ## The bookkeeping bugs, in the order they occur in practice **An `Add` whose `go` statement never runs.** A filter or an early error inside the loop body after the `Add`: ```go for _, m := range migrations { wg.Add(1) if !m.Enabled { continue // BUG: nothing will ever Done for this unit } go func() { defer wg.Done() apply(m) }() } wg.Wait() // one Add is unmatched, so this blocks forever ``` The fix is to move `wg.Add(1)` to the line immediately before `go`, so the increment and the launch cannot be separated by any control flow. This is why `Add` directly above `go` is a convention and not just a style. **A `Done` that is not deferred.** `wg.Done()` at the bottom of the worker is skipped by every early `return` and by any panic that is recovered further out. `defer wg.Done()` as the first statement is the only form that survives every exit path. **A group copied by value.** A helper taking `sync.WaitGroup` instead of `*sync.WaitGroup` decrements a copy, so the caller's counter never moves and the symptom is identical. `go vet`'s copylocks check reports this, and `go test` runs vet by default — so check the vet output before you go stack-diving. **Bulk `Add(len(units))` over a loop that skips units.** Same shape as the first case, one line higher up. ## Prevention, once you have found it Go 1.25 added a `waitgroup` analyzer to `go vet` that reports the mirror-image mistake — `wg.Add(1)` written inside the goroutine it counts, which makes `Wait` return *too early* rather than never. Running vet catches that class for free. The structural fix for most of this is Go 1.25's `wg.Go(f)`: it increments, runs `f`, and decrements when `f` returns, so there is no separate `Add` to strand and no `Done` to skip. It does not help with the liveness family — a worker blocked forever still blocks `Wait` — which remains an argument for giving every worker a `context.Context` and a deadline rather than trusting it to finish. ## A note on what does not help The race detector will not find this. `-race` reports unsynchronised concurrent access to memory that actually happened; a counter that never reaches zero is correctly synchronised and perfectly deterministic. Likewise a CPU profile shows nothing, because a hung program is not burning CPU. The goroutine dump is the tool for this failure, and it is almost always sufficient on its own.

  • The dump shows the counted goroutines still alive and blocked. Why is that not a WaitGroup bug?
    Because the arithmetic is correct: each live goroutine still owes exactly one `Done`, and it will run when the goroutine finishes. `Wait` is faithfully reporting that the work is not done. The bug is wherever those goroutines are parked — a send on a channel nobody reads, a lock held elsewhere, a request with no deadline. Fix that, and the join resolves itself.
  • How do you get the same evidence from a production service you cannot restart?
    Register `net/http/pprof` and fetch `/debug/pprof/goroutine?debug=2`, which returns full stacks for every goroutine with no downtime. Go 1.27 additionally serves a goroutine-leak profile at `/debug/pprof/goroutineleak` that isolates goroutines which can no longer make progress, which is far more readable when the process holds thousands of them.
  • Why is placing wg.Add(1) on the line immediately before the go statement a rule rather than a style preference?
    Because any statement between them can skip the launch — a `continue`, an early `return`, an error check, a filter — leaving an increment that nothing will ever match, and the symptom is a hang far from the cause. Keeping them adjacent makes the pairing locally verifiable by eye. Go 1.25's `wg.Go` enforces it structurally.
  • Would running the tests with -race help find this?
    No. The race detector reports unsynchronised concurrent memory access that actually occurred during the run; it says nothing about a counter that is correctly synchronised and simply never reaches zero. A hung program also burns no CPU, so a CPU profile is empty. The goroutine stack dump is the right instrument, and usually the only one needed.

saying these in an interview costs you the question

  • Reaches for the race detector to explain a hang
  • Adds a sleep or a shorter timeout instead of reading the dump
  • Never lowers -timeout and waits ten minutes for evidence
  • Removes the Wait call to make the test pass
  • Blames the scheduler rather than counting Add against Done
  • Ignores that the counted goroutines are still alive in the dump