What makes a Go program panic with "sync: negative WaitGroup counter", and how do you avoid it?
answer
- Done is just Add of minus one
- the counter may never go below zero
- count the decrements on every path
- a defer plus an explicit call is two
basics
~20 sDone is Add(-1), and the runtime panics the moment a decrement would take a WaitGroup's counter below zero. It means more Done calls ran than Add calls — usually a deferred Done plus a second one on an error path.
solid answer
~50 s`wg.Done()` is exactly `wg.Add(-1)`, and `sync.WaitGroup` panics with `sync: negative WaitGroup counter` as soon as a decrement would take the counter below zero. So the panic always means Done ran more times than Add did. The usual causes are a `defer wg.Done()` at the top of a function that also calls `wg.Done()` explicitly on an error path, a helper that calls Done for a group its caller has already deferred Done on, or `wg.Add(len(units))` before a loop that then also does `wg.Add(1)` per unit and one Done per unit plus a stray. The panic fires in whichever goroutine called Done, and since an unrecovered panic in any goroutine takes the whole process down, it is a crash rather than a failed unit of work. The discipline that prevents it: exactly one Add per launch, exactly one deferred Done as the goroutine's first statement, and Done called nowhere else. Go 1.25's `wg.Go` makes the pairing automatic.
code
go · 8 linesfunc applyUnit(wg *sync.WaitGroup, m Migration) error {
defer wg.Done()
if err := m.Up(); err != nil {
wg.Done() // BUG: the deferred Done still runs
return err
}
return nil
}go deeper
Know that Done subtracts one and the counter cannot go below zero. If you see this panic, look for a Done that runs twice on the same path.
Explain that Done is Add(-1) and that the check fires at the decrement, then name the usual shapes: a defer plus an explicit call, or two layers both calling Done for the same launch.
Stress that it is an unrecoverable process crash raised in the worker goroutine, and give the convention that prevents it: the code that writes Add writes Done, helpers never touch the group.
Own the convention rather than the bug: pick one join idiom for the codebase, prefer the API that pairs Add and Done for you, and treat a hand-written Done outside a defer as a review blocker.
## What the panic means `sync.WaitGroup` holds a non-negative counter. `Done` is defined as `Add(-1)`, and `Add` checks the result: if the counter would become negative, it panics immediately with ``` panic: sync: negative WaitGroup counter ``` There is no ambiguity to diagnose about *what* happened — a `Done` ran without a matching `Add`. The work is finding *which* one. ## Where the extra Done comes from **A deferred Done plus an explicit one.** This is the overwhelmingly common shape, and it appears when someone adds an error path to a function that already had the defer: ```go func applyUnit(wg *sync.WaitGroup, m Migration) error { defer wg.Done() if err := m.Up(); err != nil { wg.Done() // BUG: the defer will fire too return err } return nil } ``` On the happy path the counter is decremented once. On the error path it is decremented twice, and the migration run crashes the first time any unit fails — which is exactly when you least want a crash instead of a reported error. **Two layers both taking responsibility.** The launcher writes `go func() { defer wg.Done(); applyUnit(&wg, m) }()` while `applyUnit` also calls `Done` internally. Each layer looks correct in isolation. This is the strongest argument for a single convention: the code that writes `Add` is the code that writes `Done`, and no helper touches the group. **Mismatched bulk and per-item Add.** `wg.Add(len(units))` before the loop *and* `wg.Add(1)` inside it is not the bug — that over-counts and hangs. The bug is the reverse: a bulk `Add(len(units))` where the loop skips some units but every launched goroutine plus some cleanup path calls `Done`. **Reuse across batches.** The group's counter returns to zero after `Wait`. If a straggler goroutine from batch one is still running and calls `Done` after batch two's counter has dropped to zero, that late `Done` panics. This is really a lifetime bug: the first batch was not fully joined before the second began. ## Why it is a crash, not an error The panic is raised in the goroutine that called `Done`. A panic that unwinds a goroutine without being recovered *in that same goroutine* terminates the entire program — no other goroutine, including the one blocked in `Wait`, can recover it, and there is no way to turn it into a returned error after the fact. So a negative-counter panic in a background worker takes down the whole process, printing that goroutine's stack. That stack is the good news: the top frames name the exact `Done` that went one too far. ## The related misuse panics Two neighbours produce different messages from the same type, and it is worth being able to tell them apart. Calling `Add` to raise the counter from zero while another goroutine is already blocked in `Wait` on that zero is a documented misuse and panics, as does reusing the group for a new round before the previous `Wait` has returned. Both are lifetime errors — the group is being driven for a second batch while the first is still being joined — rather than a simple double `Done`. ## The discipline that prevents all of it 1. **One `Add` per launch, in the launching goroutine, immediately before the `go` statement.** Bulk `Add(n)` is fine, but then nothing else may `Add` for that batch. 2. **One `defer wg.Done()`, as the very first statement of the goroutine's function body, and nowhere else.** If you are typing `wg.Done()` on a line that is not a `defer`, stop. 3. **Helpers do not touch the group.** They return; the launcher joins. 4. **Since Go 1.25, prefer `wg.Go(f)`.** It increments, runs `f`, and decrements when `f` returns, so there is no `Done` for anyone to write twice. It is the cleanest structural fix, not merely a shorthand. ## Finding it when it happens The panic output is a stack trace of the offending goroutine, so the frame below `sync.(*WaitGroup).Add` is the culprit function. From there, read every path through that function and count decrements. If the trace points at a deferred call, look for a second, non-deferred `Done` earlier in the same function or in something it called. Grepping the package for `Done()` and checking that each occurrence is preceded by `defer` finds most instances in seconds.
- Which goroutine does this panic kill, and can the goroutine blocked in Wait recover it?It is raised in the goroutine that called `Done`. Only a `recover` in a function deferred by *that* goroutine could stop it, and an unrecovered panic anywhere terminates the whole program. The goroutine sitting in `Wait` cannot see or recover it, and `Wait` has no error return, so this is a process crash rather than a failed unit of work.
- How does the over-counting mistake differ in symptom from the over-decrementing one?An extra `Add` with no matching `Done` leaves the counter above zero, so `Wait` simply never returns — a silent hang that surfaces as a test timeout. An extra `Done` drives the counter below zero and panics loudly and immediately. Same class of bookkeeping error, opposite failure modes, and the loud one is far easier to fix.
- Is bulk Add(len(units)) before the loop safe, or should it be Add(1) per iteration?Both are correct if used exclusively. `Add(len(units))` is fine when the loop is guaranteed to start exactly that many goroutines; if any iteration can skip the `go` statement — a filter, a `continue`, an early error — the counter over-counts and `Wait` hangs. `Add(1)` immediately before each `go` statement makes the pairing local and is harder to get wrong.
saying these in an interview costs you the question
- Thinks the panic only kills the goroutine that called Done
- Calls Done explicitly on an error path that already defers it
- Believes Wait returns an error when the counter goes negative
- Lets a helper call Done for a group its caller already deferred
- Expects recover in the waiting goroutine to catch it