skip to content

What makes a Go program panic with "sync: negative WaitGroup counter", and how do you avoid it?

level: middleimportance: should knowfreq 42%

answer

  1. Done is just Add of minus one
  2. the counter may never go below zero
  3. count the decrements on every path
  4. a defer plus an explicit call is two

basics

~20 s

Done is Add(-1), and the runtime panics the moment a decrement would take a WaitGroup's counter below zero. It means more Done calls ran than Add calls — usually a deferred Done plus a second one on an error path.

solid answer

~50 s

`wg.Done()` is exactly `wg.Add(-1)`, and `sync.WaitGroup` panics with `sync: negative WaitGroup counter` as soon as a decrement would take the counter below zero. So the panic always means Done ran more times than Add did. The usual causes are a `defer wg.Done()` at the top of a function that also calls `wg.Done()` explicitly on an error path, a helper that calls Done for a group its caller has already deferred Done on, or `wg.Add(len(units))` before a loop that then also does `wg.Add(1)` per unit and one Done per unit plus a stray. The panic fires in whichever goroutine called Done, and since an unrecovered panic in any goroutine takes the whole process down, it is a crash rather than a failed unit of work. The discipline that prevents it: exactly one Add per launch, exactly one deferred Done as the goroutine's first statement, and Done called nowhere else. Go 1.25's `wg.Go` makes the pairing automatic.

code

go · 8 lines
go
func applyUnit(wg *sync.WaitGroup, m Migration) error {
	defer wg.Done()
	if err := m.Up(); err != nil {
		wg.Done() // BUG: the deferred Done still runs
		return err
	}
	return nil
}

go deeper

for a junior

Know that Done subtracts one and the counter cannot go below zero. If you see this panic, look for a Done that runs twice on the same path.

for a middle

Explain that Done is Add(-1) and that the check fires at the decrement, then name the usual shapes: a defer plus an explicit call, or two layers both calling Done for the same launch.

for a senior

Stress that it is an unrecoverable process crash raised in the worker goroutine, and give the convention that prevents it: the code that writes Add writes Done, helpers never touch the group.

for a principal

Own the convention rather than the bug: pick one join idiom for the codebase, prefer the API that pairs Add and Done for you, and treat a hand-written Done outside a defer as a review blocker.

## What the panic means `sync.WaitGroup` holds a non-negative counter. `Done` is defined as `Add(-1)`, and `Add` checks the result: if the counter would become negative, it panics immediately with ``` panic: sync: negative WaitGroup counter ``` There is no ambiguity to diagnose about *what* happened — a `Done` ran without a matching `Add`. The work is finding *which* one. ## Where the extra Done comes from **A deferred Done plus an explicit one.** This is the overwhelmingly common shape, and it appears when someone adds an error path to a function that already had the defer: ```go func applyUnit(wg *sync.WaitGroup, m Migration) error { defer wg.Done() if err := m.Up(); err != nil { wg.Done() // BUG: the defer will fire too return err } return nil } ``` On the happy path the counter is decremented once. On the error path it is decremented twice, and the migration run crashes the first time any unit fails — which is exactly when you least want a crash instead of a reported error. **Two layers both taking responsibility.** The launcher writes `go func() { defer wg.Done(); applyUnit(&wg, m) }()` while `applyUnit` also calls `Done` internally. Each layer looks correct in isolation. This is the strongest argument for a single convention: the code that writes `Add` is the code that writes `Done`, and no helper touches the group. **Mismatched bulk and per-item Add.** `wg.Add(len(units))` before the loop *and* `wg.Add(1)` inside it is not the bug — that over-counts and hangs. The bug is the reverse: a bulk `Add(len(units))` where the loop skips some units but every launched goroutine plus some cleanup path calls `Done`. **Reuse across batches.** The group's counter returns to zero after `Wait`. If a straggler goroutine from batch one is still running and calls `Done` after batch two's counter has dropped to zero, that late `Done` panics. This is really a lifetime bug: the first batch was not fully joined before the second began. ## Why it is a crash, not an error The panic is raised in the goroutine that called `Done`. A panic that unwinds a goroutine without being recovered *in that same goroutine* terminates the entire program — no other goroutine, including the one blocked in `Wait`, can recover it, and there is no way to turn it into a returned error after the fact. So a negative-counter panic in a background worker takes down the whole process, printing that goroutine's stack. That stack is the good news: the top frames name the exact `Done` that went one too far. ## The related misuse panics Two neighbours produce different messages from the same type, and it is worth being able to tell them apart. Calling `Add` to raise the counter from zero while another goroutine is already blocked in `Wait` on that zero is a documented misuse and panics, as does reusing the group for a new round before the previous `Wait` has returned. Both are lifetime errors — the group is being driven for a second batch while the first is still being joined — rather than a simple double `Done`. ## The discipline that prevents all of it 1. **One `Add` per launch, in the launching goroutine, immediately before the `go` statement.** Bulk `Add(n)` is fine, but then nothing else may `Add` for that batch. 2. **One `defer wg.Done()`, as the very first statement of the goroutine's function body, and nowhere else.** If you are typing `wg.Done()` on a line that is not a `defer`, stop. 3. **Helpers do not touch the group.** They return; the launcher joins. 4. **Since Go 1.25, prefer `wg.Go(f)`.** It increments, runs `f`, and decrements when `f` returns, so there is no `Done` for anyone to write twice. It is the cleanest structural fix, not merely a shorthand. ## Finding it when it happens The panic output is a stack trace of the offending goroutine, so the frame below `sync.(*WaitGroup).Add` is the culprit function. From there, read every path through that function and count decrements. If the trace points at a deferred call, look for a second, non-deferred `Done` earlier in the same function or in something it called. Grepping the package for `Done()` and checking that each occurrence is preceded by `defer` finds most instances in seconds.

  • Which goroutine does this panic kill, and can the goroutine blocked in Wait recover it?
    It is raised in the goroutine that called `Done`. Only a `recover` in a function deferred by *that* goroutine could stop it, and an unrecovered panic anywhere terminates the whole program. The goroutine sitting in `Wait` cannot see or recover it, and `Wait` has no error return, so this is a process crash rather than a failed unit of work.
  • How does the over-counting mistake differ in symptom from the over-decrementing one?
    An extra `Add` with no matching `Done` leaves the counter above zero, so `Wait` simply never returns — a silent hang that surfaces as a test timeout. An extra `Done` drives the counter below zero and panics loudly and immediately. Same class of bookkeeping error, opposite failure modes, and the loud one is far easier to fix.
  • Is bulk Add(len(units)) before the loop safe, or should it be Add(1) per iteration?
    Both are correct if used exclusively. `Add(len(units))` is fine when the loop is guaranteed to start exactly that many goroutines; if any iteration can skip the `go` statement — a filter, a `continue`, an early error — the counter over-counts and `Wait` hangs. `Add(1)` immediately before each `go` statement makes the pairing local and is harder to get wrong.

saying these in an interview costs you the question

  • Thinks the panic only kills the goroutine that called Done
  • Calls Done explicitly on an error path that already defers it
  • Believes Wait returns an error when the counter goes negative
  • Lets a helper call Done for a group its caller already deferred
  • Expects recover in the waiting goroutine to catch it