skip to content

What does golang.org/x/sync/errgroup add over a sync.WaitGroup for concurrent tasks that can fail?

level: juniorimportance: must knowfreq 72%

answer

  1. one of them counts, one of them reports
  2. the task signature is the difference
  3. func() versus func() error
  4. join point that can also say why
  5. first non-nil error plus sibling cancellation

basics

~20 s

errgroup joins goroutines like sync.WaitGroup, but each task is a func() error. Group.Wait blocks until every task has returned and gives back the first non-nil error, and errgroup.WithContext also cancels the siblings when one task fails.

solid answer

~40 s

A `sync.WaitGroup` only counts goroutines. It has no place to put a failure, so you end up wiring a buffered error channel next to it, draining that channel after `Wait`, and deciding yourself which error to report. `errgroup.Group` folds all of that in: `g.Go(func() error { ... })` starts a task, and `g.Wait()` blocks until every task has returned and returns the first non-nil error one of them produced. If you build the group with `errgroup.WithContext(ctx)` you also get back a derived `context.Context` that is cancelled as soon as a task fails, so the siblings can stop instead of finishing work nobody will use. `SetLimit` bounds how many tasks run at once. The cost is that only one error survives `Wait` — if you need all of them you have to record them yourself.

code

go · 17 lines
go
var wg sync.WaitGroup
errs := make(chan error, len(urls)) // buffered: a failing task must never block

for _, u := range urls {
	wg.Go(func() { // Go 1.25+; before that: wg.Add(1) and defer wg.Done()
		if err := fetch(u); err != nil {
			errs <- err
		}
	})
}

wg.Wait()
close(errs)

for err := range errs { // and now you decide what to do with them
	log.Println(err)
}

go deeper

for a junior

Be ready to state the two method names and what they take: Group.Go takes a func() error, Group.Wait blocks and returns the first non-nil error. Then say the one thing a WaitGroup cannot do — report a failure.

for a middle

Explain the boilerplate errgroup replaces line by line: the buffered error channel, its size, the close after Wait, and the manual cancel func. Note that Go 1.25's WaitGroup.Go removes the Add/Done bug but not the missing error path.

for a senior

Show judgment about when the join needs an answer at all. Point out that only one error survives Wait, that cancellation is cooperative, and that tasks writing into distinct slots need no mutex while a shared append does.

for a principal

Frame it as a build-versus-depend call: twenty lines of join logic your team owns and tests, against a line in go.mod that lands in every importer's module graph. Say who in your organisation gets to make that call.

## What a WaitGroup actually gives you `sync.WaitGroup` is a counter with a barrier. You increment it before starting a goroutine, each goroutine decrements it when it finishes, and `Wait` blocks until the counter reaches zero. That is the whole contract. It answers exactly one question — *are they all done?* — and it has no channel through which a goroutine can say *and this one failed*. So the moment your concurrent tasks can fail, a `WaitGroup` alone stops being enough, and you write the same boilerplate every time: - a buffered error channel, sized to the number of tasks so a failing goroutine can always send without blocking (an unbuffered one deadlocks if nobody is receiving yet); - a `close` after `Wait` so the channel can be ranged over; - a drain loop that picks which of the collected errors to return; - and, if you want the siblings to stop when one fails, a `context.CancelFunc` you create, capture and call by hand. That is roughly twenty lines that are easy to get subtly wrong and that you now own and must test. ## What errgroup.Group is `errgroup.Group` from `golang.org/x/sync/errgroup` is that boilerplate packaged. Its zero value is usable, and it has a small surface: - `g.Go(f func() error)` runs `f` in a new goroutine and remembers its error. - `g.Wait() error` blocks until every function passed to `Go` has returned, then returns the first non-nil error any of them produced (or `nil` if none failed). - `errgroup.WithContext(ctx)` returns a `*Group` and a derived `context.Context`; that context is cancelled the first time a task returns a non-nil error, and also once `Wait` returns. - `g.SetLimit(n)` caps how many tasks run concurrently; `g.TryGo(f)` starts a task only if the group is currently below that cap. The crucial difference is the task signature. A `WaitGroup` goroutine is a `func()` — it has nowhere to put a result. An errgroup task is a `func() error`, so the group can do something with failure: record it, and (with `WithContext`) cancel everyone else. ## Note on sync.WaitGroup.Go Go 1.25 added `WaitGroup.Go`, which starts a goroutine and handles the `Add`/`Done` pairing for you. That removes the most common `WaitGroup` bug — an `Add` that happens inside the goroutine instead of before it, or a missing `Done` on an early return — but it does **not** close the gap: `WaitGroup.Go` still takes a `func()` with no error, and `WaitGroup.Wait` still returns nothing. Error collection and sibling cancellation remain yours to build. ## Where each one belongs Use a `WaitGroup` when the goroutines cannot meaningfully fail, or when their results already flow somewhere else — into a channel a consumer is reading, into a shared metric, into distinct slots of a slice. Use an errgroup when the join point genuinely needs an answer: *did all of this work succeed, and if not, stop the rest and tell me.* The canonical fit is a request handler that fans out to several independent backends and cannot render its response unless all of them answer. The group starts one task per backend, each task writes its own piece into a distinct field or slice index (distinct destinations, so no mutex is needed), and `Wait` collapses the whole fan-out into a single `if err != nil` at the call site. ## What errgroup does not do for you - **It keeps only one error.** The first task to return a non-nil error wins; the others are discarded. If the failures are diagnostically interesting, log or record each one inside its own task before returning. - **Cancellation is cooperative.** The context that `WithContext` hands back is only useful if the tasks actually pass it to the calls they make or select on `ctx.Done()`. A task that ignores it runs to completion, and `Wait` waits for it. - **The context dies when Wait returns.** Do not keep using the derived context after the join — derive a fresh one for anything that outlives the group. - **It is a dependency.** `golang.org/x/sync` is maintained by the Go team and has no third-party transitive requirements, but it is still a line in `go.mod` and a decision somebody owns. ## The shape to remember A `WaitGroup` answers "are they finished?". An errgroup answers "are they finished, did any of them fail, and should the rest keep going?" — which is what most real fan-outs actually need.

  • Is the zero value of errgroup.Group usable, or must you always call errgroup.WithContext?
    The zero value works: `var g errgroup.Group`, then `g.Go(...)` and `g.Wait()`. What you lose is cancellation — without `WithContext` there is no derived context, so a failure in one task does nothing to the others and `Wait` simply waits for all of them before returning the first error. Use the zero value when the tasks are short and independent; use `WithContext` when abandoning the rest is worth something.
  • Can several errgroup tasks write into the same result struct without a mutex?
    Yes, provided each task writes to a *different* field or a different slice index and nobody reads those locations until after `Wait`. Distinct memory locations are not a data race, and `Wait` gives the happens-before edge that makes the writes visible to the joining goroutine. Two tasks appending to the same slice, or writing the same field, is a race and needs a mutex — or a redesign into distinct slots.
  • What is wrong with using an unbuffered error channel in the hand-rolled version?
    If nothing is receiving yet, the first goroutine that tries to send its error blocks forever, its `Done` never runs, and `Wait` never returns — a deadlock that only appears on the failure path, which is exactly the path least covered by tests. Buffering the channel to the number of tasks guarantees every sender can deposit its error and exit.

A WaitGroup is a turnstile counter: it only tells you everyone has left the building. An errgroup is a shift supervisor: it tells you everyone has left, hands you the first incident report, and can call the rest back in when one worker hits trouble.

saying these in an interview costs you the question

  • Claiming sync.WaitGroup.Wait returns an error
  • Saying errgroup returns all the errors from Wait
  • Thinking errgroup cancels siblings without WithContext
  • Passing a func() with no error return to Group.Go
  • Using an unbuffered error channel in the hand-rolled join