Why does errgroup.Group.Wait report only one failure when a fan-out to five upstreams has two of them fail?
answer
- the group remembers exactly one
- first to return, not first to start
- which one you get is a race
- log at the source, not at the join
- distinct slots plus errors.Join
basics
~20 serrgroup.Group.Wait keeps only the first non-nil error a task returned and discards the rest, so a second failure never reaches the caller. Log each failure inside its own task, or collect per-task errors and combine them with errors.Join.
solid answer
~50 s`errgroup.Group` stores exactly one error: the first non-nil one a task returns wins, and every later error is dropped on the floor. "First" means first to *return*, not first started, so which failure you see is a race and can differ between requests. For a page assembler calling five upstreams, that means an incident where two of them were down looks in your logs like one upstream being down — and you page the wrong team. The fix is not to abandon errgroup, whose join and cancellation you still want, but to stop treating `Wait`'s error as the record of what happened: have each task log or record its own failure, annotated with which upstream it was, before returning it. If the caller needs all of them, write into a per-task slice slot and combine with `errors.Join` after `Wait`, keeping `Wait`'s error as the "did anything fail" signal.
code
go · 14 lineserrs := make([]error, len(sources))
var wg sync.WaitGroup
for i, src := range sources {
wg.Go(func() { // each goroutine writes its own index: no data race
results[i], errs[i] = load(ctx, src)
})
}
wg.Wait()
if err := errors.Join(errs...); err != nil {
// every failure, not just whichever one returned first
return err
}go deeper
Remember the rule: one error survives Group.Wait and the others are discarded. Know that errors.Join exists for combining several error values into one.
Explain that the surviving error is the first to be returned, so it is decided by latency rather than by task order, and show the per-index slot pattern that needs no mutex.
Walk through the production consequence — wrong attribution, wrong retry decision, cancelled siblings as noise — and the log-at-the-source-versus-log-at-the-join comparison that exposes it.
Set the standard: what a fan-out is required to record before it returns, whether partial success is a first-class result type in your API, and which team owns the alert when two dependencies fail at once.
## The mechanism `errgroup.Group` records a failure exactly once. The first task whose function returns a non-nil error has that error stored; every subsequent non-nil error is discarded, and `Wait` returns the stored one. There is no accumulation, no joining, no count. Two consequences follow, and the second is the one that hurts in production: 1. **You lose the other errors entirely.** They never leave the goroutine that produced them. 2. **Which one you keep is nondeterministic.** It is decided by return order, which for network calls is decided by latency, retries and timeouts. The same outage can present as upstream A on one request and upstream B on the next. ## Why it shows up in a page assembler Consider a service that assembles one response from five upstreams — profile, preferences, entitlements, recommendations, banner — each fetched by one task in a group created with `errgroup.WithContext`, each writing into its own field of a response struct. When two upstreams are unhealthy, the handler returns one error. Whatever you build on top of that single error is now wrong in a specific, expensive way: - **Attribution.** The alert names one dependency. The other one is invisible until someone reads its own service's dashboards. - **Retry policy.** If you decide retryability from the returned error and the surviving error happens to be a 500 from a service that is fine on retry, you retry into a second dependency that is hard down. - **Partial responses.** If your product allows degrading — render the page without recommendations — a single error tells you nothing about *which* sections are missing, so you cannot make that call at the join point. And there is a fourth: with `WithContext`, the first failure cancels the siblings, so several of the tasks now return `context.Canceled` too. Those are consequences, not causes. If you log them undifferentiated, one incident produces five error lines with the real cause somewhere among them. ## The diagnostic that reveals it The cheap experiment is to log inside each task and compare that against what `Wait` returned. Have every task, on failure, emit one line naming its upstream and its error before returning, then log `Wait`'s error once at the join. Under a real double failure you will see two (or five, with cancellations) task-level lines and exactly one line at the join. That mismatch is the whole bug, and it is invisible if you only ever log at the join. ## What to do instead **Keep the group; change what carries the errors.** - *Log at the source.* Each task annotates and logs its own error before returning it. Cheapest fix, requires no structural change, and gives operators complete information even though the caller still sees one error. Filter `errors.Is(err, context.Canceled)` so cancelled siblings do not drown the root cause. - *Per-task slots plus `errors.Join`.* Give each task a distinct slice index to write its error into — distinct locations, so no mutex is needed — and after `Wait` build the combined error with `errors.Join`, which ignores nil entries and returns nil when everything succeeded. `errors.Is` and `errors.As` both see through a joined error, so caller-side classification still works. - *Model partial success explicitly.* If the response is allowed to degrade, the task's job is not to return an error at all but to record a per-section outcome in the result struct. The group then fails only on errors that must fail the whole request, and `Wait`'s single error is correct again because there is genuinely only one class of fatal failure. **Do not** wrap each task so it always returns nil just to keep the group from cancelling — you throw away the cancellation you wanted. And do not add a mutex-guarded shared error slice when distinct slots would do; it is more code and more contention for no gain. ## Where errors.Join fits `errors.Join(errs...)` builds one error wrapping several, formatting them one per line, and returns nil if every argument is nil — so you can hand it the whole slice unconditionally. That makes the collect-then-join pattern about three lines on top of the group, which is why it is the default answer once someone has been burned by a silent second failure. ## The judgment to show An interviewer is checking whether you know that `Wait`'s single error is a *control-flow* signal, not an *observability* record. Use it to decide what the handler does; never let it be the only place a failure is written down.
- Between two failing tasks, which error does errgroup.Group.Wait actually keep?Whichever function returned its error first in wall-clock terms — not the first task started, not the first to fail internally. For network calls that ordering follows latency and retry behaviour, so it is effectively nondeterministic and can differ request to request. Any alerting or retry logic that assumes a stable answer will behave inconsistently under a multi-dependency outage.
- Why not have each task return nil and record failures only in the result struct?Because returning nil tells the group nothing failed, so `errgroup.WithContext` never cancels the siblings and `Wait` reports success. That is the right design only when partial success is genuinely acceptable and you inspect the per-section outcomes at the join. If some failures must abort the request, those must still be returned as errors so cancellation and the caller's error path both work.
- How do you keep cancelled siblings from cluttering the log with context.Canceled?Check `errors.Is(err, context.Canceled)` in the task's own logging path and drop or downgrade those entries to debug. Under `errgroup.WithContext` the first genuine failure cancels everyone else, so those errors are effects of the incident rather than separate faults. Keeping them at debug level preserves them for a deep dive without burying the one error that actually explains the request.
It is a hospital triage desk that records only the first patient through the door. The queue behind them still exists; nothing about it is written down, so the shift report describes a quiet night.
saying these in an interview costs you the question
- Claiming Group.Wait aggregates every task error
- Assuming the returned error comes from the first task started
- Alerting on Wait's error as the complete failure record
- Returning nil from tasks to keep siblings from being cancelled
- Guarding a shared error slice with a mutex when slots would do