Why is an errgroup.Group not a supervisor for long-lived workers that must be restarted?
answer
- the group joins, it does not manage
- Go runs the function exactly once
- Wait needs every function to have returned
- no hook fires when one worker exits
- hand the group a supervisor, not a session
basics
~20 sA group runs each function you hand to Go exactly once and joins them. Group.Wait returns after every one of them has returned, and nothing in the group restarts anything, so a worker that exits is simply gone.
solid answer
~50 s`errgroup.Group.Go` runs the function once; `Group.Wait` blocks until all of them have returned and yields the first non-nil error. With `errgroup.WithContext`, the derived context is cancelled the first time a function returns a non-nil error, or when `Wait` returns — whichever comes first. That is a one-shot join with fail-fast semantics, and there is no hook that fires when one function exits, so a long-lived session worker that dies leaves the group permanently one worker short while `Wait` keeps waiting for the rest. The fix is to change what you hand to the group: put the restart loop inside that function, so what the group joins is a *supervisor* per device that returns only when it gives up or shutdown starts. The group then does what it is good at — joining supervisors and surfacing the first fatal error.
code
go · 16 lines// The function handed to a group's Go method is the supervisor, not the
// session itself: it returns only once it has given up on this device.
func superviseDevice(ctx context.Context, dial func(context.Context) error) error {
for attempt := 1; ; attempt++ {
err := dial(ctx)
if ctx.Err() != nil {
return nil // clean shutdown, not a group-wide failure
}
if attempt >= maxRestarts {
return fmt.Errorf("device session gave up after %d restarts: %w", attempt, err)
}
if !waitBeforeRestart(ctx, attempt) { // capped backoff, cancellable
return nil
}
}
}go deeper
Know that handing a function to a group runs it once and that waiting on the group means waiting for all of them to return. Nothing restarts anything for you.
Explain the exact semantics: Go runs once, Wait joins all and reports the first non-nil error, and a derived context cancels on the first error or when Wait returns. Then say where a restart loop has to live.
Show the two failure modes in production — silent shrinkage when a worker returns nil, and one flaky peer tearing down every sibling when it returns an error — and how per-device supervisors below the group fix both.
Decide the contract for the whole service: what a supervisor's non-nil return is allowed to mean, that it should be rare enough to alert on, and that a group returning is the signal the process should exit.
## What a group actually promises `errgroup` (from the `golang.org/x/sync` quasi-standard repos) is a `sync.WaitGroup` that also carries an error and, optionally, a context. Its contract is small and worth stating exactly: - `Group.Go(f)` runs `f` in a new goroutine. **Once.** - `Group.Wait()` blocks until every function passed to `Go` has returned, then returns the first non-nil error any of them produced. - `errgroup.WithContext(parent)` returns a group and a derived context that is cancelled the first time a function passed to `Go` returns a non-nil error, or the first time `Wait` returns, whichever happens first. Every one of those clauses describes a *join*: the group's purpose is to let one place wait for a set of tasks and learn whether any of them failed. It is the structured version of "start five things, wait for all five, report the first failure". ## Why that is the wrong shape for a long-lived worker A session worker managing a socket to a remote device is not a task that finishes. Its normal life is minutes to weeks, and its return is an *event* — the device rebooted, the peer closed the connection, a read timed out — not a result. A group has no vocabulary for that event. Nothing in the group runs when a function returns except the bookkeeping that decrements its internal counter. So the two failure modes are: **Silent shrinkage.** A worker returns `nil` after its connection closed. No error, no cancellation, no restart. The process keeps running, `Wait` keeps waiting for the other workers, and that device is simply no longer being managed. Nothing anywhere says so. This is the more dangerous of the two, because the process looks perfectly healthy. **All-or-nothing failure.** A worker returns a non-nil error. Now the derived context is cancelled, so every sibling that watches the context unwinds too, and `Wait` returns that first error. One flaky device has torn down every other session. Fail-fast is exactly right for a startup sequence or a fan-out where partial results are useless; it is exactly wrong when the workers are independent and one of them is allowed to be sick. ## The fix: supervise inside the function you hand to Go The group is not the problem — the granularity is. Hand it a function that *is* the supervisor for one device: ```go g.Go(func() error { return superviseDevice(ctx, dialer(dev)) }) ``` where `superviseDevice` loops: run the session, look at why it ended, wait out a backoff, run it again. It returns only in two situations, and both of them are decisions rather than accidents: - **Shutdown.** The context is done; return `nil`, because an orderly stop is not a group-wide failure and returning an error here would make `Wait` report cancellation as the cause of death. - **Giving up.** The session has failed too many times in a row to be worth retrying. Return a wrapped error. Now the fail-fast behaviour is what you want: this is genuinely fatal, the derived context cancels the other supervisors, `Wait` returns your error, and `main` can exit non-zero. With that split, each layer does one job. The supervisor owns "is this worker allowed to fail again?". The group owns "are all supervisors still alive, and did any of them declare defeat?". ## Consequences to expect `Wait` returning is now a meaningful, rare event — it means either the whole service is shutting down or something decided the process should die. That makes it a good place to hang the last log line before `main` returns. Shutdown ordering also gets simpler: cancelling the root context makes every supervisor stop relaunching and return `nil`, and `Wait` becomes the single place that blocks until all of them have finished unwinding. And because a supervisor returns `nil` on cancellation, you keep the property that a non-nil result from `Wait` always means something abnormal happened — which is the property that makes it worth alerting on. ## The plain-WaitGroup version None of this depends on `errgroup`. A `sync.WaitGroup` plus a channel of errors gets you the same structure with more code: the point is that the *group* joins supervisors, never sessions. If your restart loop is missing, no join primitive will supply one.
- Under errgroup.WithContext, what happens to the other workers when one of them returns a non-nil error?The derived context is cancelled, so every sibling watching that context unwinds, and `Wait` returns that first error once all of them have finished. That fail-fast behaviour is right for a startup sequence, and wrong for independent device sessions where one bad peer should not take the other forty down.
- Should a per-device supervisor return the context's error when shutdown cancels it?No — return nil. Cancellation is the orderly path, and returning `ctx.Err()` makes `Group.Wait` report `context.Canceled` as the service's cause of death, drowning any real error that triggered the shutdown. Reserve a non-nil return for the case where the supervisor genuinely gave up on the device.
- A worker under a group returns nil after its connection closes. What does the group do?Nothing at all. The function has returned, so the group's counter drops and `Wait` goes on waiting for the remaining workers. That device is no longer managed and no signal says so, which is why a restart loop, or at minimum a per-worker liveness metric, has to exist below the group.
saying these in an interview costs you the question
- Believes a group restarts a function that returns
- Expects Wait to return as soon as one worker exits
- Thinks a worker returning nil cancels its siblings
- Returns ctx.Err() from a supervisor on clean shutdown
- Puts long-lived sessions directly in Go with no restart loop