With errgroup.WithContext, what cancels the derived context, and why might siblings keep running anyway?
answer
- cancellation is a signal, not a stop
- three triggers, one of them is Wait
- the task must be looking
- capturing the wrong ctx is the classic bug
- Wait still waits for the uninterruptible task
basics
~20 sThe context from errgroup.WithContext is cancelled when the first task returns a non-nil error, and when Group.Wait returns. Cancellation is only a signal: a task that never checks ctx.Done() or passes the context down runs to completion.
solid answer
~40 s`errgroup.WithContext(parent)` returns a `*Group` and a child context. That child is cancelled the first time a function passed to `Go` returns a non-nil error — and also when `Wait` returns, so the group cleans up after itself. But cancelling a `context.Context` in Go does not stop a goroutine; it only closes `ctx.Done()` and sets `ctx.Err()` to `context.Canceled`. A sibling only reacts if it passes that same context into the calls it makes, or selects on `ctx.Done()` in its loop. Two bugs follow from that: a task that captures the *parent* context instead of the derived one is never cancelled at all, and a task doing pure CPU work with no cancellation check will run to the end while `Wait` blocks on it. Cancellation buys you nothing beyond what your tasks choose to honour.
go deeper
Recall the call shape: g, ctx := errgroup.WithContext(parent), and that tasks must use that ctx. Know that cancelling a context closes ctx.Done() rather than killing a goroutine.
Name all three cancellation triggers — first task error, Wait returning, parent cancelled — and explain why a CPU-bound task with no ctx.Done() case makes the group as slow as its slowest branch.
Show how you would find the capture bug in review or in production: which branch never aborts, what the latency histogram looks like, and how you separate context.Canceled noise from the one real failure in the logs.
Own the convention: whether your codebase allows two context variables in one function, whether every internal client must accept a context, and what you will not merge without a cancellation path.
## The two things WithContext returns ``` g, ctx := errgroup.WithContext(parent) ``` `g` is a `*errgroup.Group` and `ctx` is a context derived from `parent`. From that point on, the rule is: **every task in the group uses `ctx`, not `parent`.** ## When the derived context is cancelled There are three triggers, and it is worth being able to name all three: 1. **A task returns a non-nil error.** The first one to do so cancels `ctx`. 2. **`Wait` returns.** The group cancels its own context on the way out, whether or not anything failed, so you cannot leak the cancellation. 3. **The parent is cancelled.** `ctx` is a child, so a parent deadline expiring or a caller cancelling upstream propagates down as it would for any derived context. Trigger 2 has a practical consequence people trip over: the context is dead after `Wait`. If you write `g.Wait()` and then use `ctx` for one more call — say, to write the assembled response — that call fails immediately with `context.Canceled`. Anything that outlives the group needs its own context. ## Why cancellation does not stop anything Go has no way to kill a goroutine from outside. Cancelling a context closes the channel returned by `ctx.Done()` and makes `ctx.Err()` return `context.Canceled`. That is the entire mechanism. A goroutine notices only if it looks, and there are exactly two ways to look: - **Pass the context down.** Anything that takes a `context.Context` — an HTTP request built with `http.NewRequestWithContext`, a database query, a timer via `context.AfterFunc` — will abort itself when the context is cancelled. This covers most real tasks, because most real tasks are waiting on I/O. - **Select on `ctx.Done()`.** A loop that computes, or that reads from a channel, needs an explicit case: ``` select { case <-ctx.Done(): return ctx.Err() case v := <-in: // ... } ``` A task that does neither is uninterruptible. The group's context flips to cancelled, the task keeps running, and `Wait` keeps blocking, because `Wait`'s contract is *every function passed to `Go` has returned* — not *everything has been told to stop*. A fan-out whose slowest branch ignores cancellation is exactly as slow as it was before you introduced the group. ## The capture bug The most common real defect is shadowing gone wrong: ``` g, gctx := errgroup.WithContext(ctx) g.Go(func() error { return callA(ctx) }) // BUG: parent, not gctx g.Go(func() error { return callB(gctx) }) ``` `callA` is now immune to the group's cancellation. This reads as a typo, survives review easily, and is invisible in tests where nothing fails. The usual defence is to shadow the name deliberately — `ctx, cancel := ...` style — so there is only one `ctx` in scope for the tasks to capture; some teams simply forbid two context variables in the same function. ## What returning early actually saves When cancellation is honoured, the win is concrete: in-flight upstream calls are abandoned, their connections released, the work they would have done is never billed, and `Wait` returns close to the time of the first failure instead of the time of the slowest branch. In a fan-out of five backends where one fails at 30 ms and the slowest takes 2 s, honouring cancellation turns a 2 s failed request into a 30 ms one. ## The noise it creates Every cancelled sibling now returns an error too, and it is almost always `context.Canceled` wrapped by whatever library it was calling. These are not independent faults — they are the consequence of the first failure. If a task logs its own error before returning, filter with `errors.Is(err, context.Canceled)` so your logs show one root cause rather than five. `Wait` itself is unaffected: it already returned the first real error, because that error is what triggered the cancellation in the first place. ## The short version The derived context is cancelled on first error, on `Wait`'s return, or with the parent. Cancellation is advertising, not enforcement: the tasks have to be written to notice, and the ones that are not will keep the group alive until they finish on their own.
- You call Group.Wait, then use the context from errgroup.WithContext for one final call. What happens?The call fails immediately with `context.Canceled`. The group cancels its derived context as `Wait` returns, so that context is only valid for the duration of the group. Anything that runs after the join — writing the response, a cleanup call, a fire-and-forget metric — needs a context derived from the original parent, or `context.WithoutCancel` of it if you deliberately want the work to outlive the request.
- A sibling cancelled by the group returns context.Canceled. Should that error be logged as a failure?No — it is downstream noise from the first real failure, not an independent fault. If tasks log their own errors, guard with `errors.Is(err, context.Canceled)` and drop or downgrade those entries, so one incident produces one root-cause line instead of five. `Group.Wait` already returns the real error, because that error is what cancelled the context in the first place.
- Does errgroup.WithContext give you a deadline as well?No. It derives a cancellable context only; there is no timeout in errgroup. If the fan-out needs a bound, wrap the parent first with `context.WithTimeout` and pass that into `errgroup.WithContext`, remembering to `defer cancel()` on the timeout context. The group then inherits the deadline and its cancellation stacks on top of it.
saying these in an interview costs you the question
- Saying cancellation terminates the sibling goroutines
- Believing Wait returns as soon as the context is cancelled
- Capturing the parent context inside a group task
- Assuming WithContext adds a timeout of its own
- Reusing the derived context after Wait has returned