skip to content

With errgroup.WithContext, what cancels the derived context, and why might siblings keep running anyway?

level: middleimportance: must knowfreq 66%

answer

  1. cancellation is a signal, not a stop
  2. three triggers, one of them is Wait
  3. the task must be looking
  4. capturing the wrong ctx is the classic bug
  5. Wait still waits for the uninterruptible task

basics

~20 s

The context from errgroup.WithContext is cancelled when the first task returns a non-nil error, and when Group.Wait returns. Cancellation is only a signal: a task that never checks ctx.Done() or passes the context down runs to completion.

solid answer

~40 s

`errgroup.WithContext(parent)` returns a `*Group` and a child context. That child is cancelled the first time a function passed to `Go` returns a non-nil error — and also when `Wait` returns, so the group cleans up after itself. But cancelling a `context.Context` in Go does not stop a goroutine; it only closes `ctx.Done()` and sets `ctx.Err()` to `context.Canceled`. A sibling only reacts if it passes that same context into the calls it makes, or selects on `ctx.Done()` in its loop. Two bugs follow from that: a task that captures the *parent* context instead of the derived one is never cancelled at all, and a task doing pure CPU work with no cancellation check will run to the end while `Wait` blocks on it. Cancellation buys you nothing beyond what your tasks choose to honour.

go deeper

for a junior

Recall the call shape: g, ctx := errgroup.WithContext(parent), and that tasks must use that ctx. Know that cancelling a context closes ctx.Done() rather than killing a goroutine.

for a middle

Name all three cancellation triggers — first task error, Wait returning, parent cancelled — and explain why a CPU-bound task with no ctx.Done() case makes the group as slow as its slowest branch.

for a senior

Show how you would find the capture bug in review or in production: which branch never aborts, what the latency histogram looks like, and how you separate context.Canceled noise from the one real failure in the logs.

for a principal

Own the convention: whether your codebase allows two context variables in one function, whether every internal client must accept a context, and what you will not merge without a cancellation path.

## The two things WithContext returns ``` g, ctx := errgroup.WithContext(parent) ``` `g` is a `*errgroup.Group` and `ctx` is a context derived from `parent`. From that point on, the rule is: **every task in the group uses `ctx`, not `parent`.** ## When the derived context is cancelled There are three triggers, and it is worth being able to name all three: 1. **A task returns a non-nil error.** The first one to do so cancels `ctx`. 2. **`Wait` returns.** The group cancels its own context on the way out, whether or not anything failed, so you cannot leak the cancellation. 3. **The parent is cancelled.** `ctx` is a child, so a parent deadline expiring or a caller cancelling upstream propagates down as it would for any derived context. Trigger 2 has a practical consequence people trip over: the context is dead after `Wait`. If you write `g.Wait()` and then use `ctx` for one more call — say, to write the assembled response — that call fails immediately with `context.Canceled`. Anything that outlives the group needs its own context. ## Why cancellation does not stop anything Go has no way to kill a goroutine from outside. Cancelling a context closes the channel returned by `ctx.Done()` and makes `ctx.Err()` return `context.Canceled`. That is the entire mechanism. A goroutine notices only if it looks, and there are exactly two ways to look: - **Pass the context down.** Anything that takes a `context.Context` — an HTTP request built with `http.NewRequestWithContext`, a database query, a timer via `context.AfterFunc` — will abort itself when the context is cancelled. This covers most real tasks, because most real tasks are waiting on I/O. - **Select on `ctx.Done()`.** A loop that computes, or that reads from a channel, needs an explicit case: ``` select { case <-ctx.Done(): return ctx.Err() case v := <-in: // ... } ``` A task that does neither is uninterruptible. The group's context flips to cancelled, the task keeps running, and `Wait` keeps blocking, because `Wait`'s contract is *every function passed to `Go` has returned* — not *everything has been told to stop*. A fan-out whose slowest branch ignores cancellation is exactly as slow as it was before you introduced the group. ## The capture bug The most common real defect is shadowing gone wrong: ``` g, gctx := errgroup.WithContext(ctx) g.Go(func() error { return callA(ctx) }) // BUG: parent, not gctx g.Go(func() error { return callB(gctx) }) ``` `callA` is now immune to the group's cancellation. This reads as a typo, survives review easily, and is invisible in tests where nothing fails. The usual defence is to shadow the name deliberately — `ctx, cancel := ...` style — so there is only one `ctx` in scope for the tasks to capture; some teams simply forbid two context variables in the same function. ## What returning early actually saves When cancellation is honoured, the win is concrete: in-flight upstream calls are abandoned, their connections released, the work they would have done is never billed, and `Wait` returns close to the time of the first failure instead of the time of the slowest branch. In a fan-out of five backends where one fails at 30 ms and the slowest takes 2 s, honouring cancellation turns a 2 s failed request into a 30 ms one. ## The noise it creates Every cancelled sibling now returns an error too, and it is almost always `context.Canceled` wrapped by whatever library it was calling. These are not independent faults — they are the consequence of the first failure. If a task logs its own error before returning, filter with `errors.Is(err, context.Canceled)` so your logs show one root cause rather than five. `Wait` itself is unaffected: it already returned the first real error, because that error is what triggered the cancellation in the first place. ## The short version The derived context is cancelled on first error, on `Wait`'s return, or with the parent. Cancellation is advertising, not enforcement: the tasks have to be written to notice, and the ones that are not will keep the group alive until they finish on their own.

  • You call Group.Wait, then use the context from errgroup.WithContext for one final call. What happens?
    The call fails immediately with `context.Canceled`. The group cancels its derived context as `Wait` returns, so that context is only valid for the duration of the group. Anything that runs after the join — writing the response, a cleanup call, a fire-and-forget metric — needs a context derived from the original parent, or `context.WithoutCancel` of it if you deliberately want the work to outlive the request.
  • A sibling cancelled by the group returns context.Canceled. Should that error be logged as a failure?
    No — it is downstream noise from the first real failure, not an independent fault. If tasks log their own errors, guard with `errors.Is(err, context.Canceled)` and drop or downgrade those entries, so one incident produces one root-cause line instead of five. `Group.Wait` already returns the real error, because that error is what cancelled the context in the first place.
  • Does errgroup.WithContext give you a deadline as well?
    No. It derives a cancellable context only; there is no timeout in errgroup. If the fan-out needs a bound, wrap the parent first with `context.WithTimeout` and pass that into `errgroup.WithContext`, remembering to `defer cancel()` on the timeout context. The group then inherits the deadline and its cancellation stacks on top of it.

saying these in an interview costs you the question

  • Saying cancellation terminates the sibling goroutines
  • Believing Wait returns as soon as the context is cancelled
  • Capturing the parent context inside a group task
  • Assuming WithContext adds a timeout of its own
  • Reusing the derived context after Wait has returned