skip to content

In a worker loop, how do you observe cancellation of a context.Context, and what do you return?

level: middleimportance: must knowfreq 70%

answer

  1. one channel, one error, both on the context
  2. the channel is closed, never sent to
  3. put it in the same select as the work
  4. return the reason, not nil
  5. a pure CPU loop has to poll instead

basics

~20 s

Put a case <-ctx.Done() in the select alongside the loop's other channel operations, and return ctx.Err(), which is context.Canceled once cancel has been called. Cancellation is advisory: code that never looks at Done keeps running.

solid answer

~50 s

`ctx.Done()` returns a `<-chan struct{}` that is closed when the context is cancelled, so the loop puts `case <-ctx.Done():` in the same `select` as its real work and returns `ctx.Err()` from that branch. `Err()` is `nil` while the context is live and `context.Canceled` after cancellation, so returning it gives the caller a typed reason rather than a bare `nil`. Nothing is preempted: a branch of code that never touches `Done()` or `Err()` runs to completion regardless, which is why a CPU-bound loop with no channel operations has to poll `ctx.Err() != nil` itself at a sensible granularity. For a non-blocking check, a `select` with `case <-ctx.Done():` and a `default:` works. And a context that can never be cancelled, such as `context.Background()`, returns a nil `Done()` channel, which simply blocks forever in a `select` — the case is inert rather than broken.

code

go · 13 lines
go
func worker(ctx context.Context, chunks <-chan []byte) error {
	for {
		select {
		case <-ctx.Done():
			return ctx.Err() // context.Canceled once cancel() ran
		case c, ok := <-chunks:
			if !ok {
				return nil
			}
			encodeChunk(c)
		}
	}
}

go deeper

for a junior

Know the shape by heart: a case receiving from ctx.Done() inside the loop's select, returning ctx.Err(). Remember that Done gives you a channel, not a boolean or an error.

for a middle

Explain why closing a channel is used rather than sending: a close is a broadcast every waiter observes. Be able to say that cancellation is cooperative and that a CPU-bound loop must poll ctx.Err() itself.

for a senior

Demonstrate the operational judgment: choosing a polling granularity that bounds shutdown latency, refusing to retry on context.Canceled, and recognising that select gives Done no priority when work is also ready.

for a principal

Own the convention across a codebase: every long-running loop takes a context and returns its error, so shutdown is uniform and observable rather than each service inventing its own stop channel and its own reporting.

## The two methods that matter A `context.Context` exposes cancellation through two methods: - `Done() <-chan struct{}` — a receive-only channel of empty structs. It is never *sent* to; it is **closed**. A closed channel is permanently ready to receive, which makes it a broadcast: every goroutine selecting on it wakes, not just one. - `Err() error` — `nil` while the context is live, and once done it returns a non-nil sentinel. For a context cancelled through the func from `context.WithCancel`, that sentinel is `context.Canceled`. The empty struct carries no data because no data is needed: the *event* is the closure of the channel. ## The canonical loop ``` for { select { case <-ctx.Done(): return ctx.Err() case job, ok := <-jobs: if !ok { return nil } process(job) } } ``` Three things about this shape are worth stating explicitly. **The Done case sits beside the work, not before it.** `select` blocks until one of its cases is ready and picks uniformly at random among the ready ones. That means a worker that is idle on `jobs` wakes the instant the context is cancelled — but it also means that if both a job and the cancellation are ready simultaneously, either branch may win. If shutdown must be strict, re-check `ctx.Err()` before starting expensive work, or drain deliberately. Do not assume `Done` has priority; it does not. **Return `ctx.Err()`, not `nil` and not a home-made error.** The caller can then compare against `context.Canceled` and distinguish an orderly shutdown from a real failure. Returning `nil` after a cancellation tells the caller the work succeeded, which is a lie that propagates: a job runner will mark a half-transcoded output as complete. **Between iterations is the only place cancellation is observed.** Whatever `process(job)` does, it runs to the end. If a single unit of work is long, pass `ctx` down into it so the inner layers can select too, or check `ctx.Err()` between its phases. ## Cancellation is cooperative This is the point candidates most often get wrong. Cancelling a context does not interrupt a goroutine. There is no thread kill, no injected panic, no unwinding. The runtime can *preempt* a goroutine for scheduling purposes, but that only means another goroutine gets CPU time; the preempted one resumes exactly where it was. So a loop like ``` for i := range n { heavyPureComputation(i) } ``` is untouched by cancellation until you make it look. For CPU-bound work the cheap idiom is a polled check, not a select: ``` for i := range n { if err := ctx.Err(); err != nil { return err } heavyPureComputation(i) } ``` `Err()` is a mutex-guarded field read; checking it once per outer iteration is negligible, checking it a million times in the innermost loop is not free. Pick a granularity that bounds your shutdown latency. ## Non-blocking checks When you want to test cancellation without blocking, the `select`-with-`default` form does it: ``` select { case <-ctx.Done(): return ctx.Err() default: } ``` This is equivalent to the `ctx.Err() != nil` check and is mostly a matter of taste; the `Err()` form is shorter and reads better outside a loop that already selects. ## Details that separate a confident answer - **A nil Done channel is legal.** `context.Background()` and `context.TODO()` can never be cancelled, and their `Done()` returns `nil`. Receiving from a nil channel blocks forever, so `case <-ctx.Done():` is simply never selected. The code compiles and behaves correctly; nothing special is required. - **Done can be called repeatedly and by many goroutines.** It returns the same channel each time and is safe for concurrent use, as every method on `Context` is. - **After the channel closes, receives return immediately, forever.** There is no need to store a flag; re-reading is cheap and always true. - **`Err()` is set before `Done()` is closed.** So a goroutine woken by the closure always sees a non-nil `Err()`; you never observe the ordering the other way round. - **Do not select on Done in order to do cleanup in a parked goroutine.** For a one-shot reaction there is `context.AfterFunc(ctx, f)`, added in Go 1.21, which runs `f` in its own goroutine once the context is done and returns a `stop func() bool` you can call to cancel that registration. It avoids a goroutine sitting in a `select` for the sole purpose of noticing. ## Reporting up Once the worker returns `ctx.Err()`, the layer above generally propagates it rather than retrying: `context.Canceled` means somebody asked for this to stop, so retrying is exactly the wrong reaction. Treating cancellation as a transient failure and retrying is one of the classic ways a shutdown turns into a stampede.

  • If both the Done case and a job are ready in the same select, which one runs?
    Either. `select` chooses uniformly at random among ready cases, so cancellation has no priority. If shutdown must be strict, re-check `ctx.Err()` before starting expensive work in the job branch, rather than assuming the Done case wins the race.
  • Is there a way to react to cancellation without parking a goroutine in a select?
    `context.AfterFunc(ctx, f)`, added in Go 1.21, registers `f` to run in its own goroutine once the context is done, and returns a `stop func() bool` that unregisters it. It is the right tool for one-shot cleanup you would otherwise write as a goroutine that only waits on Done.
  • What is ctx.Done() for a context that can never be cancelled?
    `context.Background()` and `context.TODO()` return a nil channel from `Done()`. Receiving from a nil channel blocks forever, so the `case <-ctx.Done():` branch is simply never selected. The loop still compiles and behaves correctly; no guard is needed.
  • Should a caller that receives context.Canceled from a worker retry the work?
    No. `context.Canceled` means somebody deliberately asked the work to stop, so retrying contradicts the request and, during a shutdown, turns one cancellation into a stampede. Propagate it. Retry logic should apply to transient failures, not to cancellation.

saying these in an interview costs you the question

  • Thinks ctx.Done() returns an error rather than a channel
  • Believes cancellation preempts or kills the goroutine
  • Returns nil instead of ctx.Err() after cancellation
  • Checks ctx.Err() once before the loop and never again
  • Assumes the Done case always wins a select against ready work
  • Retries the work after seeing context.Canceled