Why does a Go for-select loop keep processing work after its `ctx.Done()` case is ready?
answer
- both cases can be ready at once
- a closed channel is always receivable
- the choice ignores where you wrote the case
- uniform choice means a coin flip per pass
- make the priority explicit inside the body
basics
~20 sSelect picks uniformly at random among the cases that can proceed; source order gives no preference. A cancelled context's Done channel is ready, but so is the work channel, so each pass is a coin flip.
solid answer
~40 sCancelling a `context.Context` closes its Done channel, which makes `case <-ctx.Done():` permanently ready — but not privileged. If the work channel also has a value queued, both cases are eligible and Go chooses between them uniformly at random, so on average the loop handles one more item before it exits, and more when several work cases are ready. Listing the Done case first changes nothing; `select` has no priority. If cancellation must win immediately, re-check `ctx.Err()` at the top of the work case body and return when it is non-nil, which turns the coin flip into at most one extra item. If instead the queued work should be drained, then the behaviour is correct and you only need to bound how long draining is allowed to take.
code
go · 12 linesfor {
select {
case <-ctx.Done():
return ctx.Err()
case ev := <-redraws:
// both cases can be ready at once, so re-check before doing the work
if err := ctx.Err(); err != nil {
return err
}
paint(ev)
}
}go deeper
Remember that cancelling a context closes its Done channel, and that a closed channel can always be received from. That is why the case becomes ready — not why it is chosen.
Be able to state the rule precisely: among cases that can proceed, the choice is uniform and pseudo-random, so a ready Done case competes with ready work rather than beating it. Then show the ctx.Err() re-check.
Say which contract the loop is meant to honour — stop now, or finish accepted work — and show how you bound the second. Point out that exit latency is dominated by the slowest case body, not by the select.
Own the shutdown budget across services: what a component promises to abandon on cancellation, how long a drain is allowed to take, and where that promise is written down so callers can rely on it.
## The observation A long-lived goroutine looks like this: ```go for { select { case <-ctx.Done(): return ctx.Err() case ev := <-redraws: paint(ev) } } ``` You cancel the context and the loop keeps painting for a moment. In a test it shows up as a count that is one or two higher than expected; in production it shows up as a shutdown that logs work after it announced it was stopping. Nothing is broken — this is exactly what the language specifies. ## Why it happens Cancelling a `context.Context` **closes** the channel returned by `Done()`. A receive from a closed channel succeeds immediately and keeps succeeding forever, so from that instant `case <-ctx.Done():` is *always* eligible. But eligible is not the same as chosen. When more than one case of a select can proceed, the specification says the runtime chooses among them by a **uniform pseudo-random** selection. Source order is irrelevant. So on each pass of the loop: - if the work channel is empty, only the Done case is eligible and the loop exits at once; - if the work channel has a value queued, both cases are eligible and the loop exits with probability one half. The number of extra items handled after cancellation is therefore geometric with mean one, and it gets worse as you add cases: a loop with three busy work cases plus Done picks Done only a quarter of the time, so it averages three extra items. ## Why the choice is random A deterministic priority would be worse. If the first listed case always won, any channel that is usually ready would starve every case below it — the loop would serve one source and never look at the others. Uniform choice guarantees that every ready case makes progress, which is what you want for the multiplexing job select exists to do. The cost is that you cannot express "this case matters more" declaratively. ## Making cancellation win Since select gives you no priority, express the priority in code. The cheap, idiomatic version is a re-check at the top of the work case: ```go case ev := <-redraws: if err := ctx.Err(); err != nil { return err } paint(ev) ``` `ctx.Err()` returns nil while the context is live and a non-nil error once it is cancelled or its deadline has passed, so this costs one comparison per item and bounds the overrun at one *received* item that is then dropped. Note what you have traded: that item was taken out of the channel and thrown away. If losing it matters, either acknowledge it back to its producer or do not use this pattern. The other spelling is a nested select: an inner `select` with the Done case and a `default`, checked before doing the work. It reads worse than `ctx.Err()` and does the same thing. ## The other correct answer: drain Sometimes the whole point is that queued work should finish. Then the loop's behaviour is not a bug and the fix is a different shape: after Done fires, stop selecting and consume the work channel until it is empty or its producer has closed it. That shape needs its own bound — a deadline, or a maximum number of items — because a producer that keeps feeding the channel will keep the "draining" loop alive indefinitely. Decide which of the two you want before writing the loop. "Stop as fast as possible" and "finish what has been accepted" are different contracts, and reviewers read the loop as a statement of which one you chose. ## Diagnosis When you suspect a loop is ignoring cancellation, count what each case body did. A counter per case, or a short recording with `go tool trace`, shows the goroutine's timeline: how long it sat blocked in the select versus running a case body, and how much of that ran after cancellation. That evidence separates "the select kept picking work" from a different failure — a case body that simply takes a long time and cannot be interrupted at all, which no amount of re-checking `ctx.Err()` will fix. ## What not to conclude Cancellation never interrupts a case body that is already running. A context is a signal, not a preemption mechanism: code has to look at it. If your loop's case bodies do a long blocking operation, the loop's exit latency is that operation's duration, whatever the select does.
- Does listing the `ctx.Done()` case first give it any preference?No. Source order affects exactly one thing — the one-time evaluation of the case expressions on entering the select. Among cases that can proceed, the runtime chooses uniformly at random, so a Done case written first is chosen no more often than one written last.
- How does adding more work cases change the shutdown latency?It makes it worse in proportion. The choice is uniform across ready cases, so with one work case Done wins half the time, with three busy work cases only a quarter. A loop that grew from two cases to five silently quadrupled the work it does after cancellation.
- What if the loop should finish the work already queued instead of stopping at once?Then the behaviour you are looking at is correct and you should make it deliberate: once Done fires, leave the select and consume the work channel until it is empty or its producer stops. Bound that phase with a deadline or an item count, or a fast producer keeps the goroutine alive forever.
- Why did Go's designers choose a random tiebreak rather than source-order priority?Because priority starves. If the first case always won, a channel that is nearly always ready would prevent every case below it from ever running, defeating the purpose of multiplexing. Uniform choice guarantees each ready case makes progress; the cost is that priority must be written by hand.
saying these in an interview costs you the question
- Says select prefers whichever case is listed first
- Claims a closed Done channel forces an immediate exit
- Adds a sleep so cancellation has time to land
- Thinks cancelling a context interrupts a running case body
- Reorders the cases and calls the race fixed