skip to content

In a Go for-select loop, what do you check before approving a pull request that adds a new case?

level: seniorimportance: should knowfreq 40%

answer

  1. one goroutine serves every case
  2. while a body runs, nothing else is served
  3. count the cases, not just read them
  4. more ready cases, rarer cancellation
  5. who closes the channel it selects on?

basics

~20 s

Check that the new case's body never blocks, that it either returns or falls through to the next pass, and that adding it makes the cancellation case rarer, since the choice is uniform among ready cases.

solid answer

~50 s

The loop is a single-threaded dispatcher, so my first question is what the new body does: any blocking call in it stops every other source and lengthens shutdown, and slow work belongs in its own goroutine with the value handed over. Second, the exit path — each case must either `return` or fall out and go round again, and a new `return` has to run the same cleanup as the existing ones. Third, arithmetic: the choice among ready cases is uniform, so a fourth busy case means the `ctx.Done()` case is now chosen a quarter of the time rather than a third, and shutdown latency rises accordingly. Fourth, the channel itself — who sends on it, whether it is ever closed (a closed channel's case is ready on every pass), and, if it is a send case, whether a receiver is guaranteed to still be there.

go deeper

for a junior

Focus on the mechanical checks you can make yourself: does the new case body return or fall through, and does it avoid long-running work that would hold up the loop?

for a middle

Explain why the loop is safe without a mutex — one goroutine, one case at a time — and therefore why a blocking body is a whole-component problem rather than a local slowdown.

for a senior

Show the arithmetic on the uniform choice and how a new case changes shutdown latency, and name the evidence you would ask for: per-case counters, and a short trace recording of the goroutine's timeline.

for a principal

Decide where the line is between one loop with many cases and several loops that communicate, and write that down, because a component whose stop budget is set by how many cases someone added is not a component anyone can operate.

## Why adding a case is not a local change A for-select loop is the single-threaded heart of a component. Every case shares one goroutine, one stack and one set of local state, and that is the whole reason the pattern is safe without a mutex. Adding a case adds a new claim on that one goroutine, so its effects are global to the loop even though the diff is four lines. Here is what I read for, roughly in the order I find problems. ## 1. Does the body block? While a case body runs, the goroutine is not in the select. Nothing else is being served: keystrokes queue, resize events queue, and the cancellation case is not even looked at. A body that does a synchronous write, a network call, or a lock acquisition with contention converts the loop's latency for every source into that body's duration. The fix is not to make the loop concurrent — that reintroduces the races the loop was avoiding. It is to move the slow work out: hand the value to a dedicated worker goroutine over a channel and return to the select immediately. If the handover itself can block, that is the next thing to look at, because a blocking send inside a case body is the same bug wearing a different hat. ## 2. What is the exit path? Every case body must end in one of two ways: fall out of the select and go round the loop, or `return`. A new case that returns must do everything the existing return paths do — the same deferred cleanup, the same signal to whoever is waiting for this goroutine to finish. A new case that can neither return nor complete quickly extends how long the component takes to stop, which is often what a reviewer notices last and an on-call engineer notices first. ## 3. The arithmetic of the choice Among cases that can proceed, Go chooses uniformly at random. That means the probability of any particular case being chosen is one over the number of *currently ready* cases. A loop with two work cases and a cancellation case picks cancellation a third of the time when everything is busy; add a fourth and it is a quarter. The expected number of extra work items handled after cancellation goes up correspondingly. So a pull request that adds a case has, as a side effect, changed the shutdown behaviour of the component without touching the shutdown code. If the component has a stated stop budget, this is where it quietly gets spent. ## 4. Where does the channel come from, and does it close? Ask who sends on it and whether anyone closes it. A receive case on a **closed** channel succeeds immediately on every pass, forever — so a case whose channel gets closed while the loop still selects on it turns the loop into a spin that consumes a core and drowns out the useful cases. If the closure is meant to signal something, the case body must act on it and leave the loop rather than continue. If the new case is a **send** rather than a receive, the question inverts: is a receiver guaranteed to be there? A send case that can never proceed is merely never chosen, which is quiet — the value is computed and discarded on every pass and nobody notices until you look at why the loop is allocating. ## 5. What is in the case header? Case expressions are evaluated on every entry to the select, before any case is chosen. A call in the header runs every pass whether or not its case wins, and a call that changes state there drops work invisibly. A case header should name a channel and, for a send, a value that already exists. ## 6. How would we know if this is wrong in production? I want the loop to be observable before it grows. A counter per case, exported, answers most questions after the fact: a case that never increments is a channel nobody feeds; one that dominates is starving the rest by holding the goroutine. For a live investigation, a short recording with `go tool trace` shows the goroutine's timeline — how much time it spent parked in the select versus running in a case body, and which stretches of wall-clock the loop was unavailable. That evidence distinguishes "the select kept choosing the other case" from "one case body ran for 400 ms", which need completely different fixes. ## The summary I give the author If the case is a fast, non-blocking handler for an event that genuinely belongs to this component's state, add it. If it blocks, if it needs its own error handling, or if it is really a second responsibility, it wants its own goroutine and its own loop — and then the two loops communicate over a channel, which is a change the reviewer can reason about locally.

  • How would you confirm in production which case the loop is actually taking?
    Export a counter incremented at the top of each case body — that alone shows a case nobody feeds or one that dominates. For a live look, take a short recording with `go tool trace`: its goroutine timeline shows how long the loop sat parked in the select versus running inside a body, which is what separates a choice problem from a slow-handler problem.
  • The new case does a 20 ms synchronous write. What do you ask the author to change?
    Move the write off the loop. Hand the value to a writer goroutine over a buffered channel and go straight back to the select, so the loop's latency for every other source stays flat. If the handover can block when the writer falls behind, decide explicitly what to do then — drop, coalesce, or apply backpressure — rather than letting the loop stall.
  • When would you say no, and ask for a second goroutine and loop instead?
    When the new case is a separate responsibility rather than another event for the same state: it has its own lifetime, its own error handling, or its own rate. Two loops that talk over a channel are easier to reason about and to stop than one loop with five cases and three exit paths.

saying these in an interview costs you the question

  • Assumes ordering the cases sets a priority
  • Approves a case body doing blocking I/O in the loop
  • Thinks the cases run concurrently with each other
  • Ignores that a new case changes shutdown latency
  • Adds a case with no exit path and no cleanup
  • Puts a state-changing call in the case header