skip to content

A Go for-select loop with a default case pins a CPU core while the service is idle. How do you confirm it and fix it?

level: seniorimportance: should knowfreq 44%

answer

  1. idle service, one core at 100 percent
  2. the loop never gets to wait
  3. profile the process while it is doing nothing
  4. a parked goroutine costs nothing
  5. delete the escape hatch, add cancellation

basics

~20 s

The default case makes the select return instantly, so the loop never waits and spins at full speed. Confirm it with a CPU profile of the idle process, then delete the default so the select blocks.

solid answer

~50 s

A `select` with a `default` never parks the goroutine, so wrapping one in a bare `for` produces a spin loop: it re-runs millions of times a second, burning one whole core, and the service looks busy while doing nothing. I would confirm it rather than guess, by taking a CPU profile of the idle process and looking at the top functions — a non-blocking poll loop shows nearly all samples inside the loop function and the runtime's select implementation, with no application work under it. The goroutine profile corroborates it: the goroutine is `running`, not parked in a channel receive. The fix is almost always to delete the `default` and let the select block, adding a `case <-ctx.Done(): return` so the loop can still exit. Blocked goroutines cost no CPU. If the loop genuinely must poll something that is not a channel, pace it with a timer case inside the select — never with a `default`.

code

go · 18 lines
go
// burns a core: the select can never wait
for {
	select {
	case d := <-datagrams:
		handle(d)
	default:
	}
}

// parks until there is work or the daemon is cancelled
for {
	select {
	case d := <-datagrams:
		handle(d)
	case <-ctx.Done():
		return
	}
}

go deeper

for a junior

Know the symptom-to-cause link: a select with a default never waits, so putting one in a bare loop uses a whole core doing nothing.

for a middle

Be ready to explain why the blocking version is cheaper and lower-latency, and why Gosched, a bigger buffer or more GOMAXPROCS do not address the spin.

for a senior

Show the diagnosis path, not just the answer: profile the idle process, read the top frames, corroborate with goroutine states, then remove the default and add a cancellation case.

for a principal

Weigh the fleet cost — one wasted core per replica inflates CPU requests and autoscaling — and decide what stops it recurring: a review rule on empty default branches, plus an idle-CPU check in the release process.

## The shape of the bug Consider a datagram receiver daemon whose main loop was written like this so that it "never gets stuck": ```go for { select { case d := <-datagrams: handle(d) default: } } ``` The intent reads as "take a datagram if there is one, otherwise loop round". The effect is a spin. Because `default` makes the select non-blocking, the loop body completes in tens of nanoseconds and starts again, so the goroutine is *runnable at all times*. It occupies one P for its whole scheduling quantum, gets rescheduled, and repeats — one core saturated for as long as the process lives, at its worst when the service is completely idle. The reason this survives review is that the code looks defensive. "Non-blocking" sounds strictly safer than "blocking", and the daemon does work correctly; it just costs a core to do nothing. ## Confirming it, not guessing The engineer who notices an idle service pinning a core has several cheap confirmations available: 1. **A CPU profile.** Collect one from the idle process and inspect it with `go tool pprof`. A poll loop is unmistakable: essentially all samples sit in one function, with the runtime's select implementation (`runtime.selectgo`) directly beneath it, and nothing that resembles application work. A healthy idle service produces a nearly empty profile, because parked goroutines are not sampled at all. 2. **The goroutine profile.** Fetch a goroutine dump and read the state of each goroutine. A parked receiver reports a channel-receive wait state; the spinning one is simply `running`, and its stack points at the loop. 3. **An execution trace.** `go tool trace` shows the goroutine occupying a processor continuously with no blocking events — useful when you want to see that it is never yielding rather than merely that it is hot. The diagnostic discipline matters because "one core pinned" has other causes — a tight retry loop, a compression or regexp hot path, garbage collection — and the profile distinguishes them in seconds. ## The fix Almost always: **delete the `default`.** A blocking select is not a hazard; a parked goroutine consumes no CPU and is woken by the runtime the moment a channel operation becomes possible. Add a cancellation case so the loop still has a way out: ```go for { select { case d := <-datagrams: handle(d) case <-ctx.Done(): return } } ``` This is strictly better than the original on every axis: idle cost drops to zero, latency improves (the runtime hands the value straight to the parked receiver rather than waiting for the next poll iteration to notice it), and shutdown becomes deterministic. Things that do **not** fix it: - **Calling `runtime.Gosched()` in the `default` branch.** It yields the processor so other goroutines can run, which may make the symptom less visible on a multi-core box, but the loop is still runnable and still consumes every cycle the scheduler gives it. The core stays busy. - **Raising `GOMAXPROCS`.** More parallelism does not make a spin loop stop spinning; it just leaves more cores for everything else while one is still wasted. - **Adding a buffer to the channel.** Buffering changes when sends block, not whether this select waits. - **Sleeping in the `default` branch.** A `time.Sleep` does cut the CPU burn, but it converts the design into a poll with a fixed latency floor and a wake-up per interval. If a channel is the source of the work, blocking on it is better in every respect. ## When polling really is required Occasionally the thing you are waiting on is not a channel — a file's mtime, a device, a shared flag — and you have no way to be woken. Then the loop must be *paced*, and the way to pace a select is a timer channel as a real case, so the select still blocks between passes: ```go for { select { case <-tick.C: poll() case <-ctx.Done(): return } } ``` The goroutine is parked between ticks and the CPU cost falls to nothing. The general rule is worth stating plainly: **`default` belongs only in code that already has other useful work to do on this pass.** In a loop whose only purpose is to wait, a `default` is always a bug. ## Why it is worth catching in review The cost is real but easy to under-price on a laptop: one saturated core per replica, multiplied by every instance you run, showing up as inflated CPU requests, noisy autoscaling, and a machine that never enters a low-power state. It is also perfectly invisible to functional tests — correctness is unaffected — so the only defences are a profile taken at idle and a reviewer who asks what the empty `default` branch is for.

  • What would the CPU profile of a correctly blocking version look like at idle?
    Almost empty. Parked goroutines are not executing, so the sampler collects nearly nothing and total CPU time is close to zero. That contrast is the confirmation: the spinning version attributes essentially all samples to one loop function with the runtime's select code beneath it.
  • Does calling runtime.Gosched() in the default branch fix it?
    No. Gosched yields the processor so other goroutines get a turn, which can mask the symptom on a many-core machine, but the loop remains runnable and keeps consuming every cycle the scheduler gives it. The core stays pinned; only the fairness changes.
  • How would you tell this apart from a hot regexp or compression path also pinning a core?
    The profile names the culprit directly. A poll loop shows the loop function and the runtime's select machinery on top with no application work below; a hot computation shows the real functions. A goroutine dump adds confirmation — the spinner is running with a trivial stack rather than deep in library code.
  • When is a poll loop legitimate, and how do you pace it?
    When the thing you wait on cannot wake you — a file, a device, an external flag. Then keep the select blocking and make a timer channel one of its cases, alongside a cancellation case. The goroutine parks between passes, so idle CPU returns to zero and the interval is explicit rather than accidental.

A watchman who checks the door ten million times a second is not resting between checks; he exhausts himself while nothing happens. A doorbell wakes him only when someone arrives.

saying these in an interview costs you the question

  • Adds runtime.Gosched to the default branch as the fix
  • Blames the Go scheduler rather than the loop
  • Raises GOMAXPROCS to spread the spin
  • Says a blocking select risks getting stuck
  • Guesses the cause without taking a profile