skip to content

A Go multiplexer nils each select case as its stream ends, then dies with a fatal all-goroutines-are-asleep deadlock. What went wrong?

level: seniorimportance: nice to knowfreq 30%

answer

  1. every case was switched off
  2. no default, so nothing can proceed
  3. the dump names what each goroutine awaits
  4. victims are parked sending, not selecting
  5. the loop guard must exit, not select

basics

~20 s

The loop disabled its last case. Once every channel in the select is nil and there is no default, the select can never proceed, so the goroutine parks forever. The loop needed its own exit condition.

solid answer

~50 s

Nilling a channel variable retires a `select` case, but it does not end the `select`. When the final stream drains and its variable is set to nil, the statement has no case that can ever become ready and no `default`, so it blocks and the goroutine is parked with nothing able to wake it. If nothing else in the process is runnable — the situation in a focused test — the runtime aborts with `fatal error: all goroutines are asleep - deadlock!` and dumps every goroutine with the reason it is parked. The multiplexer's goroutine shows `[select]` with the loop's own frame on top, which points at the disabling logic. The fix belongs to the loop: run it only while some source variable is non-nil. In a live server the same bug wedges one stream instead of crashing.

code

go · 18 lines
go
func pump(ctrl, data <-chan Frame, out chan<- Frame) {
	for { // BUG: no exit once every case has been disabled
		select {
		case f, ok := <-ctrl:
			if !ok {
				ctrl = nil
				continue
			}
			out <- f
		case f, ok := <-data:
			if !ok {
				data = nil
				continue
			}
			out <- f
		}
	}
}

go deeper

for a junior

Remember the single fact that carries this: a select with no ready case and no default waits, and a nil channel can never become ready.

for a middle

Explain the sequence step by step, from the comma-ok drain to the final nil to the unsatisfiable select, and write the loop condition that terminates it.

for a senior

Work the diagnosis: read wait reasons in the dump, separate the parked victims from the one stuck loop, and explain why a test crashes where the live service merely wedges.

for a principal

Decide the standard your team codes to: whether hand-rolled multiplexers are allowed at all, and what test must exist for any loop whose cases can be disabled at runtime.

## What the code does A stream demultiplexer reads length-prefixed frames off one connection and fans them into per-stream channels, with a `select` loop that watches several of them at once. As each peer finishes, its channel is closed, the loop's comma-ok receive reports `ok == false`, and the loop sets that channel variable to nil to retire the case — the standard idiom, and correct as far as it goes. The bug is in the surrounding loop, written as `for { select { ... } }`. Disabling a case does not remove it and does not end the statement. When the last variable becomes nil, the `select` is left with cases that can never be ready and no `default`, so it blocks. A `select` in that state is not an error; it is simply a wait for an event that cannot occur. The goroutine parks and no code path exists to wake it. ## Why it becomes a fatal error rather than a hang The Go runtime aborts the process when it can see that no goroutine will ever run again — every goroutine is parked on a channel or a synchronisation primitive with nothing left to signal it. The report is `fatal error: all goroutines are asleep - deadlock!`, followed by a dump of every goroutine and its wait reason. That is why a targeted unit test on the multiplexer crashes loudly. In the full service, other goroutines are alive and doing work, so the same defect shows up as one wedged stream and a goroutine that never returns — quieter, and usually noticed later. Recognising that the crash and the wedge are the same bug is the main thing this scenario tests. ## Reading the dump when none of the frames look familiar The dump lists each goroutine with a bracketed wait reason, and the reasons are specific enough to steer you: ```text goroutine 1 [select]: main.pump(...) goroutine 8 [chan send]: goroutine 9 [chan receive (nil chan)]: ``` - `[select]` on a goroutine whose top frame is your loop is the signature of this bug: the statement is live, but nothing it waits on can fire. - `[chan receive (nil chan)]` or `[chan send (nil chan)]` names a bare operation on a nil channel outside a `select` — usually a channel that was never given to `make`. - `[chan send]` on peers is a consequence, not a cause: they are blocked writing into a demultiplexer that has stopped draining. The practical technique when the dump is long and none of the top frames are yours is to sort by wait reason rather than reading top to bottom. Goroutines parked on `chan send` are almost always victims; look for the single goroutine parked in `select` or on a nil channel and start there. It is also worth remembering that this report is a fatal runtime throw, not a panic, so no deferred function runs and no `recover` intercepts it — you get the dump and the process exits. ## The fix Make termination the loop's responsibility, so that the condition mirrors the disabling: ```go for ctrl != nil || data != nil { select { ... } } ``` With a variable number of streams held in a map or slice rather than in named variables, keep an integer count of live sources, decrement it where you would have nil'd, and return when it hits zero. If the loop must outlive its inputs — a multiplexer that expects new streams to be registered later — then it needs a case that is *always* live, such as a registration channel or a cancellation signal, so the select is never left with only disabled cases. ## How to keep it from coming back - Treat "every case disabled" as a state the loop must handle explicitly, the same way you would handle an empty work set anywhere else. - Give the loop a test that drains every source and asserts the function returns. That is the test that turns this bug into a red build instead of a production wedge, because in a small test the runtime's deadlock report fires immediately. - Prefer a bounded shape over a hand-rolled one when the source count is dynamic: a counter plus a return is easier to review than a chain of nil comparisons. - When you review a `select` inside a bare `for {}`, check straight away whether any case can be disabled at runtime. If one can, the loop is missing an exit. ## The summary answer The idiom is right and the loop is wrong. Disabling cases with nil is how you retire finished sources; deciding that there is nothing left to wait for is a job the `select` cannot do for you, and if you leave it to the `select`, the best case is a fatal error in a test and the worst case is a silently stuck goroutine in production.

  • The same code in the running server does not crash. Why not?
    The fatal report requires that no goroutine anywhere can make progress. A live server has accept loops, handlers and background workers still running, so the process carries on and the defect appears as one stream that stops being served plus a goroutine that never returns. Same bug, quieter symptom.
  • In the dump, several goroutines are parked on chan send. Where do you look first?
    Not at them. Goroutines blocked sending into the multiplexer are victims of the one that stopped draining. Look for the goroutine parked in select whose top frame is the loop, or one showing a nil-channel wait reason, and work outward from there.
  • How would you write the exit condition when the number of streams is dynamic?
    Keep an integer count of live sources alongside the map of channels, decrement it at the point where you would set a variable to nil, and return when it reaches zero. If the loop must survive its current inputs, give it a permanently live case such as a registration or cancellation channel so the select is never left fully disabled.
  • Could a deferred cleanup function have logged something useful here?
    No. The all-goroutines-are-asleep report is a fatal runtime throw rather than a panic, so deferred functions do not run and recover cannot intercept it. The goroutine dump the runtime prints is the entirety of what you get, which is why the wait reasons in it are the evidence you work from.

saying these in an interview costs you the question

  • Thinks a select ends by itself once all its cases are nil
  • Blames the goroutines parked on chan send as the cause
  • Proposes adding a default case, turning the wedge into a busy loop
  • Expects a deferred function or recover to handle the fatal report
  • Concludes the nilling idiom itself is wrong and rips it out