A multiplexer goroutine has been parked on a channel send for hours. What could ever wake it?
answer
- wakes are caused, never discovered
- no deadline is stored when a goroutine parks
- the fatal deadlock check needs every goroutine asleep
- its stack is still a GC root
- capacity one, or a cancellation case
basics
~20 sOnly another goroutine receiving from that channel, or closing it. A park carries no timeout and Go's scheduler never revisits waiting goroutines, so once the caller has gone, the send stays parked for the life of the process.
solid answer
~50 sNothing in the runtime will ever wake it on its own. A park is undone only by a counterpart operation on the same channel — a receive, or a close that wakes the sender into a panic. There is no timeout inside the park, Go's goroutine scheduler never re-examines waiting goroutines, and the whole-program deadlock check only fires when *every* goroutine is asleep, which is not the case in a service still handling traffic. In a multiplexer that delivers a frame into a per-request reply channel, the usual cause is that the caller gave up — it timed out, returned, and left an unbuffered channel with no receiver. The delivering goroutine parks forever, holding its stack, its wait record on the channel's send queue, and everything its frames reference. Capture a trace and look at per-goroutine blocking to confirm the wait reason is `chan send` at the delivery site; the fix is to make delivery unable to park.
code
go · 14 linestype mux struct {
mu sync.Mutex
waiting map[uint64]chan frame // request id -> reply channel
}
func (m *mux) deliver(f frame) {
m.mu.Lock()
ch := m.waiting[f.id]
m.mu.Unlock()
if ch == nil {
return
}
ch <- f // unbuffered: parks forever if the caller has already returned
}go deeper
The takeaway to hold on to: a send on a channel with no receiver waits forever, and no timer or scheduler rescues it. Always know which goroutine is going to receive what you send.
Explain the mechanism — a park is undone only by a counterpart operation on the same channel, so a missing receiver is permanent — and that a partial hang never triggers the runtime's all-asleep fatal error.
Show the whole diagnosis: correlate a climbing goroutine count with grouped parked stacks and an execution trace's per-goroutine blocking view, then fix the delivery so it cannot park, rather than shortening the wait.
Own the invariant across the codebase. Make "every send names its receiver, or carries capacity or a cancellation case" a review rule, and require goroutine-count trending as a released signal, since this class of failure produces no errors at all.
## The shape of the hang A connection multiplexer keeps a map from in-flight request id to the channel the waiting caller will read its reply from. One reader goroutine pulls frames off the connection and delivers each one to the channel registered for its id: ```go m.mu.Lock() ch := m.waiting[f.id] m.mu.Unlock() ch <- f // parks if nobody is receiving ``` If every caller sticks around to receive, this is fine. The moment a caller can leave early — a context deadline, a cancelled request, an error return higher up — the delivery can find a channel with no receiver. On an unbuffered channel the send parks, and the goroutine's state becomes *waiting* with reason `chan send`. ## Why nothing wakes it This is the part worth being precise about, because the instinct is to assume something eventually notices: - **The park itself has no timeout.** When the runtime parks a goroutine it records a reason and a wait record and stops there. There is no deadline field that a timer could fire on. - **The scheduler never looks at waiting goroutines.** Runnable goroutines live on run queues; a parked one is on a channel's wait queue and is simply not a scheduling candidate. Nothing scans channels for waiters that could be satisfied. - **The monitor thread does not rescue it.** The runtime's background monitor preempts goroutines that run too long and retakes contexts stuck in syscalls; a goroutine that is *voluntarily* waiting is exactly what the design expects, so it is left alone. - **The global deadlock check does not fire.** The runtime's "all goroutines are asleep" fatal error requires that no goroutine anywhere is runnable. A server still accepting connections has plenty, so a partial hang is silent by construction. - **Garbage collection does not free it.** A parked goroutine's stack is a GC root and its wait record keeps it reachable. Even if the last reference to the map entry and the channel is dropped, the goroutine is not collected — the goroutine is what keeps the channel alive, not the other way round. The only thing that can end the wait is another goroutine performing the counterpart operation on that exact channel: a receive, which completes the transfer and readies the sender, or a `close`, which readies it into a panic. Neither can happen if the caller has already returned and dropped its end. ## Cost of one parked goroutine It is not just a slot in a counter. The parked goroutine holds its stack — grown to whatever depth the delivery path reached — and, because that stack is a GC root, everything reachable from its frames: the frame payload it is trying to deliver, the buffer it was sliced from, and any request state still referenced. On a busy multiplexer, one permanently parked delivery per abandoned request turns into a steadily climbing goroutine count and a heap that grows in proportion, with no error anywhere in the logs. The heap looks like a leak; the goroutine profile explains it. ## Confirming it rather than guessing The explanation you want to be able to give a colleague has three parts: which goroutines are parked, on what, and who was supposed to wake them. - **Count first.** A goroutine total that climbs and never falls, tracked over time, tells you the parks are permanent rather than merely slow. - **Group the stacks.** A dump or goroutine profile groups parked goroutines by stack, and a large group all sitting at the same `ch <- f` line in the delivery function is the whole diagnosis in one screen; the wait reason on each of them will read `chan send`. - **Use an execution trace for the timeline.** Capture a trace over a window with `go tool trace`, then look at the per-goroutine view: it shows how long each goroutine spent blocked and where it blocked, so you can show that these goroutines went to sleep at delivery and never ran again, while the rest of the service kept working. That distinguishes "parked forever" from "slow consumer" — which is the real question, because the two look identical in a single snapshot. ## Making the delivery unable to park The durable fixes all remove the possibility of an unwakeable park rather than shortening it: - **Give the reply channel capacity 1** and send to it exactly once. The delivering goroutine can then always complete the send, whether or not the caller is still there; the value is simply dropped when the channel is collected. - **Select on the caller's cancellation.** Send in a `select` with a `case <-ctx.Done():` (or a per-request done channel) so an abandoned request lets the deliverer walk away instead of parking. This requires the deliverer to hold something that is cancelled when the caller leaves. - **Deregister on departure.** Have the caller remove its id from the map when it gives up — necessary anyway to stop the map growing, though on its own it only narrows the race rather than closing it, since the deliverer may already hold the channel. - **Never let one abandoned caller stall the reader.** Whatever the mechanism, the goroutine reading frames off the connection must not be able to block on a single consumer, or one departed caller stops delivery for every other in-flight request on that connection. The general rule this scenario teaches: whenever you write a channel send, you should be able to name the goroutine that will receive it and say why that goroutine cannot leave first. If you cannot, the send needs capacity or a cancellation case, because the runtime is not going to supply either.
- Why doesn't the runtime report a deadlock for this program?Its fatal "all goroutines are asleep" check fires only when no goroutine anywhere is runnable. A service still accepting connections and serving other requests always has runnable goroutines, so a partial hang — even a permanent one — is invisible to that check by design. Detecting it is your job, not the runtime's.
- If nothing references the reply channel any more, is the parked goroutine collected?No. The goroutine is a GC root in its own right, and its wait record keeps the channel reachable — the direction of liveness is the opposite of what people expect. It keeps its stack and everything its frames reference alive until it is woken or the process exits.
- What is the smallest change that makes that delivery unable to park?Give each reply channel capacity 1 and send exactly one value to it. The delivering goroutine's send then always completes, whether or not the caller is still receiving, and the abandoned value is collected with the channel. It costs one buffered slot per in-flight request and removes the park entirely.
- How would you tell a permanently parked sender from a merely slow consumer?A single snapshot cannot: both show goroutines blocked on the same send. Look at change over time — a goroutine count that only climbs, the same stacks present across repeated dumps — or capture an execution trace and check whether those goroutines ever ran again in the window. A slow consumer drains; a parked one never moves.
saying these in an interview costs you the question
- Expects the runtime to time out a long-blocked channel send
- Says the background monitor thread will retake a parked goroutine
- Thinks garbage collecting the channel frees the blocked goroutine
- Assumes a hang would trigger the all-goroutines-asleep fatal error
- Treats it as a slow consumer and adds buffer capacity as the fix
- Says the parked goroutine costs only a few kilobytes, ignoring what its stack keeps alive