skip to content

Why can a Go library's Close deadlock when its background goroutine is blocked on a channel send?

level: seniorimportance: should knowfreq 45%

answer

  1. the goroutine is parked, not looping
  2. the top-of-loop check guards nothing below
  3. both sides are waiting for each other
  4. make the send a two-case select
  5. capacity only delays the same hang

basics

~20 s

Because Close waits for a goroutine that is parked in a channel send nobody will receive. The goroutine never reaches its cancellation check, so both sides wait forever. Make the send itself cancellable with a select.

solid answer

~50 s

A goroutine only notices cancellation where it looks for it, and a bare `w.events <- e` looks nowhere. Once the consumer stops reading — often because it just called `Close` — the send parks, the goroutine never gets back to the top of its loop to see `ctx.Done()`, and a `Close` that joins parks behind it. The fix is to make every send a `select` with the stop signal as the other case, so shutdown wins and the event is abandoned. Then decide and document the rest of the contract: the sending goroutine closes the events channel on its way out, and the caller is told either to keep draining until the channel is closed or that events may be dropped during shutdown. A capacity on the channel only postpones the hang; it never removes it.

code

go · 9 lines
go
func (w *Watcher) Close() error {
	w.cancel()
	<-w.stopped // waits for the goroutine
	return nil
}

func (w *Watcher) emit(e Event) {
	w.events <- e // parks forever once nobody receives
}

go deeper

for a junior

Remember that a channel send blocks until someone receives, so a goroutine sitting in a send is not running its loop and cannot notice that it was told to stop.

for a middle

Explain why a cancellation case at the top of a loop does not protect a send at the bottom, and write the two-case select that makes the send abandonable. Know that capacity only postpones the block.

for a senior

Walk the whole shutdown ordering: which side closes the events channel, what Close guarantees, and what the caller must do. Show that you audit every blocking operation in an owned goroutine against the stop signal.

for a principal

Own the choice between dropping in-flight events and requiring callers to drain, and make sure that choice is documented on the exported API rather than discovered by whichever team hits the hang first.

## The shape of the hang A library exposes a `Watcher`: `New` starts a goroutine, the goroutine produces events onto a channel the caller ranges over, and `Close` stops the watcher. The goroutine looks reasonable: ```go for { select { case <-ctx.Done(): return case raw := <-w.source: w.events <- decode(raw) // no cancellation here } } ``` The cancellation check exists, but it guards only the *top* of the loop. The send at the bottom is unguarded, and an unbuffered (or full) channel send blocks until a receiver is ready. A caller that stops ranging over `events` and calls `Close` produces the deadlock precisely: the goroutine is parked in the send, `Close` has cancelled the context and is parked in `<-w.stopped`, and the only party that could unblock the send — the caller — is inside `Close`. It is worth naming both failure modes, because they are the two ends of the same design question: - **Close blocks forever** when it joins a goroutine that cannot reach its cancellation check. - **Close returns too early** when it does not join at all, and the goroutine goes on sending, logging or writing after the caller believed it had shut down — the version that produces "writes after close" and test flakes. ## Making the send cancellable Every blocking operation inside an owned goroutine has to be selectable against its stop signal. For a channel send that is a two-case select: ```go select { case w.events <- e: case <-ctx.Done(): return } ``` Now shutdown always wins: a cancelled context makes `ctx.Done()` permanently ready, so the send can be abandoned and the goroutine returns, which lets `Close` return in turn. Note what this decision costs — the event `e` is dropped. That is a deliberate policy choice, and it is the right default for shutdown, but it must be a choice you make rather than one the deadlock makes for you. What you must not write instead is a check-then-send: ```go if ctx.Err() != nil { return } w.events <- e // cancellation between the two lines still hangs ``` The window between the check and the send is exactly where cancellation lands. Nor should you reach for `default:` to make the send non-blocking unless dropping events under *normal* load is acceptable too — a `default` arm drops whenever the consumer is momentarily slow, not just during shutdown. ## Who closes the events channel The goroutine that sends is the one that closes, on its way out — `defer close(w.events)` in the producing goroutine. This gives the consumer a clean `for e := range w.events` that terminates by itself, and it means the consumer never has to guess whether more events are coming. It also keeps `Close` honest: if `Close` joins the goroutine, then by the time `Close` returns the events channel is already closed, and the caller's range loop has already ended. ## Buffering is not a fix Giving the events channel capacity — 128, 4096, whatever — changes only how long it takes to hang. The producer runs until the buffer is full, then blocks in exactly the same way. Capacity is a smoothing device for bursty consumers; it is not a shutdown mechanism, and choosing a large one mostly converts a fast, obvious deadlock into a slow, confusing one, plus a chunk of retained memory. ## The contract you have to write down Once the mechanics are right, the remaining work is the documented contract, because the caller's behaviour is half of it: - *Either* the caller must keep receiving from the events channel until it is closed, and `Close` may block until they do — a strong ordering that suits a pipeline stage. - *Or* the library abandons in-flight events on shutdown so `Close` never depends on the caller, at the cost of dropping whatever was in flight — the safer default for a library imported by teams you cannot review. Whichever you pick, say it in the doc comment for `Close` and for the events accessor. A caller who does not know which of the two applies will write the one you did not implement. ## Reviewing for it The review heuristic is short: inside a goroutine your package owns, find every operation that can block — channel sends and receives, mutex acquisition, network reads, waits — and ask what makes each of them return when the stop signal fires. Any one of them that has no answer is the operation your `Close` will eventually hang on, at the worst possible time.

  • Why is `if ctx.Err() != nil { return }` followed by a plain send not a fix?
    Because cancellation can happen in the gap between the check and the send. The check passes, the consumer stops reading a microsecond later, and the send parks exactly as before. Only a select makes the check and the send a single atomic choice among ready cases.
  • Would giving the events channel a large capacity solve it?
    No — it delays it. The producer fills the buffer and then blocks identically, so the deadlock reappears later and looks less obviously related to shutdown. Capacity smooths bursts between a producer and a consumer that are both still running; it is not a cancellation mechanism, and a big one just holds memory.
  • Should the caller or the library close the events channel?
    The library, from the goroutine that sends on it, in a defer as that goroutine returns. The sender is the only party that knows no more values are coming, and a caller closing it would risk a panic on the next send. With a joining Close, the channel is therefore already closed by the time Close returns.
  • Your library must not drop events during shutdown. What changes?
    The contract moves onto the caller: they must keep draining until the events channel is closed, and Close is documented as blocking until they do. Alternatively Close takes a deadline and returns an error when in-flight events could not be delivered in time, which keeps the hang bounded and visible instead of silent.

The stop order is posted on the door, but the worker is holding a parcel and waiting for someone to take it. Nobody comes, and the person who posted the notice is outside waiting for the worker to leave.

saying these in an interview costs you the question

  • Adds capacity to the channel and calls the deadlock fixed
  • Checks ctx.Err() before a blocking send and calls it cancellable
  • Uses a default arm, silently dropping events under normal load too
  • Has the caller close the channel the library sends on
  • Makes Close non-joining to dodge the hang, creating writes after close