skip to content

A goroutine parked in sync.Cond.Wait ignores context cancellation — how do you make shutdown work?

level: seniorimportance: nice to knowfreq 26%

answer

  1. there is no WaitContext
  2. you cannot select on a condition variable
  3. cancellation has to become state
  4. flip a flag under the lock, then wake everyone
  5. context.AfterFunc can drive the Broadcast

basics

~20 s

sync.Cond has no context-aware or deadline-aware Wait, so only Signal or Broadcast can release a parked goroutine. Shutdown must set a closed flag under c.L and then Broadcast, with every waiter re-checking that flag after Wait returns.

solid answer

~50 s

`sync.Cond` offers only `Wait()` — there is no `WaitContext`, no deadline, and no way to `select` on it, because the goroutine parks in the runtime's notify list rather than on a channel. So a cancelled `context.Context` does not touch a waiter at all: the only thing that can release it is a `Signal` or `Broadcast` on the same `*sync.Cond`. The fix is to make cancellation part of the predicate. Add a `closed` (or `err`) field guarded by `c.L`, have every waiter loop on `for !ready && !closed`, and give the type a `Close` that sets the flag under the lock and then `Broadcast`s. To bridge an existing `context.Context`, `context.AfterFunc` (Go 1.21) runs a function when the context is done — have it take the lock, set the flag and `Broadcast`, and call the returned `stop` when you finish normally. If callers genuinely need per-call deadlines, that is the signal to replace the condition variable with a channel the caller can `select` on.

code

go · 13 lines
go
func (q *queue) take() (string, error) {
	q.mu.Lock()
	defer q.mu.Unlock()
	for len(q.items) == 0 && !q.closed {
		q.notEmpty.Wait()
	}
	if len(q.items) == 0 {
		return "", errors.New("queue closed")
	}
	v := q.items[0]
	q.items = q.items[1:]
	return v, nil
}

go deeper

for a junior

Take away one fact: Wait takes no context and no deadline, so the only thing that can wake a parked goroutine is a Signal or Broadcast on that same condition variable.

for a middle

Be able to write the fix: a closed flag guarded by the same mutex, a wait loop that includes it, and a Close method that sets the flag under the lock and then Broadcasts.

for a senior

Demonstrate the operational side — recognising stuck waiters in a goroutine profile, bridging an existing context with context.AfterFunc and calling its stop on the happy path, and testing that shutdown really releases all N waiters.

for a principal

Own the call on whether an old condition-variable-based package gets patched or migrated to channels, weighing the callers who now expect per-request deadlines against the cost and risk of rewriting a component other teams depend on.

## The problem, stated precisely `sync.Cond`'s entire waiting API is `func (c *sync.Cond) Wait()`. There is no variant taking a `context.Context`, no variant taking a deadline, and no channel you could put in a `select`. A goroutine inside `Wait` is parked on the condition's internal notify list, and the *only* events that move it are `Signal` and `Broadcast` on that same `*sync.Cond`. That has a blunt consequence for anyone maintaining a package written before `context` existed: cancelling the context you passed in does nothing to a goroutine already parked. `ctx.Done()` closes, `ctx.Err()` becomes non-nil, and the waiter sleeps on regardless. In a graceful shutdown that tries to drain workers and join them, the join hangs; in a goroutine dump you see the stuck goroutines sitting in `sync.(*Cond).Wait` with an ever-growing wait duration. ## Remedy 1 — make cancellation part of the predicate This is the idiomatic fix, and it is the one to lead with. A condition variable is only ever a notification about lock-protected state, so express cancellation *as* lock-protected state: 1. Add a field — `closed bool`, or better an `err error` recording why — guarded by the same mutex as the rest of the state. 2. Every waiter's loop condition includes it: `for len(q.items) == 0 && !q.closed`. 3. After the loop, distinguish the two exits: real work available, or shut down (return an error). 4. Shutdown takes the lock, sets the flag, unlocks, and calls `Broadcast` — not `Signal`, because the flag concerns every parked goroutine. This turns "cancel" into an ordinary state transition that the existing machinery already handles correctly, and it needs no new primitive. ## Remedy 2 — bridge a context into that flag When the caller already hands you a `context.Context` and expects it to be honoured, connect the two explicitly. Since Go 1.21 `context.AfterFunc(ctx, f)` runs `f` in its own goroutine as soon as `ctx` is done, and returns a `stop` function you call to detach it when you no longer need it: ``` stop := context.AfterFunc(ctx, func() { q.mu.Lock() q.closed = true q.mu.Unlock() q.notEmpty.Broadcast() }) defer stop() ``` The bridge goroutine is the thing that watches the context; the condition variable still only ever responds to `Broadcast`. Call `stop()` on the normal path so the watcher does not outlive the work — and note that this cancels *everyone* waiting on that `Cond`, which is right for a shared resource being shut down but wrong if you wanted a per-caller deadline. ## Remedy 3 — per-caller deadlines mean the wrong primitive If what you actually need is "this one caller gives up after 200ms while the others keep waiting", `sync.Cond` cannot express it, and no amount of flag-setting will make it. That requirement is a channel requirement: hand each caller something it can `select` on alongside `ctx.Done()`. Rewriting a small internal queue that way is usually a day's work and removes a whole class of shutdown bugs. The honest interview answer names this as the real endpoint rather than layering tricks on the condition variable. Anti-patterns worth naming so you can reject them: a background goroutine that `Broadcast`s on a ticker so waiters "periodically get a chance to notice" — that turns a precise wake-up into polling and still gives no deadline guarantee; and relying on process exit, which does release the goroutines but only by killing everything, so any in-flight work and any deferred cleanup in those goroutines is lost. ## Proving it, and diagnosing it The behaviour is testable without any timing luck. Start N waiters, each incrementing a parked counter under the mutex before calling `Wait`; spin until the counter reaches N so you know all N are actually parked; then cancel the context (or call `Close`) and require all N to return. Run the same harness with `Signal` instead of `Broadcast` and watch exactly one return — the difference is then a fact your test suite enforces rather than a comment. In production the diagnostic is the goroutine profile. Waiters blocked here appear with `sync.(*Cond).Wait` in their stack, and because a parked goroutine keeps alive everything its frame references, a leaked set of them is visible both as a rising goroutine count and as retained memory. A goroutine count that never returns to baseline after a shutdown request is the tell that some `Cond` never got its `Broadcast`. ## The reviewer's checklist When this shape reaches a code review, three questions catch nearly everything: is there any state change at all that can make the wait loop exit besides the happy path; is the shutdown notification a `Broadcast` rather than a `Signal`; and does every waiter distinguish "woke up with work" from "woke up because we are shutting down" instead of falling through as if it had work.

  • Why can't you simply select on a sync.Cond alongside ctx.Done()?
    `select` operates only on channel operations, and `sync.Cond` exposes no channel — `Wait` parks the goroutine on the condition's internal notify list inside the runtime. There is nothing to put in a `select` case. If callers need to combine waiting with cancellation in one statement, the primitive has to be a channel.
  • What does context.AfterFunc give you that a hand-rolled watcher goroutine does not?
    It runs your function when the context is done without you writing and owning a `select` goroutine, and it returns a `stop` function that detaches the callback so nothing outlives the work. Hand-rolling it is a goroutine you must remember to terminate on the happy path — a common source of the very leak you were trying to avoid.
  • How do you spot leaked sync.Cond waiters in a running service?
    Take a goroutine profile: parked waiters show `sync.(*Cond).Wait` in their stacks, and a count that never falls back to baseline after a shutdown request means some condition variable never received its Broadcast. Because each parked goroutine retains everything its frame references, the leak also shows up as memory that will not be reclaimed.
  • Is a background goroutine that Broadcasts on a ticker an acceptable way to add timeouts?
    No. It converts a precise wake-up into polling, wakes every waiter repeatedly to find nothing changed, and still gives no per-caller deadline — the waiter simply gets more chances to notice a flag that may never flip. If per-caller deadlines are the requirement, replace the condition variable with a channel the caller can select on.

A parked waiter is like someone asleep in a room with no phone and no window: the fire alarm outside is irrelevant, and the only thing that gets them out is someone opening the door and shouting.

saying these in an interview costs you the question

  • Expects a cancelled context to wake a goroutine parked in Wait
  • Looks for a WaitContext or Wait with a deadline on sync.Cond
  • Tries to select on a sync.Cond alongside ctx.Done()
  • Uses Signal to release waiters during shutdown
  • Polls with a ticker that Broadcasts periodically
  • Relies on process exit to clean up parked waiters