skip to content

How should a Go producer avoid blocking forever when the bounded channel it feeds stays full?

level: middleimportance: must knowfreq 55%

answer

  1. the normal path still blocks
  2. give the send some exits
  3. whose deadline is it?
  4. the event you did not enqueue still exists
  5. count it or return it, never swallow it

basics

~20 s

Wrap the send in a select that also waits on ctx.Done() and, if you need a deadline, on a timer channel. The send then either succeeds, is abandoned because the caller gave up, or is abandoned after a bounded wait — and the give-up branch must do something explicit with the event.

solid answer

~50 s

A bare `events <- e` on a full channel parks the producer until a receiver appears, which can be forever if the consumers are wedged. Put the send in a `select` alongside `case <-ctx.Done(): return ctx.Err()` so cancellation from the caller unblocks it, and add a `case <-timer.C:` from a `time.NewTimer` if the producer needs its own deadline rather than the caller's. That turns an unbounded stall into a bounded wait with three defined outcomes. The important part is the third branch: the event you did not enqueue still exists, so the code must return an error upstream, or record it in a counter that is exported and alerted on. Silently discarding it in the give-up branch trades an obvious hang for an invisible data loss, which is worse to debug. Use `defer timer.Stop()` so a per-send timer is released as soon as the send wins.

code

go · 13 lines
go
func enqueue(ctx context.Context, events chan<- Event, e Event) error {
	timer := time.NewTimer(200 * time.Millisecond)
	defer timer.Stop()
	select {
	case events <- e:
		return nil
	case <-ctx.Done():
		return ctx.Err()
	case <-timer.C:
		shed.Add(1) // var shed atomic.Int64, exported as a metric
		return errQueueFull
	}
}

go deeper

for a junior

Know that a send on a full bounded channel waits, and that putting the send in a select next to a ctx.Done() case is how you let it be given up on. Recall that ctx.Err() reports why.

for a middle

Explain each branch: the send is the normal path, ctx.Done() inherits the caller's deadline, a time.NewTimer case adds your own, and the give-up branch must return an error or count the abandoned event. Mention defer timer.Stop().

for a senior

Show judgment about the timeout value and the failure mode it creates: a short deadline sheds under load, a long one keeps propagating the stall upstream. Be ready to say how you would see abandoned events on a dashboard.

for a principal

Frame the deadline as configuration owned by the service, not a constant in the call site, and be clear that shedding on timeout is a data-loss decision the data's owner has to have agreed to before the incident.

## The problem with a plain blocking send A bounded intake channel is what gives a Go service backpressure: when the workers are behind, `events <- e` parks the producing goroutine. That is usually exactly what you want — the stall propagates upstream. But an unconditional park has no upper bound. If the consumers are deadlocked, wedged on a dependency, or simply gone, the producer waits forever. Its caller — an HTTP handler, an RPC method, a supervising loop — waits with it, holding whatever resources it holds, and the failure presents as a hang rather than as an error. So the pattern is: keep the blocking send as the normal path, and give it exits. ## The cancellation exit ```go select { case events <- e: return nil case <-ctx.Done(): return ctx.Err() } ``` `ctx.Done()` returns a channel that is closed when the `context.Context` is cancelled or its deadline passes. A closed channel is always ready to receive, so the moment the caller gives up, this `select` completes through the second case and the producer returns. `ctx.Err()` then tells the caller which happened — `context.Canceled` or `context.DeadlineExceeded`. This exit is the one to reach for first, because it costs nothing and it inherits the deadline someone upstream already set. If the request that produced this event has a two-second budget, the enqueue attempt should not outlive it. ## The deadline exit Sometimes the producer needs its own limit — a background reader with no request context, or a policy that says "we will wait 200 ms for room, then this event is not going in". ```go timer := time.NewTimer(200 * time.Millisecond) defer timer.Stop() select { case events <- e: return nil case <-ctx.Done(): return ctx.Err() case <-timer.C: return errQueueFull } ``` `time.NewTimer(d)` returns a `*time.Timer` whose `C` field delivers one value after `d`. `defer timer.Stop()` releases it as soon as the send wins rather than letting it run to completion. `time.After(d)` is the one-liner equivalent but allocates a fresh timer on every call and gives you no handle to stop, which matters in a hot loop. ## The branch people forget The give-up branch is where the design actually lives. Reaching it means an event exists that is not in the queue and never will be. Three defensible things to do with it, and one indefensible one: - **Return an error to the caller**, letting whoever produced the event decide — retry, buffer at their end, fail the request. This is the right default when the producer is a request handler, because the client learns the truth. - **Record it and move on**, incrementing an exported counter so the number of abandoned events is visible on a dashboard and can be alerted on. This is defensible for data whose owner has agreed it may be lost. - **Both** — count it and return a distinguishable error like `errQueueFull` so the metric and the caller agree. The indefensible one is a bare `return nil` in that branch. It converts a visible hang into invisible data loss, and nothing in the system will ever tell you it happened. ## What the bounded wait actually buys It does not increase throughput. Under sustained overload the queue is full most of the time, so a timeout on the send simply chooses *how long the producer is willing to participate in the stall* before giving up. What you gain is: - a bounded worst-case latency for the producer's caller; - an explicit, countable shed event rather than a silent pile-up; - liveness under consumer failure — a wedged worker pool degrades your service instead of freezing it. A short timeout sheds aggressively and protects latency; a long one preserves data and lets the stall keep reaching upstream. The number is a policy choice, and it belongs in configuration rather than baked into the call site. ## Details worth getting right If the channel has room *and* the context is already cancelled, both cases are ready and the choice between them is not specified — do not write code whose correctness depends on which one wins. Recreating a timer for every single event on a very hot path is measurable; reusing one `*time.Timer` with `Reset` is possible but easy to get wrong, and the usual answer is to make the timeout coarser or to attach it to a batch rather than to each item. Finally, a timeout on the send is not a substitute for bounding the channel. It is the escape hatch on top of a bound, not an alternative to one.

  • The channel has room and ctx is already cancelled. Which select case runs?
    It is not specified — when more than one case is ready the choice among them is arbitrary, so the event may be enqueued or the cancellation may win. Never write logic that depends on the outcome. If cancellation must always take priority, check `ctx.Err()` before entering the select, and accept that even then a cancellation arriving mid-select can lose the race.
  • Why prefer time.NewTimer with defer Stop over time.After in this send?
    `time.After` allocates a new timer on every call and gives you no handle, so on a hot enqueue path you create one per event. `time.NewTimer` plus `defer timer.Stop()` releases the timer as soon as the send succeeds. Since Go 1.23 an unreferenced timer can be collected before firing, so this is an efficiency habit rather than a leak fix.
  • What is wrong with returning nil from the timeout branch?
    It tells the caller the event was accepted when it was dropped. You have traded a visible hang for silent data loss that no metric or log records, and the first person to notice is whoever audits the downstream data days later. Return a distinguishable error, increment an exported counter, or both.

saying these in an interview costs you the question

  • Leaves the give-up branch empty so events vanish silently
  • Returns nil after failing to enqueue, hiding the loss
  • Assumes the Go runtime times out a blocked send by itself
  • Believes the runtime always panics on a stalled send
  • Uses a send timeout instead of bounding the channel at all
  • Depends on which case wins when several are ready