skip to content

Why does a drain budget built with context.WithTimeout(sigCtx, ...) fire instantly when sigCtx is an already-cancelled signal.NotifyContext context?

level: middleimportance: must knowfreq 55%

answer

  1. cancellation only travels one direction
  2. the shutdown path starts after the root died
  3. a child of a dead parent is dead
  4. Err says Canceled, not DeadlineExceeded
  5. reparent on Background or WithoutCancel

basics

~20 s

Cancellation flows from parent to child. A context derived from an already-cancelled parent is born cancelled: its Done channel is closed at once and its Err is context.Canceled, not context.DeadlineExceeded. Parent the budget on a live context.

solid answer

~40 s

`signal.NotifyContext` returns a context that is cancelled when the signal arrives, and the shutdown path only starts *after* that cancellation. Deriving the drain budget from it with `context.WithTimeout(sigCtx, 20*time.Second)` gives a child of a dead parent: cancellation propagates downward, so the child's `Done` channel is already closed before the timer is ever consulted, and `Err()` reports `context.Canceled` rather than `context.DeadlineExceeded`. The drain then "completes" in microseconds and everything in flight is abandoned, which looks exactly like a fast clean shutdown in the logs. The fix is to parent the budget on something still live: `context.Background()`, or `context.WithoutCancel(sigCtx)` when you want to keep the values carried on the signal context. The same mistake shows up wherever a shutdown routine passes the cancelled root context down to the last outbound calls it still needs to make.

code

go · 13 lines
go
<-sigCtx.Done() // the signal has fired, sigCtx is cancelled

// wrong: born cancelled, Err() is context.Canceled, no drain happens
bad, cancelBad := context.WithTimeout(sigCtx, 20*time.Second)
defer cancelBad()

// right: a fresh root
good, cancelGood := context.WithTimeout(context.Background(), 20*time.Second)
defer cancelGood()

// right: keeps the values on sigCtx, drops its cancellation
keepValues, cancelKeep := context.WithTimeout(context.WithoutCancel(sigCtx), 20*time.Second)
defer cancelKeep()

go deeper

for a junior

Remember the direction: cancelling a parent context cancels all of its children, and never the other way round. A context created from a parent that is already cancelled is dead from birth.

for a middle

Be able to state exactly what the broken child reports — Done already closed, Err equal to context.Canceled rather than context.DeadlineExceeded, Deadline still in the future — and name both fixes, context.Background and context.WithoutCancel.

for a senior

Show that you would find this in production: the log says the drain was clean and fast while the queue keeps redelivering. Talk about branching on Err to tell a genuine deadline from an inherited cancel, and about the other post-cancellation calls that break the same way.

for a principal

Make the rule explicit for the codebase — the shutdown path gets its own bounded context, and passing the cancelled root context into shutdown work is a review-blocking defect — and require a test that drives the real cancellation order rather than calling the drain with a fresh context.

## The bug in one line ```go sigCtx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM) defer stop() <-sigCtx.Done() // a termination signal arrived: sigCtx is now cancelled drainCtx, cancel := context.WithTimeout(sigCtx, 20*time.Second) // wrong parent defer cancel() <-drainCtx.Done() // returns immediately; there is no 20-second drain ``` The program looks like it drains for twenty seconds. It drains for microseconds. ## Why the child is born cancelled A `context.Context` derived from a parent forms a tree, and cancellation only ever travels **downward**. When `context.WithTimeout` (or `WithCancel`, or `WithDeadline`) attaches a new child, it first checks whether the parent is already done. If it is, the child is cancelled on the spot with the parent's error and the timer never matters. So after the signal has fired: - `drainCtx.Done()` is a closed channel from the moment it exists — every `select` on it takes that branch immediately; - `drainCtx.Err()` is `context.Canceled`, **not** `context.DeadlineExceeded`; - `drainCtx.Deadline()` still reports a deadline twenty seconds in the future, which is why reading the deadline is a poor way to check whether the budget is alive. This is not special to signal contexts. It is the ordinary rule that the earliest cancellation in a chain wins, applied at the one moment when the root of the chain is guaranteed to be dead: the shutdown path. ## Why it is so easy to write Everywhere else in the program, `ctx` is the right thing to pass along — that is the whole convention. The shutdown path is the one place where the ambient context is deliberately a *finished* one, so the reflex to "pass ctx down" produces work that cannot run. The symptom is also friendly: nothing panics, nothing errors, the process exits quickly and the log says shutdown was clean. The only visible evidence is on the queue side, where items that were reported as drained keep being redelivered. ## The fix Parent the drain budget on a context that is still live: ```go // simplest: a fresh root drainCtx, cancel := context.WithTimeout(context.Background(), 20*time.Second) // or keep the values carried on the signal context, drop its cancellation drainCtx, cancel := context.WithTimeout(context.WithoutCancel(sigCtx), 20*time.Second) ``` `context.WithoutCancel` returns a context that carries the parent's values but is never cancelled when the parent is. That matters when the signal context was built on top of a context holding request-scoped or process-scoped values — a logger, a trace identity, a tenant — that the drain's final outbound calls still need. If the signal context carries no values, `context.Background()` is clearer. ## The same trap further down the shutdown path Once the root is cancelled, *any* work started during shutdown needs a fresh, bounded context, not the cancelled one: - acknowledging the last processed items back to the queue; - a final flush of buffered metrics or logs to a collector; - deregistering from service discovery; - committing an offset or a checkpoint. All of those are outbound calls made after cancellation. Handed the cancelled context, they fail instantly with `context.Canceled`, usually into an ignored error, and the shutdown quietly loses the very bookkeeping that would have made it recoverable. ## Distinguishing the two outcomes Because the two failures produce different errors, the drain can tell them apart and should: ```go switch { case errors.Is(drainCtx.Err(), context.DeadlineExceeded): // the budget genuinely ran out with work in flight case errors.Is(drainCtx.Err(), context.Canceled): // the budget was cancelled by something else - usually the wrong parent } ``` If a drain reports `context.Canceled` at all, treat it as a bug in the shutdown wiring until proven otherwise: the budget was supposed to end by deadline or not at all. ## How to catch it It does not show up in a unit test that calls the drain function directly with a fresh context, because the fresh context is not cancelled. It shows up when the test drives the real path: cancel the signal-shaped context first, then assert that the drain still waited — that a slow in-flight item was allowed to finish, or that the elapsed time was on the order of the budget rather than zero.

  • What does drainCtx.Deadline() report in that broken case?
    It still reports a time twenty seconds in the future, because the deadline was recorded when the child was created. Only `Done` and `Err` reflect the inherited cancellation, which is why checking the deadline is a useless liveness test and `Err()` is the thing to inspect.
  • When would you use context.WithoutCancel instead of context.Background as the parent?
    When the signal context sits on top of values the drain still needs — a logger, a trace identity, configuration stashed in the context. `context.WithoutCancel` keeps those values while detaching the cancellation. If there are no such values, `context.Background()` says the same thing more plainly.
  • What else in the shutdown path breaks for the same reason?
    Every outbound call made after cancellation: acknowledging processed items, flushing buffered telemetry, deregistering from discovery, committing a checkpoint. Handed the cancelled root context they fail immediately with `context.Canceled`, usually into an ignored error, so the shutdown silently loses its own bookkeeping.
  • How do you write a test that would have caught this?
    Drive the real path rather than the drain function in isolation: cancel the signal-shaped root first, then run the shutdown with one deliberately slow in-flight item and assert the item was allowed to finish, or that the elapsed wait was on the order of the budget instead of about zero.

saying these in an interview costs you the question

  • Thinks the child's own timer overrides an already-cancelled parent
  • Expects context.DeadlineExceeded when the parent was cancelled
  • Passes the cancelled root context to the final queue acknowledgements
  • Checks Deadline() instead of Err() to see whether the budget is alive
  • Believes cancellation can propagate upward from child to parent