skip to content

Why is an error that matches context.Canceled never worth retrying in Go?

level: juniorimportance: should knowfreq 50%

answer

  1. who decided the work should stop?
  2. the condition never clears by itself
  3. a cancelled context fails the next call instantly
  4. errors.Is against the context sentinel
  5. abandoned is not the same as bad

basics

~20 s

context.Canceled means someone deliberately called the work off, so the cause never clears. Any attempt reusing that context fails immediately, and a fresh one produces a result nobody is waiting for. Stop and return the error.

solid answer

~50 s

`context.Canceled` is the sentinel a cancelled `context.Context` reports, and it says nothing about the downstream system being unhealthy — it says the caller walked away, usually because the request was abandoned or the process is shutting down. Retrying with the same context is pure waste: the context is already cancelled, so the very next call fails on entry. Retrying with a fresh context is worse, because you are now doing work whose result has no consumer, and during shutdown you are holding the process open to do it. Because drivers and helpers wrap the cause, check with `errors.Is(err, context.Canceled)` rather than `==`, and treat that branch as "abandoned": stop attempting, release whatever you hold, and return. In a queue worker, abandoned is also not the same as a bad message — the message should go back for redelivery, not to a dead-letter queue.

code

go · 8 lines
go
if err := handle(ctx, msg); err != nil {
	if errors.Is(err, context.Canceled) {
		// The caller walked away. Not a bad message, not a sick
		// dependency: leave it unacknowledged for redelivery.
		return err
	}
	return retryOrDeadLetter(msg, err)
}

go deeper

for a junior

Be ready to say what cancellation means: the caller stopped the work on purpose, so no attempt can change the outcome. Know the check is errors.Is against the context.Canceled sentinel, not a string match.

for a middle

Explain why a cancelled context fails the very next call, and why detaching to a new context turns a wasted retry into an unwanted side effect. Show that wrapping is why you use errors.Is.

for a senior

Demonstrate the third disposition: abandoned work should be redelivered, not dead-lettered and not counted against the retry budget. Tie it to a concrete incident such as a rolling restart poisoning the dead-letter queue.

for a principal

Own the shutdown contract for the service: what in-flight work is allowed to finish, what is abandoned, and how the queue's acknowledgement semantics make abandonment safe rather than lossy.

## What the error actually means Go carries cancellation as a value: a `context.Context` derived from `context.WithCancel`, `context.WithTimeout` or `context.WithDeadline` becomes "done", and the reason is exposed as an error. When the cancellation was explicit — someone called the cancel function, or a parent context was cancelled and the cancellation propagated down — that error is the package-level sentinel `context.Canceled`. That is a statement about the **caller**, not about the **dependency you were calling**. A dial failure, a connection reset or an overloaded downstream are statements about the world: they may be different a second from now. `context.Canceled` is a statement about intent: the work was called off. Intent does not heal on its own, so no number of attempts changes it. ## Why the retry is not merely useless but harmful There are two ways a worker gets this wrong. 1. **Retrying with the same context.** Every context-aware call in the standard library checks `ctx.Done()` before and during the operation. A cancelled context short-circuits, so the retry loop spins as fast as it can, burning CPU and, if you count attempts, exhausting the message's retry budget in milliseconds. To an operator this looks like an error spike with no corresponding downstream traffic. 2. **Retrying with a fresh context.** This is the more dangerous mistake, because it "works". You detach from the cancellation and complete the operation — but the caller has gone, so nobody consumes the result, and during a shutdown you have just extended the process's lifetime doing work the shutdown was meant to stop. Worse, if the operation has side effects (a charge, an email, a write), you have performed one that was explicitly cancelled. ## How to detect it correctly By the time the error reaches your retry decision it has usually been wrapped several times — a driver may return its own error type whose `Unwrap` chain ends at the context error. So the check is: ```go if errors.Is(err, context.Canceled) { return err // abandoned: do not retry } ``` `errors.Is` walks the `Unwrap` chain and compares against the sentinel at every level, which `err == context.Canceled` cannot do. Matching on the message text (`strings.Contains(err.Error(), "canceled")`) is worse still: the text is not part of any package's contract, and different layers spell and decorate it differently. ## Three dispositions, not two Once you have this branch, the natural shape of a worker's error handling is three outcomes rather than a retryable/terminal binary: - **retryable** — the world might be different next time (a connection was refused, the dependency returned an overload signal). Attempt again, subject to a budget. - **terminal** — the input itself is wrong (the payload does not unmarshal, a required field is missing). No attempt will succeed; send it to a dead-letter queue with the error attached. - **abandoned** — cancellation. Neither of the above: you never learned whether the work would have succeeded, so do not consume the message's retry budget and do not dead-letter it. Leave it unacknowledged so it is redelivered to another worker, or to this one after restart. Collapsing abandoned into terminal is a common and expensive bug: a rolling deploy cancels in-flight contexts on every pod, and a fleet's worth of perfectly good messages lands in the dead-letter queue during what should have been a routine restart. ## What an interviewer is listening for That you can say cancellation is caller intent rather than dependency health; that you check it with `errors.Is` against the sentinel because the error is wrapped; and that you can name a concrete consequence of getting it wrong — a spin loop that eats a retry budget, or work completed after the requester disappeared.

  • Why use errors.Is(err, context.Canceled) instead of err == context.Canceled?
    By the time the error reaches your retry decision it has usually been wrapped — a driver or client returns its own error whose `Unwrap` chain ends at the context sentinel. `==` compares only the outermost value and misses every wrapped case, so the cancellation is silently classified as some unknown failure and retried. `errors.Is` walks the chain and compares at each level.
  • During a rolling restart your worker dead-letters every in-flight message. What went wrong?
    Shutdown cancels the workers' contexts, the in-flight handlers return an error matching `context.Canceled`, and the code has only two buckets — retry or dead-letter — so cancellation falls into terminal. The messages were never actually attempted to completion, so the fix is a third disposition: leave them unacknowledged for redelivery and do not count the attempt.
  • Is it ever right to finish the work with a fresh context after cancellation?
    Only when the work must complete for correctness regardless of the requester — flushing a buffer, releasing a lease, writing an audit record. That is cleanup, not a retry, and it needs its own bounded context so shutdown still terminates. Re-running the original operation on a detached context is a bug: it performs an effect that was explicitly called off.

A retryable failure is a busy signal — call back and it may ring. context.Canceled is the other party hanging up on purpose; redialling does not change their mind.

saying these in an interview costs you the question

  • Retrying with the same cancelled context and spinning
  • Treating cancellation as a transient network failure
  • Detaching to a fresh context so the work "succeeds"
  • Matching cancellation on err.Error() text
  • Dead-lettering messages that were merely abandoned at shutdown