skip to content

Your RPC spans show context.DeadlineExceeded at one hop and context.Canceled at the next — what does each value tell you about where the deadline was enforced?

level: seniorimportance: should knowfreq 50%

answer

  1. the two values point in opposite directions
  2. one hop timed out, one was abandoned
  3. no error at all is the worst case
  4. the value never crosses the wire
  5. a deadline is only as real as its observer

basics

~20 s

context.DeadlineExceeded marks the hop whose own context chain ran out of time; context.Canceled marks a hop torn down because someone upstream had already given up. The deadline was enforced at the first hop, and the second was collateral.

solid answer

~50 s

Read the two sentinels as a direction. A hop reporting `context.DeadlineExceeded` had a deadline in *its own* context chain expire — either one it derived or one it inherited, because a child context adopts its parent's error when the parent finishes. A hop reporting `context.Canceled` was stopped by a decision: its caller went away and the transport cancelled the request context it was serving. So the deadline was enforced at the `DeadlineExceeded` hop, and the `Canceled` hop is downstream fallout, not the culprit. Two follow-ups matter operationally: a hop that logged *no* error at all while its caller timed out never observed its context and is still burning CPU on an abandoned request; and when several deadlines are stacked in one process, `context.DeadlineExceeded` alone cannot say which one fired — attaching a named cause and recording `context.Cause` is what makes that legible.

code

go · 12 lines
go
func terminalClass(ctx context.Context, err error) string {
	switch {
	case errors.Is(err, context.DeadlineExceeded):
		return "deadline_exceeded" // this hop's own budget ran out
	case errors.Is(err, context.Canceled):
		return "canceled" // someone upstream gave up first
	case err != nil:
		return "failed"
	default:
		return "ok"
	}
}

go deeper

for a junior

Know that both values mean the work stopped early, and that context.DeadlineExceeded means time ran out while context.Canceled means someone gave up on purpose.

for a middle

Explain that a child context inherits its parent's error inside a process, that the error value never crosses a network boundary, and that each service derives its own from its own transport.

for a senior

Walk the diagnosis end to end: which hop's budget bound, why the cancelled hop is fallout rather than cause, and how you spot the hop that recorded no error because it never watched its context.

for a principal

Own the convention that makes this readable across teams — what terminal class every service records, whether budgets are named with a cause, and how the error-rate SLI separates abandonment from expiry.

## Reading the two sentinels as a direction of travel When a request crosses several services and each one records `ctx.Err()` on its span, the pattern of sentinels is a map of who gave up first. **`context.DeadlineExceeded` at a hop** means a deadline in that process's own context chain expired. That is either a deadline the hop derived for itself, or one it inherited: within a single process, when a parent context finishes, its children are cancelled with the parent's error, so `DeadlineExceeded` propagates down the local tree. It is the value of the hop that ran out of time. **`context.Canceled` at a hop** means a decision arrived. Nothing timed out locally; something stopped wanting the answer. In a server, that is normally the transport noticing the caller's connection has gone and cancelling the context of the request being served. The work was abandoned from above. So a trace showing `DeadlineExceeded` at hop A and `Canceled` at hop B, downstream, reads: A's budget ran out, A dropped the call, B's transport saw the caller vanish and cancelled B's request context. **A is where the deadline was enforced; B is collateral.** Investigating B's slow query as "the cause of a cancellation" is the classic wrong turn. ## The error value does not cross the wire The crucial mechanic behind that reading: a context's error value is process-local. Cancelling a context on one machine does not transmit a Go error to another. Each process computes its own `ctx.Err()` from its own context tree, and whatever cross-service signal exists — a connection closing, a stream being reset, a deadline the transport encoded into the request — is translated back into a *local* cancellation by that service's transport layer. That is why the sentinels differ across hops even though it is conceptually one timeout. Do not expect the same value everywhere and do not conclude from a `Canceled` downstream that no deadline was involved. ## The three-case classification worth writing down Record `ctx.Err()` on the span at every hop and the terminal class of each becomes one of: - **DeadlineExceeded** — this hop's own budget was binding. If it is the edge, your overall SLO budget is too tight or a dependency got slow. If it is a middle hop, look at the deadline that hop derived: a per-attempt budget shorter than the work. - **Canceled** — this hop was abandoned. Normal traffic in a healthy system; a *rise* in it means something upstream is timing out or clients are disconnecting. - **No error recorded, yet the caller timed out** — the ugly one. The callee never looked at its context: it did not thread `ctx` into the call it was blocked on, or it was blocked on something that ignores contexts entirely. It ran to completion, wrote its result into a void, and held a connection and memory the whole time. This is a timeout enforced at the wrong layer — the caller gave up, the callee never learned. That third case is the one worth hunting. A caller-side deadline is not a cancellation mechanism; it is only a promise about how long the caller will wait. Real cancellation requires the callee to observe `ctx.Done()` and to have passed `ctx` down to whatever it is waiting on. ## Telling stacked deadlines apart Within a single process there are often several deadlines in one chain — the overall request budget, a per-attempt budget, a per-dependency budget. All of them produce the identical `context.DeadlineExceeded`, and the sentinel alone cannot say which fired. Two things fix this. First, `context.Cause(ctx)`, when the deadline was created with a cause attached, returns that named error while `ctx.Err()` still reports `context.DeadlineExceeded` — so you can tag the span with "per-attempt budget" and keep every existing `errors.Is` check intact. Second, `ctx.Deadline()` returns the deadline and an `ok` flag, and comparing it against what this layer intended tells you whether a shorter deadline from above is the binding one. A hop whose effective deadline is far earlier than the one it set is being governed by its caller, which is usually correct and always worth knowing. ## For the engineer arriving from another language If you expect a timeout to arrive as a thrown exception that unwinds the callee's stack, the Go picture will look broken. Nothing is thrown, nothing is interrupted, and a goroutine that is not watching `ctx.Done()` is not affected in any way by a cancellation. The sentinels are reports, not events: `ctx.Err()` tells you what the context decided, and your code returning it is the only reason a caller ever sees it. The operational consequence is the whole point of the leaf. A deadline is only as real as the code that observes it, and the way you audit that is exactly this: put the terminal error class of each hop on its span, and look for the hop that finished successfully while its caller had already gone.

  • A downstream hop records no error at all, yet its caller reported context.DeadlineExceeded. What does that mean?
    The callee never observed its context. It either did not thread `ctx` into the call it was blocked on, or it was blocked on something that ignores contexts. It finished the work, wrote the result nobody was waiting for, and held a connection and memory for the whole extra duration. The timeout was enforced only at the caller.
  • Several deadlines are stacked in one process and all report context.DeadlineExceeded. How do you tell which fired?
    The sentinel cannot distinguish them, so attach a named cause when the deadline is created and record `context.Cause(ctx)` beside `ctx.Err()`. `Err()` stays `context.DeadlineExceeded` so nothing that classifies on it breaks, while the cause names the layer. Comparing `ctx.Deadline()` with what this layer intended also shows when a caller's shorter budget is the binding one.
  • Why do two hops in one request report different sentinels for what is conceptually one timeout?
    Because a context's error is process-local and never travels over the network. Each service computes its own `ctx.Err()` from its own context tree; the cross-service signal is a connection closing or a stream being reset, which that service's transport turns into a local cancellation. The hop whose clock ran out sees `DeadlineExceeded`; the one that was dropped sees `Canceled`.
  • Should a rise in context.Canceled at your service alarm you?
    A baseline of it is normal — callers disconnect. A *rise* is a symptom rather than a cause: it usually means an upstream hop started timing out, or clients are abandoning slow requests. Investigate the hop reporting `context.DeadlineExceeded`, not the one reporting cancellations, and keep the two as separate signals rather than one error rate.

saying these in an interview costs you the question

  • Treating a cancelled hop as the cause of the timeout
  • Expecting the same sentinel at every hop of one request
  • Assuming a context error travels across the network
  • Believing a caller's deadline stops the callee's work
  • Folding cancellations and deadline expiries into one error rate