skip to content

Your handler sets a 2s context.WithTimeout, but its outbound HTTP and database calls still run for 30s. Why?

level: seniorimportance: must knowfreq 52%

answer

  1. the deadline fired; nobody was listening
  2. nothing interrupts a goroutine from outside
  3. look at the argument lists, not the timer
  4. the shorter function name is ctx-free
  5. Query versus QueryContext

basics

~20 s

Almost certainly the context was derived and then never passed to the calls. A Go deadline enforces nothing by itself; it only bites when the ctx reaches the operation, via http.NewRequestWithContext, QueryContext and the other ctx-taking variants.

solid answer

~50 s

A context deadline is advisory: it closes a channel and sets an error, and nothing else. Nothing in the runtime interrupts a goroutine or aborts a socket read because a deadline passed - the work stops only if the `ctx` reached the code doing it. So check plumbing first. Did the handler build its request with `http.NewRequest` (or call `http.Get`) rather than `http.NewRequestWithContext`? Did it call `db.Query` rather than `db.QueryContext`? Is there a helper in the middle taking no `context.Context` that quietly starts a fresh one from `context.Background()`? Any of those severs the chain and the derived context becomes a variable nobody reads. Confirm it with a per-hop latency histogram: the handler's own span walls off at 2 s while the downstream span runs to 30 s. The fix is threading `ctx` as the first parameter down to the call.

code

go · 7 lines
go
ctx, cancel := context.WithTimeout(r.Context(), 2*time.Second)
defer cancel()

// ctx is never passed, so neither of these observes the 2s budget.
resp, err := http.Get(u)
rows, dbErr := db.Query(q, id)
_, _, _, _ = resp, err, rows, dbErr

go deeper

for a junior

Remember that a context only helps if it is actually handed to the call. Learn the paired names: Query and QueryContext, NewRequest and NewRequestWithContext.

for a middle

Explain that Go cancellation is cooperative - the deadline closes a channel and sets an error, and the work stops only where some code observes it. Name the standard library pairs where the shorter name silently drops your budget.

for a senior

Diagnose it from evidence rather than reading code at random: a per-hop latency histogram that walls off at the deadline for the handler but not for the downstream call, plus a connection-pool gauge showing work outliving the request. Then explain the capacity cost that makes it urgent.

for a principal

Set the standard that makes this unrepresentable: ctx as the first parameter across every request-path package, review rules for context.Background outside main and tests, and a decision on whether unbounded legacy helpers get fixed or fenced behind a bounded wrapper.

## The premise to correct first Deriving a context with a deadline does exactly two observable things: it arranges for a channel to close at a moment in time, and it arranges for `Err()` to return `context.DeadlineExceeded` after that. It does not stop a goroutine. It does not close a socket. It does not roll back a query. **Go has no mechanism for interrupting a goroutine from the outside.** Cancellation is entirely cooperative, and a deadline is cooperative cancellation with a timer attached. So when a 2-second budget visibly fails to bound a 30-second call, the question is not "why did the deadline not fire?" - it almost certainly fired on schedule - but "who was supposed to be watching, and were they given the context at all?" ## The usual causes, roughly in order of frequency **1. The ctx-free variant was called.** The standard library offers pairs, and the shorter name is the one without a context: - `http.Get`, `http.Post` and a request built by `http.NewRequest` carry a background context; only `http.NewRequestWithContext` attaches yours. - `db.Query`, `db.Exec`, `db.QueryRow` and `db.Begin` ignore your context; `db.QueryContext`, `db.ExecContext`, `db.QueryRowContext` and `db.BeginTx` honour it. The convenience function is shorter to type, autocompletes first, and appears in every quickstart, so it wins by default. This is the single most common cause. **2. A helper in the middle drops the chain.** A function whose signature is `func fetchUser(id string) (*User, error)` cannot receive your deadline. Somewhere inside it there is a `context.Background()` or a `context.TODO()` - often added years ago to get something compiling - and everything below that line is running on an unbounded budget. The derived context in your handler is then a local variable that nothing reads. Sometimes the compiler will not even complain, because the ctx is still passed to *something* (a log line, a metric) while the expensive call underneath is detached. **3. The ctx is passed but the work is asynchronous.** The handler hands `ctx` to a goroutine it never waits on, or pushes an item onto a queue whose consumer has its own lifetime. The deadline correctly stops the handler and returns an error to the client, while the background work continues to completion - which looks, from a resource graph, exactly like the timeout failing. **4. The deadline is set after the expensive part.** The clock on `context.WithTimeout` starts at the call site. If the handler decodes a large body, acquires a lock, or waits on a semaphore before deriving the context, none of that time is inside the budget. ## How to confirm it rather than guess A **per-hop latency histogram** settles this quickly. Instrument the handler's total duration and each outbound call separately. The signature of an unplumbed context is unmistakable: the handler's own measured span piles up at exactly the deadline value - a hard wall at 2 s, because that path *is* honouring the budget - while the downstream call's histogram runs on to 30 s with no wall at all. If instead both spans stop at 2 s and the client still sees 30 s, the deadline is being observed and the problem is somewhere else entirely, such as a queue in front of the handler. Two cheap corroborations: - Grep the request path for the ctx-free names (`db.Query(`, `http.Get(`, `http.NewRequest(`) and for `context.Background()` / `context.TODO()` outside `main` and tests. In practice this finds it in minutes. - Check the connection pool's in-use gauge. If the handler returned at 2 s but the pool is still holding the connection at 30 s, work is outliving the request - which is exactly the capacity cost that makes this bug matter. Be aware that the standard toolchain will not find it for you. `go vet` includes `lostcancel`, which catches a `CancelFunc` that is never called, but nothing in the standard checks flags a function that had a context available and used the non-context call anyway. ## Why this bug is expensive rather than merely annoying When a client gives up at 2 s and the server keeps working for 30, every timed-out request leaves behind a live connection, an in-flight query, and a goroutine. Under load the arithmetic is brutal: the arrival rate is unchanged, but the effective service time is fifteen times what your capacity model assumed, so the pool saturates and even fast requests start queuing behind abandoned ones. The user-visible symptom - everything slow - has no obvious link to the one handler with a missing `ctx` argument, which is why the plumbing check belongs early in the incident checklist rather than late. ## The fix, and how to keep it fixed Thread `context.Context` as the **first parameter** of every function on the request path, from the handler to the call that actually talks to the network, and use the `...Context` variant at the bottom. A helper that cannot receive a context is a helper that cannot be bounded, and changing its signature is the fix - not passing a fresh background context into it. As a review habit: any `context.Background()` outside `main`, a test, or a deliberately detached cleanup path deserves a comment explaining itself.

  • Suppose ctx is plumbed correctly everywhere and the handler still returns at 30 seconds. What do you look at next?
    Whether the deadline is derived before or after the expensive part - body decoding, lock acquisition and waiting for a pool slot all sit outside the budget if they happen first. Then whether anything is queued in front of the handler, since a request can wait a long time before your `WithTimeout` line ever executes. Comparing handler-internal time with end-to-end time separates the two.
  • Why is this bug worse under load than the raw 30 seconds suggests?
    Because abandoned work keeps its resources. Each timed-out request still holds a connection, a query and a goroutine for the full 30 s, so effective service time is many times what capacity planning assumed. The pool saturates, healthy requests queue behind dead ones, and the whole service degrades from one missing argument.
  • Will go vet catch a call that ignores an available context?
    No. `go vet` ships `lostcancel`, which flags a `CancelFunc` that is never called, but the standard checks do not flag using `db.Query` where `db.QueryContext` was available. Finding it is a grep for the ctx-free names plus for `context.Background()` and `context.TODO()` outside `main` and tests, backed by a code-review habit.

saying these in an interview costs you the question

  • Assumes the runtime kills the goroutine when the deadline passes
  • Says the deadline did not fire, without checking plumbing
  • Thinks deriving a context is enough without passing it
  • Blames clock skew or the timer rather than the argument list
  • Fixes a helper by giving it its own context.Background()