When a client's request times out and it stops waiting for a response, does the server-side work triggered by that request actually stop too? What has to be true for a timeout to translate into real cancellation of in-flight work, rather than just the caller giving up while the callee keeps computing?
answer
- client timeout is local, doesn't stop server by itself
- transport cancel signal: gRPC RST_STREAM / context cancel
- cooperative deadline check at checkpoints
- zombie work wastes CPU/DB connections even after client gives up
- Go context.Context Done()+Deadline() combines both
basics
~20 sNot automatically. The client giving up just means it stops waiting; the server keeps working unless something explicitly tells it to stop, like a cancellation signal sent over the still-open connection, or the server checking a shared deadline or context object during its own work and bailing out early.
solid answer
~50 sA timeout on the client side is purely local; it just means the client stops waiting for a response and frees its own resources. For that to actually stop the server's work, either (a) the transport itself supports cancellation propagation, like gRPC sending a cancel signal or HTTP/2 sending a RST_STREAM when the client closes the connection, and the server framework checks for that signal at points during its work; or (b) the server independently derives its own deadline from a propagated header and periodically checks 'have I passed my deadline' at cooperative checkpoints (before each downstream call, in a loop, etc.) and aborts early on its own. Without either mechanism, the server has no idea the client left and keeps consuming CPU, memory, connections, and downstream capacity on work whose result will be discarded, a correctness-neutral but resource-wasteful outcome that gets dangerous under load.
go deeper
Should intuit that just because the client stopped waiting doesn't necessarily mean the server stopped working, even without naming a mechanism.
Should be able to name at least one concrete mechanism (a cancel signal or a deadline check) that turns a timeout into real cancellation.
Should explain both transport-level cancellation and cooperative deadline checking, their respective limitations, and the zombie-work/resource-exhaustion consequence of neither being present.
Should reason about combining both mechanisms as a platform-level default (e.g., via a shared context or middleware library), identify where cancellation silently fails to propagate (across thread pools, queues, non-gRPC legacy hops), and connect this to load-shedding capacity planning during incidents.
## A client timeout is a purely local event A timeout firing on the client side is, by itself, an entirely local event: it means the calling thread stops blocking on the socket read, the client library throws a timeout exception or returns an error, and the caller's own resources (the thread, the connection pool slot) are freed to move on. Nothing about that local event automatically reaches into the server process and tells it 'stop what you're doing.' If nothing else happens, the server keeps executing the request exactly as if the client were still patiently waiting — CPU cycles, memory, open database connections, and any downstream calls the server itself makes all continue to be consumed — and when the server eventually finishes, it writes a response into a socket that the client has already closed or stopped reading from, and that response is simply discarded. This is 'wasted work' or 'zombie work,' and it's a subtle but real production hazard: under load, a burst of timed-out requests doesn't reduce load on the server the way you'd hope, the server keeps doing the full amount of work for every one of them, so timeouts alone don't shed load, they just hide the problem from the client's perspective while the server silently keeps burning capacity that could have gone to requests still worth serving. ## What turns a timeout into real cancellation For a timeout to become real cancellation, actually stopping the in-flight work, freeing the resources it holds, and not consuming further downstream capacity, one of two mechanisms has to be in place. ## Transport-level cancellation propagation The first is transport-level cancellation propagation: protocols like gRPC (built on HTTP/2) let a client that gives up send an explicit signal, closing the stream, which the transport turns into a `RST_STREAM` frame or an equivalent cancel notification, that the server-side framework receives asynchronously, independent of whatever the request-handling code is doing. If the server's handler is written to observe that signal, it can abort early: - stop a long-running database query; - release a lock; - stop calling further downstream services. In gRPC, that means checking `context.Err()` in Go, or `Context.isCancelled()` in Java, or having the framework interrupt a blocked call automatically. Plain HTTP/1.1 has a weaker version of this, the server can sometimes detect the client closed the TCP connection, but there's no standardized in-band cancel signal the way HTTP/2 and gRPC have, so cancellation propagation is notably better-supported in gRPC-based systems than classic REST-over-HTTP/1.1 ones. ## Cooperative deadline checking The second mechanism, which works even without transport-level cancellation, is cooperative deadline checking: the server derives its own absolute deadline from the propagated timeout or deadline header on the inbound request (this is the same deadline-propagation mechanism used to compute downstream call budgets), stores it somewhere accessible during request handling (a context object, a thread-local, a request-scoped variable), and the request-handling code itself periodically checks 'has this deadline already passed?' at natural checkpoints: - before starting each downstream call; - at each iteration of a processing loop; - before doing an expensive computation step. It bails out early with a deadline-exceeded error if so. This doesn't require the client to still be connected or to have sent any explicit cancel signal; it only requires the server to know its own deadline and check it periodically. The cost is that it's cooperative, a single long, uninterruptible operation (a slow synchronous database call with no query timeout of its own, a CPU-bound loop with no deadline check inside it) won't be interrupted mid-flight, so the deadline check is only as effective as how finely-grained the checkpoints are. ## Robust systems combine both In practice, robust systems combine both: transport-level cancellation (so a client's abandonment interrupts work fast, without waiting for the next checkpoint) plus cooperative deadline checks at each downstream call boundary (so even non-gRPC hops, or work that transport-level cancellation can't reach directly, still bail out on their own). Go's `context.Context`, which carries both a `Done()` channel signaled by cancellation and a `Deadline()` value, is the canonical implementation of this combined pattern, and it's why idiomatic Go service code is expected to select on `ctx.Done()` inside loops and pass `ctx` into every downstream call. ## The symptom when the discipline is missing The concrete failure mode when this discipline is missing shows up as exactly the kind of incident postmortem where a client-facing dashboard shows request timeouts recovering (clients gave up and moved on) while server-side CPU and database connection-pool utilization stay pinned at the same elevated level, because the server never got the message that the work it was doing was no longer wanted.
- Why doesn't a client timing out and closing its connection automatically free the server-side resources the request was using?Closing the connection is a client-local action; unless the transport explicitly notifies the server (e.g., gRPC's cancel signal over HTTP/2) and the server code checks for that notification, the server has no way to know the client left and keeps running the handler to completion, holding whatever CPU, memory, and connections it was using the whole time.
- What's the practical limitation of cooperative deadline checking as a cancellation mechanism?It only stops work at the checkpoints the code explicitly checks; a long, uninterruptible operation like a single slow database call with no internal timeout, or a tight CPU-bound loop with no deadline check inside it, will run to completion regardless of the deadline having passed. Its effectiveness depends entirely on how finely-grained and how comprehensive those checkpoints are.
- Why might a service's server-side CPU and database connection usage stay elevated even after client-facing timeout error rates recover during an incident?If cancellation propagation isn't wired up, clients giving up doesn't stop the corresponding server-side work from continuing to run to completion; the client sees fewer errors because it stopped waiting, but the server is still doing the full amount of underlying work for every one of those abandoned requests, so its resource usage doesn't drop in step with the client-visible error rate.
Hanging up the phone doesn't stop the person on the other end from still looking something up for you, unless the phone system tells them the line went dead, or they periodically check whether you're still there.
saying these in an interview costs you the question
- Assumes a client timeout automatically stops server-side work
- Doesn't distinguish transport-level cancellation (RST_STREAM/gRPC cancel) from cooperative deadline checking
- Has no answer for why server resource usage can stay high even as client timeout rates drop
- Thinks HTTP/1.1 has the same in-band cancel signal as gRPC/HTTP2
- Never mentions checking a deadline or context at loop or call boundaries