When a web framework times out a request, why can the handler keep holding a worker, and what decides whether it stops?
answer
- response slot and execution slot are different
- nothing forcibly stops running code
- the blocking call owns the clock
- signal helps only if propagated
- bound the call, not just the response
basics
~20 sNothing forcibly stops running code, so an abandoned handler keeps its worker until its current call returns. It stops early only if the framework signals cancellation and the handler reaches a point where it can observe it.
solid answer
~40 sThe timeout frees the response, not the execution slot. In a thread-per-request model the worker running the handler cannot be safely killed from outside, so it stays occupied until the blocking call it is parked in returns on its own — meaning a handler blocked for 30 seconds holds its worker for 30 seconds even though the client was answered at 2. In an awaitable model the framework can cancel the request's task, but the handler only notices at an await point, and a blocking call made without offloading is just as unstoppable there. The practical decider is whether the slow operation itself has a bound: a per-call timeout on the outbound request or the database call is what actually returns the worker.
go deeper
Hold on to one fact: answering the client does not end the work. A handler waiting on a slow call keeps its worker until that call returns by itself.
Be able to explain why forcible termination is unsafe, why an awaitable is only cancellable at its suspension points, and why a per-call bound is what actually returns the worker.
Demonstrate the diagnosis: in-flight count climbing while responses still flow, worker stacks parked in one dependency, and the fix applied at the call rather than at the pool size.
The tradeoff to argue is where the ceiling belongs — bounded concurrency per dependency and deliberate load shedding versus large pools that turn a fast failure into a long brownout.
## Two separate resources, one timer Every in-flight request holds two different things: the **response slot** (the client is waiting for bytes) and the **execution slot** (a worker, a task and whatever connections and buffers the call chain has acquired). A request timeout releases the first. The second is released only when the handler's own call stack unwinds. That asymmetry is the reason a service can answer every request inside its timeout and still saturate. Capacity is bounded by *concurrent execution slots*, not by *outstanding responses*. Once a timed-out handler no longer appears in any latency graph but still occupies a slot, the two numbers diverge and the graphs stop describing the machine. ## Why the worker cannot simply be taken back - **Forcible termination is unsafe.** Stopping a worker at an arbitrary instruction can leave a half-written record, an unreleased lock or a connection returned to a pool in an unusable state. Runtimes therefore do not offer a safe "stop that now" primitive, and frameworks do not fake one. - **Interruption is a request, not an order.** Where a runtime offers an interruption-style flag, it only sets a marker; code parked in a network read frequently ignores it entirely, so setting it does not guarantee a prompt return. - **The slow call owns the clock.** A handler parked in a socket read returns when that read returns. If the read has no bound of its own, the handler is stuck for as long as the peer takes, regardless of what the framework decided about the response. ## How the execution model changes the picture | Model | What the timeout can reach | When the handler actually stops | |---|---|---| | Thread-per-request | The dispatch that was waiting on the handler's result | When the current blocking call returns on its own; a bounded call returns sooner, an unbounded one holds the worker indefinitely | | Awaitable handlers (futures or coroutines) | The task representing the request, which can be cancelled | At the next suspension point the handler reaches; code between suspension points, and any blocking call made inline, is not interruptible | | Blocking work offloaded to a side pool | The awaiting side only | The offloaded job runs to completion on its own worker even after the awaiting side has been cancelled, unless the job itself is bounded | The row that surprises people is the last one: offloading keeps the request-handling loop responsive, but it moves the stuck work rather than shortening it. The side pool becomes the new ceiling and the new place to look when saturation appears. ## What a cancellation signal is worth here Many frameworks expose a per-request cancellation signal — a token, a scope or a callback — that is triggered on timeout and on client abort. It is genuinely useful, but only as far as the code that receives it: 1. The handler must **pass it down** to every call that supports it, otherwise the signal stops at the handler's first line. 2. A call that does not support cancellation must be **bounded another way**, with its own timeout, so it returns on its own within a known time. 3. Cleanup on the way out — releasing connections, undoing partial writes — must still run, so a cancellation is an exit path that needs the same care as an error path. A framework that offers the signal and a codebase that never propagates it produce exactly the same behaviour as a framework with no signal at all. ## Diagnosing and fixing it - Compare **in-flight request count** against **response rate**. If responses are flowing while the in-flight count only climbs, abandoned work is accumulating. - Sample where workers are parked. A saturated pool whose stacks all sit in the same outbound call names the missing bound directly. - Add the bound at the source: a read timeout on the outbound call, a statement or query timeout on the database call, a bounded queue in front of the side pool. - Size the pool as a deliberate ceiling rather than as a cushion. A larger pool only buys time; it converts a fast failure into a slow one. The summary an interviewer is listening for: **a timeout is a bound on the answer, and only the work's own bound is a bound on the work.**
- Does moving a blocking call to a side pool solve the capacity problem?It protects the request-handling side but does not shorten the stuck work. The offloaded job keeps running on its own worker even after the awaiting side is cancelled, so the side pool becomes the new ceiling. It is a fix for responsiveness, not for capacity; the capacity fix is still a bound on the call itself.
- Why do a framework's cancellation signals often appear to do nothing?Because the signal has to be carried into each call to have any effect. If the handler never passes it to its outbound client or its query, the signal is set and observed by nobody, and the work continues exactly as before. Cancellation is cooperative all the way down, not a property the framework can impose.
- How would you show, from metrics alone, that timed-out work is piling up?Plot concurrent in-flight requests against completed responses per second. A healthy service holds the two in proportion; when responses continue while the in-flight count only climbs, requests are being answered without their work ending. Worker-pool utilisation staying pinned while throughput falls tells the same story.
- Is raising the worker count a legitimate response to this saturation?Only as breathing room. More workers means more simultaneous stuck calls and more pressure on the dependency that is already slow, which usually deepens the incident. The durable fixes are a bound on the slow call, a limit on concurrent calls to that dependency, and shedding load once the limit is reached.
saying these in an interview costs you the question
- Says the framework kills the handler thread on timeout
- Thinks cancelling an awaitable stops code between suspension points
- Believes offloading a blocking call makes it stoppable
- Assumes a cancellation signal works without being passed down
- Treats a bigger worker pool as the fix for stuck handlers