skip to content

In a thread-per-request web framework, what caps the number of requests in flight, and what happens beyond that cap?

level: middleimportance: must knowfreq 62%

answer

  1. capacity measured in worker-seconds
  2. one worker held for the whole exchange
  3. extra time is queue time, not handler
  4. held-open sockets pin workers
  5. bound the queue, shed early

basics

~20 s

The handler worker count is the ceiling: one request occupies one worker for its whole exchange, waiting included. Beyond it requests sit in the server's connection or dispatch queue, so clients see queue time, then connection refusals or timeouts, while handler timings still look healthy.

solid answer

~50 s

In this model each in-flight request holds one worker from dispatch until the response is written, so **concurrent requests can never exceed the pool size**, regardless of how fast the handlers are. Arrivals beyond that wait in the accept backlog or a dispatch queue; once the queue or backlog is full the server stops accepting, and clients see connection timeouts or resets. The telltale symptom is that per-handler timings look unchanged while end-to-end latency explodes, because the extra time is *queue* time the handler never sees. Two things push the ceiling down invisibly: handlers that spend most of their time waiting on downstream calls (utilisation is low, occupancy is total), and **long-lived responses** such as streamed or held-open connections, which pin a worker for the life of the socket. Once you serve those, the number of open sockets, not the request rate, is what you must size for.

go deeper

for a junior

Remember that one request holds one worker for its whole life, including time spent waiting, so the pool size is the number of requests that can be in progress at once.

for a middle

Explain the queueing chain past the ceiling and why latency rises while handler timings stay flat, and derive throughput from worker count divided by occupancy.

for a senior

Show how you would confirm occupancy exhaustion in an incident and defend the ceiling with bounded queues, downstream deadlines and early load shedding.

for a principal

Own the tradeoff: which endpoints deserve isolated capacity, what fraction of traffic you are willing to shed, and when the ceiling is a reason to change execution model rather than raise a number.

## Occupancy, not utilisation The unit of capacity in a thread-per-request framework is **worker-seconds**, and a request consumes them for its entire lifetime: header parsing, body read, your handler, every downstream wait inside it, response serialisation and the write back to the client. Processor usage is irrelevant to that accounting. A handler that does nothing but wait 500 ms on another service uses almost no processor time and still occupies a worker for 500 ms. That gives a hard ceiling: - **Maximum concurrent requests = the number of handler workers.** Nothing about faster handlers or a bigger machine changes that number by itself. - Throughput follows from occupancy: with `W` workers and an average occupancy of `T` seconds per request, the server tops out near `W / T` requests per second. - The first thing that saturates is therefore usually *downstream latency*, not processor capacity. A downstream service that slows from 50 ms to 500 ms cuts the ceiling tenfold without anything on the server getting busier. ## What happens past the ceiling Requests do not fail immediately; they **queue**, usually in layers, and each layer has its own limit: 1. **The dispatch queue** in front of the workers, if the server has one. Requests wait here for a free worker. 2. **The accept backlog** — connections the operating system has completed but the server has not yet picked up. 3. **Everything upstream** — the load balancer's own queue and the client's connection pool. As each fills, the symptom changes: - First, **latency rises without any handler getting slower**. This is the signature of queueing, and it is why a dashboard of handler durations can stay flat through an outage. - Then, once the backlog is full, new connections are **refused or dropped**, and clients report connection timeouts or resets rather than slow responses. - Finally, the queue itself becomes harmful: requests that have already been abandoned by their client are still dispatched, and the server spends its scarce workers producing responses nobody will read. ## Where the ceiling leaks away | Situation | Effect on the ceiling | |---|---| | Handler waits on a slow downstream call | Occupancy rises; effective throughput falls proportionally | | Response is streamed over a long period | One worker held for the whole stream, not for a few milliseconds | | Connection held open to push events | One worker pinned per open socket in designs that bind the two | | A large request or response body over a slow link | The worker is held for the transfer, not just the logic | | Nested synchronous calls to the same server | Two workers per logical operation; the ceiling halves | The streaming and held-open rows are the ones that surprise teams. A framework whose worker is attached to the exchange rather than to a burst of work cannot serve ten thousand held-open connections with two hundred workers — the two-hundred-and-first client simply waits. Frameworks differ here: some release the worker between requests on a kept-alive connection and only hold it while a request is actually being processed, others keep a worker attached for as long as the connection is open, and that difference decides whether idle keep-alive connections count against your ceiling. Read your server's documentation for which it does before sizing anything. ## Sizing and defending the ceiling - **Measure occupancy, not processor usage.** Average in-flight requests, queue wait time and the ratio of busy to total workers tell you where you are; a processor graph will not. - **Bound every queue.** An unbounded wait queue converts overload into unbounded latency, which is worse than a fast rejection: clients retry, and retries multiply the load that caused the problem. - **Reject early when saturated.** Returning `503` with a `Retry-After` while the queue is deep protects the workers you still have; shedding a fraction of traffic keeps the rest healthy. - **Apply a deadline to downstream calls.** Occupancy is bounded by the slowest thing a handler waits on, so a call with no timeout is an unbounded lease on a worker. - **Separate long-lived work from ordinary requests.** If some endpoints hold sockets open for minutes, give them their own capacity so they cannot consume every worker the fast endpoints need. - **Do not size by intuition.** Raising the worker count raises memory use and context switching, and it cannot help at all when the true bottleneck is downstream; past a point it just moves the queue from the server into the dependency. ## How to recognise it in an incident The pattern is consistent: end-to-end latency climbs and eventually times out, handler-duration metrics stay roughly flat, in-flight or busy-worker counts sit pinned at their maximum, processor usage is unremarkable, and error rates start with client-side timeouts rather than server exceptions. Read that shape as *occupancy exhaustion*, then look for what each occupied worker is waiting on — it is nearly always one dependency, one lock, or one class of long-lived response.

  • Why can handler-duration dashboards look healthy during this kind of saturation?
    Those timers usually start when the handler is dispatched, so all the waiting for a free worker happens before measurement begins. End-to-end or client-side latency includes it; handler duration does not. Track queue wait time and busy-worker count alongside handler duration to see the gap.
  • Does raising the worker count fix a saturated server?
    Only when the bottleneck is genuinely the worker count. If handlers are waiting on a downstream dependency, more workers simply push more concurrent load onto that dependency and make it slower, while costing memory and context switching. Fix the occupancy or protect the dependency first.
  • Why is an unbounded request queue worse than rejecting requests?
    An unbounded queue turns overload into unbounded latency. Clients time out, retry, and add load, while the server still spends workers on requests nobody is waiting for. A bounded queue plus an early `503` keeps served requests fast and makes the overload visible.

saying these in an interview costs you the question

  • Sizes the pool from processor usage instead of occupancy
  • Assumes more workers always increase throughput
  • Thinks queue time shows up in handler-duration metrics
  • Forgets that streamed or held-open responses occupy a worker throughout
  • Leaves downstream calls without timeouts, making occupancy unbounded
  • Prefers an unlimited queue so that no request is ever rejected