skip to content

A web service's request timeout fires on schedule, yet its in-flight count climbs until every worker is busy. How would you diagnose and fix it?

level: seniorimportance: should knowfreq 50%

answer

  1. answered but not finished
  2. in-flight versus completions
  3. sample where the workers are parked
  4. bound the call, not the pool
  5. cap concurrency per dependency, then shed

basics

~20 s

Timed-out handlers are still running. Confirm by comparing in-flight count with response rate and sampling where workers are parked, then bound the slow call itself, propagate cancellation into it, and cap concurrency to that dependency rather than enlarging the pool.

solid answer

~40 s

The symptom is the signature of a response-side timer with no matching bound on the work. Confirm it by plotting **concurrent in-flight requests** against **responses per second**: if responses keep flowing while in-flight only climbs, handlers are being abandoned but not ending. Then sample worker state — if most stacks sit in the same outbound call or query, that call is the missing bound. Fixes, in order: give the slow call its own timeout so it returns by itself; propagate the framework's per-request cancellation signal into it, and into the client-abort path too; cap concurrent calls to that dependency and shed load beyond the cap so a slow dependency cannot consume the whole pool. Enlarging the pool is the fix to avoid — it buys minutes and pushes more load onto something already failing.

go deeper

for a junior

The takeaway is that a timeout ends the wait, not the work, so a service can answer every request on time and still run out of workers.

for a middle

Be ready to explain the two metrics that expose it, in-flight against completions, and why a bound on the slow call is what actually returns a worker.

for a senior

Walk the full loop: confirm from metrics, sample worker state, bound the call, propagate cancellation, cap concurrency to the dependency, and verify that in-flight falls back.

for a principal

Argue the policy: bounded concurrency and deliberate shedding as the default posture, with pool growth as a last resort, and ownership of end-to-end latency budgets rather than per-service guesses.

## Reading the symptom Two facts are given: responses are produced on time, and the in-flight count keeps climbing. Together they say the framework is meeting its promise to callers and nothing is meeting a promise about the work. The timer ends the *waiting*; the handler ends when its current call returns. Every request that times out therefore leaves behind an execution slot that no metric derived from responses can see. This is why the incident looks contradictory on a dashboard: latency percentiles are flat at the timeout value, error rate is a clean band of timeout responses, and the process is nonetheless approaching a hard ceiling. ## Diagnosis, in order 1. **Compare in-flight against completions.** Concurrent in-flight requests plotted against responses per second. Healthy services keep them proportional; a climbing in-flight count with steady responses means work is accumulating behind answered requests. 2. **Sample where workers are parked.** Take several samples of worker state. A saturated pool whose stacks converge on one outbound call or one query names the missing bound immediately. 3. **Check the layering of timers.** If the layer in front of the framework fires first, the framework never runs its own error stage, and you lose both the mapped response and the log line that would have explained it. 4. **Check whether cancellation is wired at all.** If the framework exposes a per-request cancellation signal, find out whether anything downstream receives it. A signal nobody reads behaves identically to no signal. 5. **Check the side pools.** If blocking work is offloaded, the visible request path can look healthy while the offload pool is the real saturated resource. ## Fixes, strongest first | Fix | What it changes | Why it works | |---|---|---| | A bound on the slow call itself | The call returns by itself within a known time | The only mechanism that reliably returns the execution slot, independent of any signal | | Propagating the cancellation signal | Timeout and client abort reach the call in progress | Ends work early where the call supports cancellation, and covers aborts as well as timeouts | | Bounded concurrency per dependency | Caps how much of the pool one dependency can occupy | Contains a slow dependency instead of letting it consume the whole service | | Shedding load past the cap | Fast rejection instead of an unbounded queue | Keeps latency honest and keeps the healthy paths of the service alive | | Enlarging the worker pool | More simultaneous stuck calls | The fix to avoid: it delays saturation and adds pressure to the failing dependency | ## The part candidates miss A request timeout and a bound on the work are **not** the same control, and configuring only the first is the defect here. The distinction is worth stating explicitly in an interview: - The request timeout answers **the caller**. It is the outer promise, and it should sit inside whatever the layer in front allows. - The per-call bound answers **the server**. It is what keeps the execution slot from being held indefinitely. - The cancellation signal is the **bridge** between them, and it is only as good as its propagation. It also carries client aborts, so wiring it once serves both causes of abandoned work. ## Verifying the fix - In-flight count should now fall back toward the response rate rather than ratcheting upward. - Worker-state samples should stop converging on a single call. - The dependency's own load should drop once concurrency is capped, which often improves its latency and dissolves the original slowdown. - Rejections appear at the cap during a dependency incident. That is the intended behaviour: a bounded fast failure for some traffic instead of an unbounded brownout for all of it. ## Guardrails to leave behind - Treat any outbound call without a bound as a defect, and make that reviewable rather than remembered. - Alert on in-flight or worker utilisation, not only on latency and error rate, since the first two are the metrics that move during this failure. - Keep aborts and timeouts on distinct counters so the next incident can be told apart from a wave of impatient callers at a glance.

  • Why is raising the worker count the wrong first move here?
    Because each extra worker becomes another simultaneous call into a dependency that is already slow, which usually degrades it further. It postpones saturation instead of preventing it and converts a fast, visible failure into a long brownout. Capacity should be added only after the work has a bound.
  • Which metric would have caught this before saturation?
    Concurrent in-flight requests, or worker-pool utilisation, plotted next to completions. Latency and error rate both look deceptively clean during this failure because they are measured when a response is written, and the abandoned work never writes one. In-flight count is the number that actually moves.
  • How do client aborts fit into the same failure?
    They produce the same shape: the exchange ends while the handler keeps working. If the cancellation channel is wired for timeouts it usually covers aborts too, which matters most on slow endpoints, because the callers most likely to abort and retry are the ones being served slowly.
  • What do you tell callers while the cap is rejecting requests?
    Reject quickly with a clear retryable signal rather than holding the request in a queue. A caller that is told to back off can shed its own load; one that is left waiting builds the same queue on its side. Pair it with jittered backoff guidance so the recovery is not synchronised.

saying these in an interview costs you the question

  • Raises the worker pool size as the primary fix
  • Insists the timeout must already have stopped the work
  • Trusts flat latency graphs as proof the service is healthy
  • Sets a request timeout but leaves outbound calls unbounded
  • Lets one slow dependency draw from the whole shared pool
  • Queues excess load indefinitely instead of rejecting past the cap