skip to content

How should you decide what value to set a downstream call's timeout to, and what specifically goes wrong if a caller's timeout is set longer than the callee's own internal processing timeout (or its thread-pool queue wait)?

level: seniorimportance: should knowfreq 50%

answer

  1. derive timeout from p99 latency + margin, not guesswork
  2. caller timeout should be <= callee's own worst-case time
  3. queue-wait timeout distinct from processing timeout
  4. too-generous caller timeout removes back-pressure, worsens cascading failure
  5. Tomcat/Jetty unbounded queue + bounded thread pool classic bug

basics

~20 s

Base the timeout on the callee's real latency (like its 99th-percentile response time) plus margin, not a round number picked by guessing. If the caller waits longer than the callee itself would ever take, the caller just blocks pointlessly once the callee has already given up or hung.

solid answer

~50 s

A well-chosen timeout is derived from the callee's observed latency distribution, often its p99 or p99.9, plus a safety margin, so that the vast majority of legitimate slow-but-successful responses aren't falsely killed, while genuinely hung calls are caught reasonably fast. Setting it by guesswork (a round number like 5 seconds with no data behind it) risks either false timeouts on normal tail latency or being needlessly generous. If the caller's timeout exceeds the callee's own worst-case processing time, or, worse, exceeds how long the callee's server can even hold a request in its queue before its own timeout or circuit fires, the caller ends up blocked waiting for a response the callee already gave up producing, and the caller's overly generous timeout lets the callee's queuing get worse (more concurrent waiting callers) instead of shedding load quickly.

go deeper

for a junior

Should know timeouts shouldn't be totally arbitrary and ideally relate to how long the call normally takes, even without naming percentiles.

for a middle

Should suggest using observed latency (e.g., 'a bit above what it normally takes') to set timeouts and recognize a caller shouldn't wait far longer than the callee ever legitimately needs.

for a senior

Should explicitly reason about p99/p99.9-based timeout selection, the queue-wait vs processing-timeout distinction, and how a too-generous caller timeout removes back-pressure and worsens cascading overload.

for a principal

Should treat timeout calibration as an ongoing, data-driven operational practice (e.g., timeouts derived from live SLO dashboards, revisited as latency profiles shift) and design service-wide defaults (queue-wait limits, load shedding) that don't rely solely on well-behaved callers to apply back-pressure.

## Choosing the value is a statistical decision Choosing a timeout value is fundamentally a statistical decision, not an arbitrary one: the right timeout for a call to a given downstream endpoint should be derived from that endpoint's actual observed latency distribution, commonly its p99 or p99.9 response time under normal, healthy load, plus a modest safety margin, rather than a round number picked because it 'feels long enough.' The reasoning is direct: | Where the timeout is set | What follows | |---|---| | at, say, the median (p50) response time | you'll routinely kill perfectly healthy requests that simply landed in the slower half of the normal distribution, producing a steady background rate of false-positive failures that have nothing to do with the dependency actually being broken | | at p99 plus margin | you accept that roughly 1% of genuinely healthy requests might still be slow enough to threaten the timeout, but the vast majority of legitimate variance is tolerated, while a request that's an order of magnitude slower than even the worst normal case is reliably caught and failed fast | Teams that skip this and hardcode an intuitive round number, 'let's just say 5 seconds', usually end up either: - **far too generous** — masking real problems, allowing threads to block for much longer than the dependency ever legitimately takes, delaying detection of an actual outage; - or, less often but just as harmful, **too tight for a legitimately bursty endpoint** — generating alert fatigue from constant false timeouts that erode trust in the alerting until real incidents get ignored too. ## When the caller waits longer than the callee ever would The specific failure mode of a caller's timeout being longer than the callee's own internal timeout or queueing tolerance is a particularly insidious one because it's invisible from the caller's side until you look closely. Imagine a callee service that itself enforces, say, a 2-second internal processing timeout on any single request (perhaps it queries a database with a 2s query timeout), but the calling service's HTTP client is configured with a 10-second read timeout. When the callee is healthy, this mismatch does nothing visible, most calls finish well under 2s and the caller's generous timeout never matters. But under degraded conditions, say the callee's thread pool starts backing up because of a slow database, new inbound requests to the callee sit in its own request queue before a worker thread even picks them up. If the callee has no queue-wait timeout of its own (only a processing timeout that starts once a thread picks up the work), a request can sit queued for many seconds, and the calling service's overly generous 10-second timeout does nothing to stop it from continuing to wait, worse, it does nothing to stop the caller from sending more concurrent requests into that same backed-up queue, because from the caller's perspective nothing has failed yet. The caller's threads pile up waiting on the callee's queue, and the callee's queue keeps growing because the caller isn't failing fast enough to signal back-pressure. This is exactly the mechanism behind many real 'slow death' cascading failures: a caller with too-generous a timeout doesn't merely tolerate a slow callee, it actively participates in making the callee slower, by not applying back-pressure early enough. ## The fix The fix has two parts. 1. **First, on the caller.** The caller's timeout for a given call should generally be set at or below what the callee's own worst reasonable processing time is, informed by the callee's published SLA or observed p99 latency, so that the caller never usefully waits longer than the callee would ever legitimately take to answer. 2. **Second, and just as important, on the callee.** The callee itself should enforce its own internal bounds, a request-queue-wait timeout in addition to a processing timeout, so it fails fast internally on requests it can't get to promptly, returning a fast, explicit error (which the caller can act on, e.g., fail over elsewhere or open a circuit breaker) instead of silently holding the request in an unbounded queue. ## Where this bug lives in real fleets Systems like Envoy expose both a per-route timeout and separate idle-timeout or queue-related settings for exactly this reason, and Java thread-pool-backed servers (Tomcat, older Spring MVC deployments) have historically been a common source of this exact bug, because a bounded worker thread pool with an unbounded request queue in front of it lets requests wait indefinitely for a free thread with no timeout ever firing on the queueing itself, only on the processing that starts once a thread is finally free. A concrete real-world instance of this class of incident is well documented in postmortems from companies running large Java service fleets on Tomcat and Jetty during traffic spikes, where callers configured with generous multi-second timeouts kept sending traffic into services whose request queues had already grown into the tens of seconds, amplifying rather than containing the outage.

  • Why is setting a timeout at the median (p50) response time of a downstream endpoint usually a mistake?
    By definition, roughly half of all legitimately healthy requests take longer than the median, so a p50-based timeout would routinely kill perfectly normal, successful requests, producing a steady stream of false-positive failures. Basing it on a higher percentile like p99 plus margin tolerates normal variance while still catching genuinely abnormal slowness.
  • What's the difference between a processing timeout and a queue-wait timeout on a server, and why does a service need both?
    A processing timeout bounds how long a worker thread spends actively handling a request once it starts; a queue-wait timeout bounds how long a request can sit waiting for a free worker thread before being picked up at all. A service with only a processing timeout and a bounded thread pool can let requests wait unboundedly long in the queue before a thread ever starts the bounded processing clock, so without a separate queue-wait timeout, total request latency is effectively unbounded even though 'a timeout' technically exists.
  • How does a caller's overly generous timeout actively worsen a callee's overload, rather than just tolerating it?
    Because the caller doesn't fail fast, it keeps waiting on, and doesn't stop sending new requests into, a callee whose queue is already backing up, so the caller provides no back-pressure signal; more concurrent requests keep piling into the same struggling callee, growing its queue further and making the underlying slowdown worse, whereas a tighter, well-calibrated timeout would fail fast and let the caller shed load or fail over sooner.

It's like setting how long you'll wait for a table at a restaurant based on how long tables actually take to turn over on a normal night, not a round number you made up, and if the restaurant itself has no cap on how long a party can wait to be seated, your patience alone won't stop the line from growing forever.

saying these in an interview costs you the question

  • Picks a timeout value by intuition or round numbers rather than from observed latency data
  • Doesn't know the difference between a processing timeout and a queue-wait timeout
  • Believes a longer caller timeout is always 'safer'
  • Can't explain how a too-generous timeout removes back-pressure and worsens cascading overload
  • Assumes a bounded thread pool automatically implies a bounded total wait time

context