skip to content

How would you choose timeout values for a service that sits in the middle of a call chain - an inbound caller, your service, and three downstream HTTP dependencies?

level: principalimportance: should knowfreq 40%

answer

  1. Budget inwards from the caller's deadline, not outwards from dependency latency
  2. Sequential adds, parallel shares wall clock
  3. Worst case = timeout x attempts; retries live inside the budget
  4. Inner timeout < outer timeout; propagate cancellation
  5. Little's law: long timeouts x load = pool exhaustion; add breakers and shedding

basics

~20 s

Start from the caller's deadline, not from downstream latency. Subtract your own processing, then give each downstream a timeout that fits the remaining budget including retries. Every inner timeout must be shorter than the outer one, and connect timeouts stay small.

solid answer

~60 s

Work **inwards from the caller's deadline**, never outwards from what a dependency happens to take. 1. Establish the inbound deadline (an SLO, or an explicit deadline propagated by the caller). 2. Subtract your own processing time and a safety margin - that is the budget for downstream work. 3. Allocate it across dependencies: sequential calls split the budget, parallel calls share the wall clock. 4. Size each per-call timeout as measured p99.9 latency plus margin, but **capped by the remaining budget** - and remember the worst case is timeout x (retries + 1), so retries must fit too. 5. Keep connect timeouts small (hundreds of milliseconds in-datacentre) since they are pure network setup. The cardinal rule is **inner timeout < outer timeout**. If your downstream timeout exceeds the caller's, you hold threads and connections doing work nobody will receive - the caller has already given up. Pair timeouts with cancellation propagation so abandoned work actually stops, and with pool sizing: by Little's law a long timeout under load exhausts pools and turns a slow dependency into your own outage. Add circuit breakers and load shedding so repeated timeouts fail fast rather than queueing.

go deeper

for a junior

Know that each downstream call needs a timeout and that it must be shorter than the time your own caller is willing to wait.

for a middle

Split the caller's deadline across dependencies, distinguish sequential from parallel calls, and account for retries multiplying the worst case.

for a senior

Derive values from measured p99.9 per endpoint, enforce inner-under-outer, propagate cancellation, and tie the numbers to pool sizing and retry budgets.

for a principal

Treat the deadline as a first-class contract: explicit propagation, admission control and shedding when the budget is already spent, circuit breakers, degraded responses, and periodic re-derivation as latency distributions move.

## Deadlines, not durations The common mistake is to set each timeout from the dependency's observed latency in isolation. That produces a system where the total possible latency is the sum of everyone's generosity, usually far beyond what any caller will wait for. The correct model is a **deadline budget**: a request enters with a finite amount of time, and every component spends from it. Ideally the deadline is propagated explicitly - a header or RPC metadata carrying "this expires at T" - so each hop computes its remaining budget rather than restarting a fresh clock. Without propagation, approximate it from the inbound SLO, but recognise the weakness: a request that was already queued for 300 ms starts your service with less time than you assume. ## Allocating the budget Suppose the inbound SLO is 1000 ms at p99 and your own compute costs 50 ms. Roughly 900 ms remains after a margin. - **Sequential dependencies** consume the budget additively. Three sequential calls at 300 ms each already exhaust it, and a retry on any of them blows it. - **Parallel dependencies** share wall clock: three concurrent calls can each be given close to the full remaining budget, and the slowest determines the outcome. Parallelising is often the cheapest way to buy budget. - **Optional dependencies** should get a deliberately small slice and a fallback, so an enrichment call cannot consume the budget of the critical path. Size each timeout from measured latency - p99.9 plus margin, per endpoint rather than one global value, since a search endpoint and a health check have nothing in common - and then cap it by the remaining budget. If p99.9 does not fit in the budget, that is a design signal: cache, precompute, parallelise, or degrade, rather than quietly setting a timeout you know will be exceeded. ## Retries inside the budget Worst-case time for a call is `timeout x (attempts)`. A 500 ms timeout with two retries is 1500 ms, so it must be budgeted as 1500 ms, not 500. Practical patterns: retry only the fastest dependencies; retry with a *shorter* timeout on the second attempt; or use hedged requests (issue a second attempt after p95 and take the first answer) where the operation is safe to duplicate. Always check the remaining deadline before starting a retry - retrying with 50 ms left is pure waste. ## The inner-outer rule and its consequences If your downstream timeout is longer than your caller's, then after the caller gives up your service keeps a thread, a connection, and downstream capacity busy producing an answer that will be discarded. Under load this compounds: work-in-progress grows, pools saturate, and the system does more useless work exactly when it can least afford it. Enforce inner < outer at every hop, and implement **cancellation propagation** so that when the inbound connection is closed or the deadline passes, in-flight downstream calls are actually aborted. ## Coupling with concurrency Little's law again: concurrency = arrival rate x latency. Timeouts set the worst-case latency, and therefore the worst-case concurrency you must provision. At 200 requests per second, a 10-second timeout implies up to 2000 concurrent in-flight requests during a downstream stall - far beyond a typical thread pool or connection pool. Short timeouts are not only about user experience; they are the mechanism that bounds resource consumption. Timeout, pool size, and retry policy must be chosen as one set, and the arithmetic should be written down. ## Failing fast instead of waiting Timeouts are the last line of defence, not the first. Add: - **Circuit breakers** - after a threshold of timeouts, fail immediately for a cooldown so you stop spending the budget on a known-dead dependency. - **Load shedding / admission control** - reject at the edge when queue delay already consumes the budget, rather than accepting work that will time out anyway. - **Graceful degradation** - a cached or partial answer within the deadline beats a perfect answer after it. ## Operating the choice Document every timeout with its justification ("p99.9 = 180 ms, timeout 300 ms, budget slice 350 ms"), export the deadline-exceeded rate per dependency, and revisit whenever latency distributions shift - after a schema change, a region move, or a traffic pattern change. Values chosen once and never revisited become either too tight (spurious failures after a legitimate latency increase) or so loose that they never fire before the caller gives up, which is functionally the same as having no timeout at all.

  • What goes wrong when a downstream timeout is longer than the timeout your own callers apply to you?
    Once the caller gives up, your service keeps a thread, a connection and downstream capacity occupied producing a result nobody will read. Under load that work-in-progress accumulates, pools saturate, and useless work crowds out useful work exactly when capacity is scarce. Inner timeouts must always be shorter than outer ones, and cancellation should propagate so abandoned work is actually stopped.
  • How does deadline propagation improve on each service setting its own fixed timeouts?
    A propagated deadline tells each hop how much time actually remains, rather than each hop restarting a fresh clock from an assumed total. That prevents the sum of locally reasonable timeouts from exceeding the caller's patience, lets a hop skip work it cannot finish in time, and makes queue delay visible as consumed budget. Fixed local timeouts silently break whenever a request was delayed before reaching you.
  • How do timeout values interact with connection and thread pool sizing?
    By Little's law, concurrency equals arrival rate times latency, and the timeout sets worst-case latency. At 200 requests per second a 10-second timeout implies up to 2000 concurrent in-flight requests during a stall, which no ordinary pool can absorb. Short timeouts are therefore a capacity control as much as a latency control, and pool size, timeout and retry count must be chosen together.

A connecting itinerary: the whole journey has a fixed arrival time, so each leg gets a slice of it. A leg allowed to run longer than the total trip is not a plan, it is a guarantee of a missed connection - and the passenger has already left.

saying these in an interview costs you the question

  • Deriving timeouts from what a dependency happens to take instead of from the caller's deadline
  • Setting one global timeout for every endpoint regardless of its latency profile
  • Ignoring that the worst case is timeout multiplied by the number of attempts
  • Allowing an inner timeout to exceed the outer one, so work continues after the caller has left
  • Treating timeouts as the whole strategy, with no circuit breaker, cancellation or degraded response

context