skip to content

You own a user-facing API with a 500 ms latency target that fans out to several downstream services, some of which call further services. How would you assign timeout and deadline budgets across that call graph?

level: principalimportance: should knowfreq 30%

answer

  1. allocate top-down from the user number; subtract reserves
  2. sequential splits, parallel shares
  3. slice sized from measured p99 + margin
  4. enforce on dequeue; shed expired; cap accepted deadlines
  5. classes of traffic; optional deps degrade to defaults

basics

~20 s

Start from the user-visible budget and divide it downward: subtract fixed overhead and a cleanup reserve, split the rest across sequential hops, share it among parallel ones, and propagate the remainder as a deadline every hop enforces and sheds on. Size each hop from its measured p99, not from a guess.

solid answer

~60 s

Treat the 500 ms as a budget to be **allocated**, not a value to be repeated. 1. **Reserve first.** Subtract edge overhead (TLS, auth, serialization), a slice for the API's own work, and a small cleanup and response-assembly reserve. Perhaps 380 ms remain for downstream work. 2. **Allocate by structure.** Sequential hops split the budget in series; parallel hops each get the full remaining budget since they overlap. The critical path is what you are budgeting. 3. **Propagate a deadline**, not per-hop constants. Each hop uses min(local cap, remaining), and rejects immediately when the deadline has already passed — checked on dequeue, so queueing counts. 4. **Size from data.** A hop's cap should sit above its p99 including retries; below it you manufacture failures, far above it you waste budget. If measured p99s do not fit the target, that is a design finding, not a config problem. 5. **Class the traffic.** Interactive, batch, and background get different budgets. 6. **Degrade.** Decide per dependency: optional ones return a default when their slice expires; only the critical path fails the request. Then measure: timeouts by hop, deadline-expired-on-arrival, and budget headroom.

go deeper

for a junior

Show the basic subtraction: the user budget minus overhead, divided across the calls on the path, with each call getting a limit that fits.

for a middle

Distinguish sequential from parallel allocation, derive limits from measured p99, and propagate the remaining time rather than fixed constants.

for a senior

Add server-side enforcement and shedding on dequeue, retry and hedge budgets inside the allocation, per-dependency degradation policy, and headroom metrics.

for a principal

Treat budgets as a cross-team contract with declared need and promise, mechanically checked; classify traffic; decide which dependencies leave the request path entirely; and call out when the target is structurally unachievable rather than tuning configuration.

## Budgeting is subtraction, not repetition The common failure is picking one plausible number and configuring it everywhere. Budgets are allocated top-down from the only number that exists for a real reason: what the user (or the calling contract) will tolerate. Everything else is derived. ```text 500 ms user target - 40 ms edge overhead (TLS, auth, parsing, serialization) - 40 ms own work + response assembly - 40 ms reserve (cleanup after timeout, jitter, safety margin) = 380 ms available for downstream work on the critical path ``` The reserve matters more than people expect: post-timeout cancellation and cleanup happen *after* the budget is spent, and without slack the API blows its own target while tidying up. ## Allocate by graph structure Walk the call graph and budget the **critical path**. - **Sequential** calls divide the remaining budget in series: auth 40 ms, then profile 120 ms, then recommendations 180 ms, with 40 ms slack. - **Parallel** calls each receive the same remaining budget, because they overlap; the hop is bounded by the slowest, and a fan-out of ten with a 200 ms slice costs 200 ms, not 2 s. - **Nested** hops receive what is left when they are called, not a fixed constant — which is exactly why a deadline is propagated rather than a duration configured. Size each slice from measured behaviour: a dependency's p99 (including its own internal retries) plus margin. A cap below its p99 manufactures failures for requests that would have succeeded; a cap far above it wastes budget that a sibling could have used and delays failure detection. If the sum of realistic p99s along the critical path exceeds the target, no configuration fixes it — the answer is caching, precomputation, parallelizing sequential hops, dropping a dependency, or renegotiating the target. ## Enforce the deadline on the server side too A deadline propagated but not enforced is documentation. Each service should read the incoming deadline, check it **when the request is dequeued** (so time spent queued is charged), and reject expired requests instantly with a distinct "deadline exceeded" outcome. This is among the cheapest and most effective load-shedding mechanisms available: under overload the queue drains itself of work nobody is waiting for, instead of the server spending its capacity producing answers that will be discarded. Services should also cap the deadline they will accept. A caller handing down 60 s to a service whose work is always sub-second is either buggy or abusive, and honouring it ties up capacity. ## Classes of traffic, not one global number Interactive requests, asynchronous jobs and bulk backfills have genuinely different budgets. Model them as named classes (for example interactive 500 ms, background 30 s, batch minutes) carried with the request, so a shared downstream service can prioritize and shed by class. Without classes, protecting interactive latency forces you to give batch work the same tiny budget, and it fails constantly. ## Retries, hedging, and degradation live inside the budget Retries consume the same allocation: a hop with 180 ms might attempt twice at ~80 ms with jittered backoff, or once at 180 ms — a decision the retry loop makes from the remaining time rather than from a hard-coded count. Hedging likewise shares the deadline, never extends it. Decide per dependency what expiry *means*: - **Critical** — its failure fails the request. Give it the generous slice. - **Optional** — recommendations, badges, personalization. Give it a short slice and a default value on expiry, so the response degrades rather than fails. - **Fire-forward** — analytics, audit. Move it off the request path entirely into a queue with its own lifetime, so it never competes for the user's budget. This is usually where the largest wins are: much of a fan-out is optional, and treating it as such buys the critical path a bigger slice. ## Guardrails and review Guard against the classic inconsistencies: a downstream timeout larger than the caller's remaining budget (guaranteed wasted work), a connect timeout longer than the whole request budget, retry counts multiplied across layers, and any unbounded call at all — clients without a default timeout should fail a build or a lint rule, not a review. ## Measure and iterate Export, per hop: p50/p99 latency, timeout rate, deadline-expired-on-arrival rate, and headroom (slice minus observed p99). A hop whose headroom is negative is a scheduled incident; a hop whose headroom is enormous is holding budget hostage from the critical path. Re-derive the allocation whenever the topology changes, and treat the budget as an interface contract between teams: each service publishes the deadline it needs and the deadline it promises, so the graph can be checked end to end rather than negotiated incident by incident.

  • The measured p99s along your critical path already add up to more than the 500 ms target. What do you do?
    Recognize it as a design problem, not a configuration problem — no allocation of timeouts can make a chain fit a budget it structurally exceeds. The levers are architectural: parallelize hops that are only accidentally sequential, cache or precompute a stage, remove or make optional a dependency on the critical path, move work off the request path, or renegotiate the target. Setting timeouts below the p99 instead just converts slow requests into failed ones.
  • How do you keep budgets consistent as teams independently change their services?
    Make the budget an explicit contract rather than scattered configuration: each service declares the deadline it needs to serve its promise and the deadline it promises to callers, and those declarations are checked mechanically — in a service catalog or at build time — against the graph. Pair it with runtime enforcement and headroom metrics per hop, so drift shows up as shrinking headroom before it shows up as an incident, and forbid unbounded clients in CI so a new call cannot silently join the graph without a budget.

saying these in an interview costs you the question

  • Configuring one timeout value everywhere instead of allocating a budget top-down.
  • Giving parallel calls a divided budget as if they were sequential, or giving sequential calls the full budget each.
  • Setting a hop's timeout below its measured p99, manufacturing failures for requests that would have succeeded.
  • Propagating deadlines but never enforcing them server-side, so expired work is still processed.
  • Leaving no reserve for cleanup and response assembly, so the API misses its own target while unwinding.
  • Treating an optional enrichment dependency as critical, letting it fail the whole request.

context