Compare giving each call in a request chain its own fixed timeout with carrying a single absolute deadline through the whole chain. What breaks with per-call timeouts, and what does a propagated deadline give you?
answer
- timeouts add, deadlines don't
- deadline = one instant, carried with the request
- remaining = deadline - now; use min(remaining, cap)
- expired on arrival -> reject immediately
- cross-machine: send remaining duration, not wall-clock
basics
~20 sPer-call timeouts add up, so the total is bounded only by the sum, not by what the user will wait. A deadline is one point in time passed down the chain: every hop computes the time remaining, so the whole request is bounded and any hop can fail fast when the time is already gone.
solid answer
~50 sA **per-call timeout** is a duration each hop applies independently. With five hops of two seconds each, the worst case is ten seconds — the guarantee composes by addition, which is not a guarantee at all. Worse, a late hop can start work the original caller has already abandoned. A **deadline** is a single answer to "by when must this whole request be done?", carried with the request. Each hop computes `remaining = deadline - now` and uses `min(remaining, its own cap)`. That gives: - **End-to-end bounding** — total latency is the deadline, regardless of depth. - **Fail-fast** — a hop whose remaining time is already zero or clearly too small rejects immediately instead of burning capacity on work nobody will read. - **Correct budget for retries and fan-out**, since every decision consults one shared clock. Across machines propagate the **remaining duration**, not an absolute wall-clock timestamp, unless clocks are tightly synchronized — skew silently shrinks or inflates the budget.
go deeper
Contrast the two clearly: a timeout is per call and they add up; a deadline is one time-by-which for the whole request, passed along.
Show the arithmetic, the min(remaining, cap) rule, and fail-fast on an already-expired deadline.
Cover propagation mechanics — automatic attachment in the client library, duration versus timestamp and skew, counting queue time, reserving budget for cleanup — and the wasted-work capacity argument.
Treat it as a platform contract: deadlines are mandatory request metadata, servers enforce and shed on them, budgets are allocated per class of traffic, and unbounded calls are a build-time defect.
## Two different things A **timeout** is a duration attached to one operation: "this call gets at most 2 seconds". A **deadline** is an instant attached to the whole request: "everything for this request must be done by T". They answer different questions, and a system needs both, but only one of them composes. ## Why per-call timeouts do not compose Consider a request that traverses five services, each with a 2-second client timeout: ```text A --2s--> B --2s--> C --2s--> D --2s--> E worst case observed by the user: ~10s (plus retries, plus queueing) ``` Each hop is individually reasonable and the total is not. Add one retry per hop and the worst case triples. Three problems follow. 1. **No end-to-end guarantee.** The bound is the sum along the deepest path, which changes whenever anyone adds a hop — so an unrelated team's refactor can blow your latency contract. 2. **Wasted work.** Suppose A gives up after 2 seconds. C, D and E know nothing about that and keep working, holding threads, connections and locks for a result that will be discarded. Under load this is how a slow dependency turns into a capacity collapse: the system is busiest doing work nobody is waiting for. 3. **Nonsense at the leaves.** A leaf service can start a 2-second query at the moment 100 ms of user patience remains. It cannot know, so it cannot refuse. ## What a deadline changes The caller computes the deadline once, at the edge: `deadline = now + budget`. Every downstream call carries it. At each hop: ```text remaining = deadline - now if remaining <= minUsefulTime: fail fast ("deadline exceeded") callTimeout = min(remaining, localCap) ``` The end-to-end bound is now the deadline itself, independent of how many hops exist. Any hop can also **reject before starting**: a request that arrives with 5 ms left is dropped instantly, freeing capacity for requests that can still succeed. This early rejection is one of the most effective load-shedding mechanisms in a distributed system, and it only exists if the deadline travels with the request. The local cap is still worth keeping. It expresses "this call has never legitimately taken more than X", which catches a hung dependency even when the overall budget is generous, and it protects an inner service from a caller that hands down an absurdly long deadline. ## Propagation mechanics Deadlines must ride on the request as metadata — an RPC header, a message attribute, an ambient context object — and be attached automatically by the client library rather than passed by hand, or someone will forget and that hop becomes unbounded. The clock question matters. An **absolute timestamp** is only meaningful if the sender's and receiver's clocks agree; a few hundred milliseconds of skew silently shortens or lengthens the budget, and the failures are baffling. Propagating the **remaining duration** avoids skew but ignores network transit time, so the budget inflates slightly at each hop; senders typically subtract an estimate of one-way latency. Within a single process, an absolute deadline read from a monotonic clock is strictly better than a duration, because it is immune to clock adjustments and to the time the request spent queued. **Queue time counts.** A request sitting 400 ms in an inbound queue has already spent 400 ms of the budget. Checking the deadline when work is dequeued — not only when it is enqueued — is what makes an overloaded server shed its backlog instead of processing stale requests. ## How the deadline splits inside a hop Sequential sub-calls share the remaining budget in series: each takes what it needs and passes on what is left. Parallel sub-calls each get the same remaining budget, since they run concurrently — the hop is bounded by the slowest. Always reserve a slice for the hop's own work: serialization, response assembly, and post-timeout cleanup. If the entire budget is handed to the downstream call, the hop has no time left to produce an answer, or even to clean up. ## Relationship to cancellation and scopes A deadline is only a number; something must act on it. Operationally, "the deadline has passed" is delivered as a cancellation of the region doing the work — the whole subtree of tasks for that request — which is exactly what a task scope names. Deadline propagation and scoped cancellation are two halves of the same mechanism: the deadline says *when*, the scope says *what*.
- Should you propagate an absolute deadline timestamp or the remaining duration between services?Across process or machine boundaries, propagate the remaining duration unless clocks are tightly synchronized, because clock skew silently corrupts an absolute timestamp in either direction. The cost is that network transit is not accounted for, so senders usually subtract an estimated one-way latency. Within a single process, use an absolute deadline read from a monotonic clock: it is immune to clock adjustments and correctly charges time spent waiting in queues.
- If a deadline bounds the whole request, why keep per-call timeout caps at all?They encode local knowledge the deadline does not have: a call that normally completes in 20 ms has clearly hung at 2 s, even when the request still has 8 s of budget left. Capping it fails fast, frees the connection, and permits a retry or a fallback within the remaining budget. The cap also defends an inner service against a caller that hands down an unreasonably long deadline. The effective timeout is the minimum of the two.
A per-call timeout is each leg of a journey promising to take under two hours. A deadline is a flight you must catch: every leg checks the clock on the wall, and if the plane has already left there is no point boarding the train.
saying these in an interview costs you the question
- Believing per-hop timeouts bound end-to-end latency — they compose by addition along the deepest path.
- Propagating an absolute wall-clock timestamp between machines without considering clock skew.
- Handing the entire remaining budget to a downstream call, leaving no time for the hop's own work or cleanup.
- Checking the deadline only when a request is accepted, never when it is dequeued, so an overloaded server processes a stale backlog.
- Thinking a deadline stops work by itself — it is just a number until something cancels the work region.