How do you set per-fetch timeouts in a federated router so one slow subgraph cannot stall a request?
answer
- The client's clock, not the service's
- Sequential steps add up
- Every step spends from one budget
- Abandoned work keeps running downstream
- The wrapper decides what expiry costs
basics
~20 sGive the whole operation one deadline, derive each fetch's timeout from the budget left rather than a fixed constant, size it from that subgraph's own latency distribution, and propagate the deadline downstream so an abandoned fetch stops working.
solid answer
~50 sStart from the client-visible budget, not from the services. A plan's sequential steps consume it in series, so three steps at 1,850 ms each is a 5,550 ms worst case even though every individual timeout looks reasonable. Give the operation one deadline and derive each fetch's timeout from what remains, so a slow first step shortens the later ones instead of blowing the total. Size the per-subgraph value from that service's own distribution — a documents service at 2,310 ms p99 and a billing service at 940 ms should not share one number. Propagate the deadline downstream, or the router abandons a fetch while the subgraph keeps executing and its database work keeps running. And remember that expiry lands as a field failure: the timeout number and the field's nullability are one decision, because a tight timeout on a Non-Null branch converts slowness into a blank response.
code
pseudocode · 12 linesdeadline = now() + operationBudget # e.g. 2600 ms
for step in plan.steps: # sequential steps
remaining = deadline - now() - reserve
if remaining <= 0:
fail(step, reason = "budget exhausted")
continue
timeout = min(remaining, subgraphCeiling[step.service])
result = fetch(step, timeout, headers = { deadlineMillis: remaining })
if result.expired:
nullFieldsOf(step) # becomes a field failure
addError(path = step.responsePath)go deeper
Know that a router fetch can be given a time limit and that hitting it makes the field fail like any other error. Be able to say why a request with no limit at all is worse than one that gives up early.
Explain the arithmetic: sequential plan steps consume one budget in series while parallel steps overlap, so per-step limits must be derived from what remains rather than set as independent constants.
Show the production judgement — per-service ceilings from measured distributions, deadline propagation so abandoned work actually stops, and the comparison between the router's timer and the subgraph's own that tells a slow service from a tight budget.
Own the latency budget as an organisational contract: who allocates milliseconds across eleven teams, how a new subgraph earns a slice, and how the budget is defended when a team wants to add another sequential hop to a core screen.
## What a timeout is deciding A federation router's unit of work is a **plan step** — one request to one subgraph. A timeout on that step is not really about the subgraph. It is a decision about what the client sees: at expiry the fetch is treated as failed, the fields it was to fill go null, and either the client gets a response that arrives half-empty or, if any of those fields is Non-Null, a much emptier one. So timeout configuration and schema nullability are the same conversation held twice. ## Budget first, service second The usual mistake is per-subgraph timeouts chosen in isolation. In an 11-service legal case-file graph, a plan for one case screen might be: fetch **Cases**, then fetch **Documents** using the case key, then fetch **Deadlines** using a document key from the previous step. Three sequential steps. If each carries a 1,850 ms timeout, the worst-case client-visible latency is 5,550 ms plus router overhead — and no single setting looks wrong. Work the other way. Fix the operation deadline from the product requirement — say 2,600 ms — and give each fetch `remaining budget minus a reserve`. A first step that burns 1,400 ms leaves the second with roughly 1,100, not a fresh 1,850. Parallel steps share the same remaining window, so a step's cost is the slowest branch, not the sum. This is the only arrangement in which the number the client experiences is the number you configured. ## Sizing the per-subgraph ceiling Inside the budget, each subgraph still needs its own ceiling, drawn from its own measured distribution rather than a house default. A billing service with a 940 ms p99 and a documents service with a 2,310 ms p99 have nothing in common; one global 3,000 ms value gives the first a timeout it will never hit — so a hung billing call runs to three seconds before anyone notices — and gives the second a false alarm rate under normal load. A workable starting point is p99 plus a margin, re-derived when the service's profile changes, with two sanity checks. It must be comfortably shorter than any timeout upstream of the router — a client or proxy that gives up first makes every downstream second pure waste. And it must be shorter than the operation deadline, or it can never fire. ## Propagate the deadline, or you have not really timed out Abandoning a fetch stops the router waiting; it does not stop the subgraph working. The subgraph carries on executing, its database queries carry on running, and if the router's caller retries or a caseworker reloads, the same work is now in flight twice. Under load this is how a graph collapses without anyone deploying anything: the router's request count is flat while the subgraph's concurrent execution count climbs. The fix is to tell the subgraph the deadline — a remaining-time or absolute-deadline header on the fetch, with the subgraph cancelling its own work when it passes. GraphQL specifies no such mechanism and neither does the composition specification; it is a plain HTTP convention you establish across your own services, and it needs the subgraph side to actually honour cancellation rather than merely receive the header. ## Expiry is a schema decision A timeout that fires on a nullable contributed field costs one hole in the response. The same timeout on a Non-Null contributed field erases the parent, so a 200 ms tightening can turn a slightly-slow afternoon into blank case screens. Before shortening any timeout, look at the nullability of the fields that step fills. The two knobs must be turned together: a short deadline is only safe behind a nullable seam. ## Observing it Three signals earn their keep. Per-subgraph fetch latency and timeout rate at the router, so you can see which service is eating the budget. The subgraph's own view of the same call — a fetch the router timed out that the subgraph logged as successful means the deadline is too tight, not that the service is broken. And operation-level completeness: the share of responses carrying an error whose `path` starts at a given branch, which is what the user actually felt. ## Where this stops Timeouts bound how long a failure takes to surface. What to do *about* a repeatedly failing dependency — retry policy, backoff, breaker thresholds, bulkheads — is general service-resilience work that applies to any inter-service call and is not specific to a graph. The federation-specific part is the mapping: a plan step is the unit that times out, and each step's expiry lands on a named branch of a response that the client will still be asked to render. ## What an interviewer is listening for Budget arithmetic over the plan's sequential depth, per-service ceilings from real distributions, deadline propagation so the abandoned work actually stops, and the link between the timeout and the nullability of the fields it kills. A candidate who only says "set a timeout on each subgraph" has configured a number, not a behaviour.
- Why is one global subgraph timeout across eleven services a bad default?Because their latency distributions have nothing in common. A value generous enough for the slowest service lets a hung call to the fastest one burn the whole budget before it fires, and a value tuned to the fastest turns normal load on the slowest into a steady stream of failed branches. The ceiling is a property of the service you are calling; the budget is a property of the operation. You need both numbers, not one.
- The router times out a fetch that the subgraph logs as a success. What does that tell you?That the deadline is too tight for the work, or that time is going somewhere the subgraph's own timer does not see — connection acquisition, queueing at the router, serialization of a very large payload. It is not evidence that the service is broken. Comparing the router's per-fetch timer against the subgraph's own execution timer is the cheapest way to separate a genuinely slow service from a badly sized budget.
- Does tightening a timeout ever make the client experience worse?Yes, in two ways. If the field it kills is Non-Null, expiry erases the parent branch, so slightly-slow becomes blank rather than partial. And if the router abandons the fetch without propagating the deadline, the subgraph keeps working, so a tighter timeout raises the concurrent load on the very service that was struggling. Tighten the timeout and the nullability seam together, or not at all.
saying these in an interview costs you the question
- Sets one timeout value for every subgraph
- Ignores that sequential plan steps sum into the client's latency
- Thinks abandoning a fetch stops the subgraph's work
- Configures a timeout longer than the caller's own
- Tightens timeouts without checking the fields' nullability
- Calls deadline propagation part of the GraphQL specification