skip to content

Two federated query plans issue the same number of subgraph fetches, but one is far slower. Why?

level: seniorimportance: should knowfreq 48%

answer

  1. Count is load, shape is latency
  2. The critical path, not the total
  3. Concurrent rounds overlap; chained rounds add
  4. Read nesting in a trace, not a list
  5. The lever is field placement

basics

~20 s

Depth, not count. End-to-end latency follows the longest chain of dependent fetches, because each round waits for the previous one to return. Six fetches in two concurrent rounds cost about two hops; six chained fetches cost six.

solid answer

~50 s

Fetch count is a load number; **plan depth** is the latency number. Fetches grouped in one concurrent round overlap in time, so a round costs roughly its slowest member, not the sum of its members. Chain those rounds and the costs add. A plan of six fetches arranged as two rounds is roughly two subgraph latencies deep; the same six chained is six, and on a job-board graph that was the difference between a 190 ms and a 540 ms p95 at a 1,200-request-per-minute peak. Depth comes from dependency: every entity fetch waiting on key values adds a round. Diagnose it from the plan itself, or from a trace where spans nest rather than sit side by side. And fix it in the schema: field placement and ownership set the depth, so no amount of router tuning shortens a chain the composition made necessary.

code

pseudocode · 9 lines
pseudocode
Plan A  (2 rounds, ~120 ms)
  Parallel [ Fetch Jobs, Fetch Applications ]
  Parallel [ Fetch Employers, Fetch Candidates ]

Plan B  (4 rounds, ~240 ms)
  Fetch Jobs
  Fetch Employers      (needs employer keys from Jobs)
  Fetch Billing        (needs account keys from Employers)
  Fetch Trust          (needs contract keys from Billing)

go deeper

for a junior

Recall the headline only: calls that go out together cost about one call's time, and calls that wait for each other add up. Counting calls is not the same as measuring how long a request takes.

for a middle

Explain why a round costs its slowest member and why chains add, and be able to point at the dependency that created each sequential edge in a plan you are shown.

for a senior

Diagnose it properly: read plan shape or trace nesting, notice that every subgraph p95 can stay flat while the composed query degrades, and identify the schema change that added the round.

for a principal

Own the budget. Decide which operations get a depth limit, get that limit checked at composition time in the pipeline, and weigh a schema migration's risk against the latency it actually buys back.

## The metric people reach for is the wrong one When a federated query gets slow the first instinct is to count subgraph calls, because that number is easy to get and it feels like the cost. It is a real number — it is what your subgraphs are being asked to serve — but it is not what the caller waits for. The caller waits for the **critical path**: the longest chain of fetches where each one cannot start until the one before it has returned. A quick arithmetic sketch on a job-board graph makes the point. Suppose each subgraph answers in about 60 ms. Six fetches issued as two concurrent rounds of three cost roughly 120 ms plus router overhead. The same six fetches chained one after another cost roughly 360 ms. Identical load on the subgraphs, identical fetch count on the dashboard, three times the latency. That is why "we reduced calls" is not automatically a performance win, and why "the plan has more fetches now" is not automatically a regression. ## What creates depth Every sequential edge in a plan is a data dependency. The router cannot address an entity fetch without the key values that identify its objects, so each hop from an object to fields another subgraph owns is one more round. Chain three such hops — a job's employer, that employer's billing account, that account's contract terms, each in a different service — and you have a four-round plan from a selection set that looks innocuous in the composed schema. Nothing about the caller's document announces this. Depth is a property of where the fields live. A second source is authored: a field declared to require sibling data owned elsewhere forces the router to fetch that data before it can ask for the field. That is a correctness mechanism and it is often the right call, but it buys correctness with a round trip and it should be recognised as a depth decision. ## A concrete regression The pattern that produces an incident looks like this. A screen on the job board ran a plan two rounds deep and a p95 near 190 ms. A schema change moved one field — an employer's verification badge — out of the Employers subgraph and into a new Trust subgraph keyed by the employer's account, which itself had to be reached through Employers. The composed schema still exposed `Job.employer.verified`; nothing about the caller's document changed; no caller was redeployed. But the plan went from two rounds to four, and at the 1,200-request-per-minute peak the p95 moved to about 540 ms while every subgraph's own p95 stayed flat. Each service looked healthy in isolation, which is exactly why this class of regression takes a while to find. ## How to see it Three tools, in order of directness. First, ask the router to show the plan for the operation; the nesting *is* the answer, and reading it takes seconds. Second, look at a distributed trace and pay attention to the **shape**, not the list: spans that start together and overlap are one round; spans that begin only when their predecessor ends are a chain. A trace UI that renders subgraph spans as a flat list of durations is actively misleading here, because two plans with identical span lists can be two rounds or six. Third, if you are instrumenting your own metric, record the plan's depth per operation rather than its fetch count — it correlates with latency in a way the count does not. ## How to fix it The lever is the schema, not the router. Options, roughly in order of how often they are the right answer: - **Move the field to where it can be reached in fewer hops**, if a team boundary genuinely allows it. - **Let the intermediate subgraph supply the field inline** when it already has the value in hand — federation has a directive for promising exactly that, which collapses a round when the promise holds. - **Duplicate a small, stable field** across the two subgraphs that need it, accepting the consistency cost in exchange for removing a hop. - **Reshape the hot operation** so the deep branch is not on the critical path of a first paint, for example by moving it behind a second request that renders later. And sometimes the answer is to accept it. A four-round plan on a rarely-visited admin screen is not a problem worth a schema migration. Depth budgets belong on the operations that actually carry your traffic; spending schema-change risk on the rest is how you turn a performance exercise into an outage.

  • Does that mean fetch count never matters?
    It matters, but as a load and reliability figure rather than a latency one. A wide round hits several subgraphs at once, consumes router concurrency, and multiplies the chance that one of them is the slow or failing one. So width is what you watch for capacity and blast radius, and depth is what you watch for response time. They are two budgets, not one.
  • How would you catch a depth regression before it reaches production?
    Plan the graph's top operations at composition time and assert on their depth. Composition happens in a pipeline before deployment, so the planned shape of a known set of documents can be computed and compared against the previous supergraph. A change that pushes a hot operation from two rounds to four then fails a check, rather than surfacing later as a latency alert nobody can attribute.
  • A trace shows four subgraph spans of 60 ms each and a total of 250 ms. What can you conclude?
    That the four ran essentially in a chain, not concurrently — the total is close to their sum. If they had formed one round the total would sit near 60 ms plus router overhead. The durations alone could not tell you this; only the start times and the nesting can, which is why the shape of a trace matters more than its span list.

Six errands run by six people in parallel take one errand's time; the same six run by one person, each depending on the last, take six. The count on the list is identical.

saying these in an interview costs you the question

  • Optimises for fewer fetches while leaving the chain long
  • Blames subgraph latency when every subgraph p95 is flat
  • Thinks total fetch count predicts end-to-end latency
  • Assumes concurrent fetches add their latencies together
  • Reads a flat list of trace spans without checking nesting

context