skip to content

questions

5

What is a query plan in a federated GraphQL supergraph, and what is each step?

level: juniorimportance: must knowfreq 62%

answer

  1. One document, several services
  2. Work done before any network call
  3. Each step is a whole operation
  4. A tree, not a flat list
  5. Root fetch first, then entity fetches

basics

~20 s

A query plan is the ordered set of subgraph requests a federation router derives from one client document. Each step is a complete GraphQL operation sent to exactly one subgraph, and the router merges the results into a single response.

solid answer

~50 s

A federation router receives one document written against the composed schema, and no single subgraph can answer all of it. Planning is the work it does before any network call: decide which service can supply each selected field, then emit a tree of **fetches**, each one a complete, valid GraphQL operation addressed to exactly one subgraph. Fetches that need nothing from one another are grouped so they can run concurrently; a fetch whose input comes from an earlier result is placed after it. The first fetch normally enters at a root field the client named; later ones are usually entity fetches, carrying objects the router already holds into the subgraph that owns the missing fields. The router then stitches every response into the shape the client asked for. None of this is in the GraphQL specification — it is how a router implements the Apollo Federation contract, and plan vocabulary varies between implementations.

code

graphql · 10 lines
graphql
query BoardHome {
  featuredJobs(limit: 12) {
    id
    title
    employer {
      name
      logoUrl
    }
  }
}

go deeper

for a junior

Be ready to say in one sentence that a router turns one document into several subgraph operations and merges the results. Knowing that a step is a whole operation, not a single field call, is the bar here.

for a middle

Explain the two step shapes — a root fetch entering at a field the caller named, and an entity fetch carrying representations — and why a plan is a tree rather than a flat list of calls.

for a senior

Be able to read a printed plan and say what it will cost: which fetches go out together, which wait, and what the router had to add to the first fetch to make the second one addressable at all.

for a principal

Own the framing that plan shape is a property of your schema and service boundaries, not of the router. Argue for treating plan review as part of schema review rather than as runtime tuning after the fact.

## One document, many owners In a federated setup the graph a client sees is composed from several independently deployed schemas, each owned by a different team and service. A job-board graph might split into a Jobs subgraph (postings, titles, locations), an Employers subgraph (company profiles, logos, verification status) and an Applications subgraph (who applied to what, and when). The caller knows none of that. It writes one operation against the composed schema, where `Job.employer` looks like an ordinary field returning an ordinary object. Something has to bridge the two views, and that something is the router. **Query planning** is the part of its job that happens before any network call: read the incoming document, consult the composed schema for which service can supply each selected field, and produce a small program of subgraph requests that, executed in the right order, yields exactly the response the caller asked for. ## What a step actually is The unit of a plan is a **fetch**: one complete, valid GraphQL operation addressed to exactly one subgraph, with its own variables. That is worth stating plainly, because it is where most misconceptions start. The router does not forward the caller's document to every service and hope. It does not open a call per field. It *writes new documents* — usually a handful, sometimes just one — and each of them is an operation the receiving subgraph could have accepted from any ordinary caller. Steps come in two shapes. The first fetch of a plan is normally a **root fetch**: it enters a subgraph at a root field the caller actually named. Later fetches are normally **entity fetches**: the router already holds some objects, needs fields another subgraph owns, and asks for them through the federation entity entry point, passing those objects as representations — small stubs carrying `__typename` plus the key fields that identify each one. ## The shape of a plan A plan is a tree rather than a flat list, because concurrency matters. Fetches that depend on nothing from each other are grouped so the router can issue them at the same time; a fetch that consumes another's output sits after it. A third kind of node exists purely for addressing: it names a path inside data already fetched — say every element of `featuredJobs`, then that element's `employer` — so the following fetch knows which objects it is about. That node is why a single entity fetch can cover an entire list instead of one fetch per element. Router implementations differ in the vocabulary they print for these nodes; the structure they express is the same. ## A worked example Take the board's home screen, which asks for twelve featured postings with each employer's name and logo. The composed schema makes that one selection set. The plan is two fetches deep. The first fetch asks the Jobs subgraph for the twelve postings and, for each one, an *employer stub*. Note that last part carefully: the caller asked for `employer { name logoUrl }` and the Jobs subgraph can supply neither of those fields. What it can supply is the employer's identity. The second fetch takes all twelve stubs to the Employers subgraph in one entity request, and the router merges the returned names and logos back into the twelve job objects before serialising a response. ## What the caller sees, and does not The response carries the field names and aliases the caller wrote, and nothing else. It does not expose the plan, the intermediate stubs, the injected `__typename`, or which service answered what. That opacity is the point of a composed graph — and it is also why plan-shaped problems are invisible from the calling side and have to be diagnosed at the router, from its plan output or from tracing. ## Specification status Be careful with the word "specified" here, because interviewers listen for it. The GraphQL specification says nothing about routers, subgraphs or plans; as far as it is concerned there is one schema and one execution. Apollo Federation, as a composition specification, defines the directives a subgraph uses to declare its keys and ownership, and the entity entry point a subgraph must expose. Query planning is what a router *does* with that contract. The planning algorithm, the plan format, the node names it prints, and the choice it makes when two subgraphs could both answer a field are implementation behaviour and differ between routers. Saying "the specification defines the query plan" is a wrong answer; saying "the composition specification defines the contract, and the planner is one implementation of it" is the right one.

  • Which part of a query plan is defined by a specification, and which part is not?
    The contract is specified: Apollo Federation defines the directives a subgraph uses to declare keys and ownership, and the entity entry point every subgraph exposes. The plan itself is not. Its algorithm, node vocabulary, and tie-breaking when two subgraphs could answer the same field are router implementation behaviour, and two routers over the same composed schema may produce different plans for the same document.
  • The composed schema shows twelve jobs each with an employer. Why is that not twelve employer fetches?
    Because a plan step is addressed at a path, not at an object. The router walks to the path holding every job's employer stub, collects all of them, and issues one entity request whose representations list carries the whole set. The subgraph answers the list in one operation, and the router merges each result back to the job it came from.
  • Can a plan be a single fetch?
    Yes, and that is the common case for a document whose fields all live in one subgraph. The router still plans — it must confirm every selected field is reachable from one service — but the result is one root fetch and no merging. Plans grow steps only where a selection crosses an ownership boundary.

Like a travel agent handed one itinerary: they book each leg with the airline that flies it, in an order that respects connections, then hand you back one itinerary.

saying these in an interview costs you the question

  • Says the router forwards the caller's document to every subgraph
  • Thinks one plan step is one field resolver invocation
  • Claims the GraphQL specification defines query plans
  • Believes the caller chooses which subgraph answers a field
  • Assumes one subgraph request per item in a list

context

open as a page

In a federation query plan, what forces one subgraph fetch to wait for another?

level: middleimportance: must knowfreq 55%

basics

~10 s

A data dependency. A fetch waits only when its input comes from an earlier fetch's result — usually the key values identifying its objects. Fetches whose inputs are already available run concurrently.

open as a page

Two federated query plans issue the same number of subgraph fetches, but one is far slower. Why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Depth, not count. End-to-end latency follows the longest chain of dependent fetches, because each round waits for the previous one to return. Six fetches in two concurrent rounds cost about two hops; six chained fetches cost six.

open as a page

How do you decide which subgraph should own a field so that query plans stay shallow?

level: principalimportance: should knowfreq 34%

basics

~20 s

Start from the hot operations, not the schema. Place a field where the traffic that matters reaches it in the fewest dependent hops, and treat ownership as a team decision with a real migration cost.

open as a page

Why does a federation router add fields to a subgraph fetch that nobody selected?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

Because the next step has to be addressable. Into each fetch the router injects __typename and the key fields of any object whose remaining fields another subgraph owns, then strips those values back out before serialising the response.

open as a page