skip to content

In a federated graph, why does the router batch entity fetches instead of one call per item?

level: juniorimportance: must knowfreq 56%

answer

  1. One list, one downstream service
  2. Count network calls, not rows
  3. Keys collected before the call goes out
  4. A list-valued representations argument
  5. One call per plan step, not per item

basics

~20 s

One call per item is an N+1 across services, where each call is a full subgraph request. The router instead collects every item's key into a single _entities call carrying a list of representations, so a 143-item list costs one request.

solid answer

~40 s

A federated plan often resolves a list in one subgraph and some of that list's fields in another. Done naively that is one network call per item — the N+1 problem, except each call is a subgraph HTTP request with its own serialization, auth check and execution, so it is far costlier than a repeated database query. Apollo Federation's subgraph specification avoids it by making the reserved `_entities` root field take a **list**: the router gathers the key fields of all 143 items into one fetch, sends a single request, and splices each returned element back into its position. Cost becomes one call per subgraph per plan step, independent of list length. It bounds calls *between* services only — the subgraph still receives 143 keys and must avoid its own N+1 while serving them.

code

graphql · 10 lines
graphql
query ExhibitionFloorPlan {
  exhibition(id: "EX-1913") {
    title
    artworks(first: 143) {
      title
      accessionNumber
      conservation { status lastInspected }
    }
  }
}

go deeper

for a junior

Recall the shape: a list resolved in one service whose fields live in another costs one call per item unless the keys are gathered up first. Be able to say why a network call per item is worse than a database call per item.

for a middle

Explain the mechanism concretely: the reserved _entities field takes a list of representations, the router builds one per item, and results come back in a list it splices back into position. Know that this bounds calls per plan step, not total work.

for a senior

Show where the bound leaks in production: the subgraph's own N+1 behind a single healthy-looking call, nesting that adds plan steps, and a batched call whose size grows without limit. Say which metric would have shown each.

for a principal

Own the framing that call count and call size are two separate amplification axes, and that a graph needs a policy for both — schema-level page caps, cost limits and subgraph-side defences — rather than trusting each team to notice its own fan-out.

## What "fetch amplification" names A federated graph is one client-facing schema composed from several subgraph schemas. The router accepts the client's document, works out which subgraph owns which field, and issues its own requests to those services. Fetch amplification is the ratio between what the client asked for — one document — and what the router had to do downstream. The dangerous case is when that ratio scales with the *data*: a list of 20 items costs 20 calls, a list of 143 costs 143. Take a museum collection graph. A **Catalog** subgraph owns `Artwork` with `@key(fields: "id")` and the `Exhibition.artworks` list. A separate **Conservation** subgraph contributes `Artwork.conservation`, holding condition reports the curators keep out of the catalogue database. A curator opens an exhibition floor plan: ```graphql query ExhibitionFloorPlan { exhibition(id: "EX-1913") { title artworks(first: 143) { title accessionNumber conservation { status lastInspected } } } } ``` Catalog answers `exhibition`, `title`, `accessionNumber`. Every `conservation` value lives in another process. The naive shape is one request per artwork: 143 HTTP calls for one screen. ## Why this is worse than an in-process N+1 The classic N+1 costs a database round trip per item. Here each item costs a full service call: connection acquisition, JSON serialization on both sides, an auth check, network latency, and a second execution of the whole GraphQL pipeline in the subgraph. At an unremarkable 6 ms per call, 143 sequential calls are the better part of a second of pure overhead; issued concurrently, they arrive at Conservation as a burst that can saturate its pool and slow every other caller. The multiplier is also invisible from the client side — the document looks small, and the amplification only shows up in the subgraph's request count. ## The batched entity fetch Apollo Federation's subgraph specification exists partly to make one call enough. Every federated subgraph exposes a reserved root field: ```graphql _entities(representations: [_Any!]!): [_Entity]! ``` The argument is a **list**. The router collects the key fields of all 143 artworks, builds one representation per item — each an object carrying `__typename` plus the key fields — and sends a single query: ```graphql query($representations: [_Any!]!) { _entities(representations: $representations) { ... on Artwork { conservation { status lastInspected } } } } ``` Conservation returns one list, the router splices each element back into the position it came from, and the client sees an ordinary nested response. One plan step, one call, whatever the list length. That list-shaped argument is the whole reason cross-service amplification is bounded at all — it is a composition-specification concept, not something the GraphQL specification itself defines. ## What batching bounds, and what it does not It bounds **calls between services** for one step of the plan. Three things it does not bound: *The work inside the subgraph.* Conservation now receives 143 keys in one request. If its `conservation` resolver loads one condition report per key, the N+1 simply moved across the wire — same 143 database round trips, now hidden inside a single HTTP call that looks healthy in the router's metrics. Cross-service batching and in-service batching are separate defences, and each only fixes its own layer. *The number of steps.* Nesting creates new steps. If each artwork's conservation record exposes a conservator who owns further artworks, the plan gains another entity fetch, and its representation count is the product of the fan-out above it. Deeply nested documents — the pathological ones run to 19 levels — turn a handful of steps into a long chain, each fed by a wider list than the last. *The size of the one call.* A batched fetch converts many small requests into a single very large one. That is usually the right trade, but it moves the risk from call count to payload size and blast radius. ## What an interviewer is listening for Count network calls per plan step, not rows. Say plainly that the fix is the list-valued `representations` argument rather than "the router is smart". And name the boundary: a batched entity fetch is a promise about the number of service calls, never a promise about the number of queries the receiving service runs.

  • Does a batched entity fetch mean the subgraph runs one database query?
    No. The router promises one *service* call; what the subgraph does with the 143 keys inside it is entirely the subgraph's problem. If its resolver loads one record per key, the N+1 has simply moved across the wire and is now hidden inside a request that looks healthy in the router's metrics. Cross-service batching and in-service batching are separate defences and you need both.
  • If entity fetches are batched, where does amplification come back?
    In nesting. Each nested field owned by another subgraph adds a plan step, and that step's representation count is roughly the product of the fan-out above it: 143 artworks, each with a conservator, each with further artworks. A pathological document nesting 19 levels turns two fetches into a long sequential chain, each fed a wider list than the last. Depth, not list length, is what the batching does not fix.
  • Is the list-valued `representations` argument part of the GraphQL specification?
    No. `_entities`, the `_Any` scalar and representations come from Apollo Federation's subgraph specification — the composition standard layered on top of GraphQL. The GraphQL specification itself knows nothing about entities, keys or routers; it only defines a schema with root fields, one of which here happens to be reserved by that convention.

A courier who drives to one street 143 times, once per parcel, versus loading all 143 into one van: the parcels are the same, the trips are not.

saying these in an interview costs you the question

  • Claims federation removes N+1 everywhere automatically
  • Thinks the router opens one request per list item
  • Counts rows returned instead of subgraph calls
  • Assumes a batched entity fetch means one database query
  • Believes a response cache in the router replaces batching
  • Cannot say where the keys for the batch come from

context