skip to content

How does a supergraph get composed and how does the router plan a query across subgraphs, and where can it go wrong at scale?

level: principalimportance: should knowfreq 28%

answer

  1. compose = build-time merge + validate SDLs
  2. router builds a query plan of fetch nodes
  3. _entities joins by @key at runtime
  4. N+1 -> batch @EntityMapping/@BatchMapping
  5. @requires chains serialize -> latency

basics

~20 s

Composition merges every subgraph's SDL into one supergraph schema at build time, failing if they're incompatible. At runtime the router builds a query plan — which subgraphs to call, in what order, using _entities to join by @key — then executes and merges. Risks: N+1 entity fetches, deep dependency chains, and composition conflicts.

solid answer

~50 s

Composition is a build-time step (Apollo `rover supergraph compose` or GraphOS) that reads each subgraph's `_service` SDL and merges them into a single supergraph schema, validating that shared entities agree on keys and that types are compatible; incompatibilities fail the build before deploy. At runtime the router parses a client query against the supergraph and produces a query plan: a tree of fetches assigning each field to its owning subgraph, sequencing them by dependency (e.g. get `Book` from catalog, then batch its ids into the reviews subgraph's `_entities`). It executes fetches, calling `_entities` to resolve references by `@key`, and stitches the results. At scale the pitfalls are N+1 across subgraph boundaries (mitigate with batch `@EntityMapping`/`@BatchMapping`), long sequential `@requires` chains that raise latency, over-broad entity ownership, and composition conflicts when teams evolve shared types. Spring subgraphs address the data side via batching; the router owns planning.

go deeper

for a junior

Know composition happens before runtime and the router merges results; details not expected.

for a middle

Describe query planning at a high level and that _entities joins by key.

for a senior

Explain batch entity resolution to avoid N+1 and the compose-vs-execute split.

for a principal

Reason about subgraph boundary design, @requires latency, composition governance, and observability of query plans at scale.

## Two phases: compose (build) and execute (runtime) Federation cleanly separates a build-time and a run-time concern. ### Composition (build time) - Each subgraph exposes its schema via `_service { sdl }`. - A composition tool — Apollo `rover supergraph compose` locally, or **Apollo GraphOS** as a managed service — fetches all subgraph SDLs and merges them into one **supergraph schema** (which also embeds routing metadata: which subgraph owns which field). - **Validation** happens here: shared `@key` entities must have compatible keys and field types; a type owned by two subgraphs with conflicting definitions fails composition. This is your safety net — breaking changes are caught before they reach the router. - The resulting supergraph is what the router loads. With GraphOS you also get **schema checks** in CI: proposed subgraph changes are validated against the current supergraph and against real traffic before publish. ### Query planning (run time) - The router holds the supergraph. For each incoming client operation it builds a **query plan**: a tree of `Fetch` nodes, each targeting one subgraph, ordered by data dependencies. - Typical shape: a root fetch to the owning subgraph, then **`_entities` fetches** to other subgraphs, joining by the entity `@key`. Sibling fetches with no dependency can run in parallel; dependent ones (e.g. a `@requires` field) run in sequence. - The router executes the plan, calls each subgraph (including `_entities` with representation arrays), and merges partial results into the client response shape. ## Where it breaks at scale 1. **N+1 across the boundary.** If a query returns 100 `Book`s and each needs `reviews` from another subgraph, a naive plan sends 100 single-entity lookups. The router already batches representations into one `_entities` call, but the *subgraph* must resolve that batch efficiently — otherwise you get N+1 inside the subgraph. Mitigate with a batch-form `@EntityMapping` (List in, List out, order preserved) and `@BatchMapping` for contributed collection fields backed by DataLoader-style batching. 2. **Deep `@requires` chains.** Each `@requires` adds a sequential dependency edge: the router must fetch owner data before it can call the computing subgraph. Chained requirements serialize fetches and inflate latency. Prefer keeping computed fields near their data. 3. **Composition conflicts / coupling.** Shared entities create cross-team coupling: renaming a key field or changing a shared type can break composition. Governance (schema checks, ownership rules, deprecation windows) is essential. 4. **Over-fetching in plans.** Poor entity boundaries force the router to hop between subgraphs repeatedly for one logical object; each hop is a network call. Design subgraph boundaries around cohesive ownership. 5. **Error and partial-failure semantics.** If one subgraph fails, the router returns partial data + errors; clients and resolvers must handle nulls gracefully. `@EntityMapping` returning `null` for unresolved keys is normal and must not throw. ## Spring's role vs the router's - **Spring subgraph** controls the *data* side: efficient `@EntityMapping`/`@BatchMapping`, correct `@key`/`@requires`/`@provides` in SDL, and returning entities in input order. - **Router** controls *planning and orchestration*. You do not write query-plan logic in Spring; you make each subgraph fast and correct, and you shape entities/keys so the planner produces cheap plans. ## Design guidance (principal lens) - Draw subgraph boundaries around ownership and cohesion, not convenience, to minimize cross-subgraph hops. - Keep keys stable and minimal; treat shared-entity changes as public API changes with checks and deprecation. - Use GraphOS/rover schema checks in CI to catch composition breaks pre-merge. - Instrument query-plan depth and per-subgraph latency; watch for `_entities` fan-out hotspots and add batching there first. - Reserve `@provides` for measured hotspots; accept its consistency risk consciously.

  • How do you keep _entities resolution from becoming N+1 in a Spring subgraph?
    Use the batch form of @EntityMapping (accept a List of keys/representations, return a List/Flux in matching order) and back contributed fields with @BatchMapping / DataLoader so one request triggers one bulk query, not one per entity.
  • At what point are incompatible schema changes across subgraphs detected?
    At composition (build time). Merging subgraph SDLs validates key and type compatibility; incompatible changes fail composition before the supergraph is published, and GraphOS schema checks catch them in CI.
  • Why can heavy use of @requires hurt latency?
    Each @requires forces the router to fetch owner-owned fields first and pass them in, adding a sequential dependency edge to the query plan. Chained requirements serialize fetches instead of running them in parallel.

context