skip to content

Which federation schema do you diff for client-breaking changes before a subgraph deploys?

level: seniorimportance: should knowfreq 46%

answer

  1. Three documents, three different questions
  2. One of them is the consumer contract
  3. Composition failure is not a client incident
  4. A one-line addition can be a removal
  5. An empty diff still changes latency

basics

~20 s

Diff the API schema, because that is the only document consumers see. A subgraph SDL diff misses changes the other subgraphs decide, and a supergraph diff is full of routing churn that never reaches a client.

solid answer

~40 s

Run three checks in order. First, does the candidate subgraph still compose — composition fails closed, so a failure blocks the deploy without ever being a client incident. Second, recompose and diff the resulting **API schema** against the currently published one: removed fields, arguments made required, output fields made nullable, enum values dropped. That is the contract, and it is the only diff that catches a field removed by `@inaccessible` while the subgraph's own SDL grew by a line. Third, check the removals against recorded client operations or field-usage telemetry, because a removal only hurts if something selects it. The **supergraph** diff is a superset — every subgraph URL change and every ownership move shows up there — so it is the changelog you read to explain behaviour, not the gate you block on.

code

graphql · 11 lines
graphql
# before, in the registry subgraph
type Animal @key(fields: "id") {
  id: ID!
  herdBookNumber: String!
}

# after: additive in this file, a removal in the API schema
type Animal @key(fields: "id") {
  id: ID!
  herdBookNumber: String! @inaccessible
}

go deeper

for a junior

Know that clients only ever see the API schema, so that is the document a breaking-change check has to compare. Recognising that a subgraph's own file is not the contract is enough at this level.

for a middle

Explain why a subgraph diff is insufficient: the API schema is derived from all subgraphs together, so the same edit can be additive or a removal depending on what the others declare. Name the ordinary breaking-change classes you look for.

for a senior

Demonstrate the full gate — compose, diff the API schema, then check removals against recorded operation usage — and be able to separate a composition failure, which blocks deploys, from a shipped breaking change, which breaks a consumer.

for a principal

Own the policy: who is allowed to break the graph, what the deprecation window is relative to the slowest client, and whether the organisation collects the operation usage data without which the answer is always no.

## Three diffs, three different questions Every federated change produces three documents that could be diffed, and each diff answers a different question. Picking the wrong one is how a "purely additive" subgraph change reaches production and breaks a client. The pedigree graph's `Animal` is a 37-field type assembled from three subgraphs. Suppose the breeding team ships a change. What do you compare? ### The subgraph SDL diff answers: what did this team intend? Useful for review, useless as a safety gate, because a subgraph's SDL is only one input to the merge. Changes that look identical here have opposite consequences: * Removing a field that **another** subgraph also declares as shareable removes nothing from the API schema. * Removing a field that another subgraph `@requires` does not break clients at all — it breaks *composition*, and nothing ships. * Adding `@inaccessible` to an existing field is a one-line addition here and a **field removal** in the API schema. The subgraph diff cannot distinguish those, because the answer depends on documents the diff never looked at. ### The supergraph diff answers: what changed for the router? This is a strict superset of what changed for clients — every API-schema change shows up here, plus a great deal that never reaches a client. A subgraph URL moving, a field's ownership moving between subgraphs, a new subgraph declaring an existing value type, a composition run under a newer join version: all rewrite the supergraph and change nothing a client can observe. Gating on this diff means alerting on changes that are, by construction, invisible outside the platform. ### The API schema diff answers: what changed for consumers? This is the contract. Recompose with the candidate subgraph, derive the API schema, and diff it against the API schema currently published. The classes that matter are ordinary GraphQL breaking changes, not federation ones: * a type, field, enum value or argument removed; * an output field's type made nullable, or narrowed to something a client's existing selection cannot handle; * an argument or input field made non-null, or a new required argument added; * a member removed from a union, or an interface no longer implemented. Because the API schema is derived from the whole graph, this diff catches the `@inaccessible` case and the shareable-field case that the subgraph diff missed, and it stays quiet about routing churn that the supergraph diff would have shouted about. ## The check that comes before the diff Composition itself is the first gate, and it fails closed: if the candidate subgraph does not compose with its siblings, there is no new API schema to diff and nothing deploys. A composition failure is a broken build, not a client incident — a distinction worth making explicitly in an interview, because it separates "everyone's deploys are blocked" from "someone's app crashed". ## The check that comes after the diff A breaking change in the abstract is only an incident if something selects the removed element. With 37 fields on `Animal`, most of them have a handful of consumers and some have none. So the diff feeds a second question: over a window at least as long as the slowest client's refresh cycle, did any recorded operation select this field? That needs either a registry of the operations clients are allowed to send, or field-usage telemetry from the router. Without one of them, the only safe policy is never to remove anything, which is how a graph accumulates dead fields. Deprecation is the mechanism that makes the window work: mark, publish, watch usage fall, then remove. The deprecation reason is part of the API schema, so it is visible in exactly the document consumers introspect. ## The trap: identical API schema, different behaviour The most instructive failure on this leaf is the one where every check passes. Move a field's ownership from one subgraph to another and the API schema is byte-identical — same fields, same types, same nullability. But the query plan is not: the field now comes from a different service, possibly a fetch deeper in the plan, possibly one that must wait on a key fetch that used to be unnecessary. That is how you get a pedigree query that times out only in production. Staging has a few thousand animals and every fetch is fast; production has the real herd book, the new owning subgraph's fetch sits behind another one, and a document that used to complete comfortably now exceeds the client's timeout. The schema diff was empty. The supergraph diff — the one you were tempted to ignore as noise — was the only artefact that showed the change at all. So the honest answer to "which schema do you diff" has two halves: **the API schema is the contract, and the supergraph is the changelog.** Diff the first to decide whether consumers are affected; read the second to explain a latency regression that the first swears did not happen.

  • Give an example of a change that is additive in a subgraph but breaking in the API schema.
    Adding `@inaccessible` to a field that is already published. The subgraph SDL gains a directive and loses nothing, composition succeeds, and the supergraph still contains the field so the router can use it internally. But deriving the API schema removes it, so every client document selecting it now fails validation. Only the API-schema diff sees this.
  • The API schema diff is empty after a deploy, yet one pedigree query now times out in production. What changed?
    Something that lives only in the supergraph — most often field ownership moving to another subgraph. The field's name, type and nullability are unchanged, so the contract diff is silent, but the query plan now fetches it from a different service, possibly a step deeper and behind a key fetch that used to be unnecessary. The supergraph diff is where that change is visible.
  • Why is a composition failure treated differently from a breaking API change?
    Because it fails closed. No new supergraph is produced, the router keeps serving the last good one, and no client sees anything. The cost is that everyone's deploys through that graph are blocked until it is fixed, so it is an urgent build problem rather than an outage. A breaking API change, by contrast, ships successfully and breaks somebody else's application.
  • How long should a field stay deprecated before removal?
    At least as long as the slowest consumer takes to roll forward, which for a shipped mobile client is months rather than days. The deprecation reason is part of the API schema, so it is visible in the document consumers introspect. The decision to remove should be evidence-based: zero recorded operations selecting the field across a window longer than that refresh cycle.

saying these in an interview costs you the question

  • Diffs only the changed subgraph SDL and calls it safe
  • Treats every supergraph diff as client-facing
  • Thinks successful composition proves nothing broke
  • Forgets @inaccessible removes a field from clients
  • Assumes an unchanged API schema means unchanged behaviour
  • Removes a field without checking operation usage

context