skip to content

How would you migrate a stitched GraphQL gateway to a composed graph without a flag day?

level: seniorimportance: should knowfreq 46%

answer

  1. One thing must not change
  2. Move topology first, contract later
  3. The old gateway can be one node
  4. Renames have to become real names
  5. Replay traffic and diff, do not inspect

basics

~20 s

Treat the client-facing schema as the invariant and move one type at a time: register the remaining stitched gateway as a single subgraph of the composed graph, then lift types out of it, diffing the schema and replaying real traffic.

solid answer

~50 s

The migration has one invariant and one hard part. The **invariant** is the schema clients see: it must stay compatible at every step, so you diff it after each move and treat any change as a deliberate client change rather than a side effect. The **hard part** is not plumbing, it is that a stitching configuration contains things composition has no home for — renames, hidden fields, argument rewriting, wrapper types the gateway synthesised. Each must move into the owning service's real schema, or into a thin service that presents the transformed view. The incremental route: register the existing stitched gateway as one subgraph of the composed graph — it is itself a GraphQL server — then lift one type at a time out of it. Verify by replaying recorded operations against both paths and diffing data and error paths; schema equality will not catch an argument mapping that differed.

code

graphql · 12 lines
graphql
# Stitched today: clients see FareRow because the gateway renamed the
# fares service's colliding "Seat" type. Composition has no such rename.
#
# Migrated: the fares team renames the type in its own schema to the name
# clients already see, so the client-facing schema does not change at all.

type FareRow @key(fields: "flightNumber seatNumber") {
  flightNumber: String!
  seatNumber: String!
  amountMinorUnits: Int!
  currency: String!
}

go deeper

for a junior

Know the destination: services stop being merged by a central configuration and start declaring their own contribution. The thing that must not change during the move is the schema clients already depend on.

for a middle

Explain the incremental route rather than a cutover — the stitched gateway becomes one node of the new graph and types are lifted out of it individually, with the merged schema diffed after each move.

for a senior

Show that you have costed the parts that do not translate: renames, hidden fields and synthesised shapes, plus services that lack a stable identifier. Then describe verification by replaying real traffic and diffing data and error paths, not by reading schemas.

for a principal

Own the sequencing argument: topology and contract must not change in the same step, or no diff is interpretable. Be ready to say how long you will carry two paths, who pays for the adapter work, and what ends the migration.

## Frame it as a schema-preserving refactor Everything downstream of the gateway is changing; everything upstream must not notice. That framing gives you the plan, the verification strategy and the rollback story in one move. Keep the airline seat-map graph in view. Nine teams contribute; the gateway configuration has grown to 41 delegation rules and 6 renames, and roughly a third of the rules predate everyone currently on the platform team. ## Step 1: make the client-facing schema a tracked artefact Before moving anything, capture the merged schema the gateway serves today and put it under diff in continuous integration. Every subsequent step is judged against it. This is also the moment you discover how much of that schema exists only because of gateway configuration — the renamed types, the fields filtered out, the wrapper types nobody's service actually declares. ## Step 2: inventory what the configuration does that composition cannot express Work through the configuration and classify every rule: - **Plain delegation** — resolve this field by calling that service with the parent's key. This translates directly: the owning service declares the key and contributes the field. - **Renames** — the gateway calls the fares service's `Seat` something else because it collides. Composition defines no gateway-side rename; one type name is one type. The fix is that the fares team renames the type **in its own schema** to whatever the clients already see, so the gateway's rename becomes the real name and clients notice nothing. - **Hidden or filtered fields** — the gateway removed a field from the merged schema. Composition merges what is declared, so the field must be removed from the service's schema, or the service must keep an internal-only variant. If it was hidden for authorisation reasons, that is a genuine gap to close, not a translation. - **Argument rewriting and synthesised types** — the gateway invented a shape that no service declares. These are the expensive ones: someone has to own the shape, which means either a service adopts it or you keep a thin service whose only job is to present it. That inventory, not the router deployment, is the real project plan. ## Step 3: run both models at once, with the old gateway as one subgraph The key move that avoids a flag day: a stitched gateway is itself an ordinary GraphQL server, so it can be registered as a **single subgraph** of the composed graph. Everything not yet migrated keeps working exactly as before, behind one node in the new topology. It has to expose the composition entry points for any type another service will key into, which is a small adapter, and it will not participate in query planning any more cleverly than a black box — but that is fine, because its role is temporary and shrinking. Now migration becomes one repeatable move: pick a type, have its real owning service declare it properly, remove it from the stitched remainder, recompose, diff the client-facing schema, ship. Nine teams can do this on nine independent calendars. ## Step 4: make services federation-capable, one at a time Each service joining directly has to do real work: declare the fields that identify its objects, implement resolution by those identifying values, and be honest about which fields it actually owns versus borrows. Two problems surface here reliably. First, a service that never had a stable identifier for an object now needs one — a seat identified by a database row id that changes on reseat is not a usable identity. Second, ownership disputes that stitching papered over become explicit, because someone must declare the field. ## Step 5: verify by replay, not by inspection Schema equality is necessary and nowhere near sufficient. A delegation rule whose argument mapping quietly differed from the new declaration produces a schema that matches perfectly and data that does not. Replay a captured corpus of real operations against both paths and diff the full response — `data` **and** the error paths. Error paths in particular expose mismatched nullability and differently-shaped failures that data comparison on the happy path never touches. Do the replay with real recorded traffic rather than hand-written cases. The operations that break are almost always the ones nobody remembered were in production. ## Step 6: keep rollback trivial until the end Until the last type leaves the stitched remainder, the old gateway is still running and still correct. Route a percentage of traffic, keep the ability to route it all back, and only decommission when the remainder subgraph is empty and the schema diff has been clean for a while. ## The failure mode to name out loud The migration that goes wrong is the one where the composed graph is stood up in parallel as a *new* graph with a *better* schema. Two changes at once — new topology and new contract — and now every client difference is ambiguous: is it a migration defect or an intended improvement? Move the topology first with the contract frozen, then improve the contract afterwards, when a schema diff means exactly one thing.

  • What is the invariant that tells you the migration is on track?
    The schema clients see. Freeze it as a baseline, diff after every move, and treat any difference as a client-facing change that needs a decision rather than an artefact of the migration. If the contract and the topology change together, no diff can be interpreted.
  • How do services move one at a time rather than all at once?
    Register the existing stitched gateway as a single subgraph of the composed graph — it is itself a GraphQL server, and it needs a small adapter exposing the composition entry points for the types others key into. Everything unmigrated keeps working behind that one node, and types are lifted out of it individually until it is empty.
  • What in a stitching configuration has no equivalent on the composed path?
    Gateway-side transforms: renames, filtered or hidden fields, rewritten arguments, and wrapper types the gateway synthesised. Composition merges what services declare and specifies no central transform, so each one must be pushed into the owning service's schema or absorbed by a thin service you keep specifically to present that shape.
  • Why is comparing the two schemas not enough verification?
    A delegation rule that mapped a parent value onto a different argument than the new declaration does produces an identical schema and different data. Replay recorded production operations against both paths and diff the full response, including error paths, which is where nullability and failure-shape differences show up first.

saying these in an interview costs you the question

  • Plans a single big-bang cutover of every service
  • Assumes gateway renames carry over automatically
  • Changes the client contract during the migration
  • Verifies with schema comparison only
  • Expects the backing services to need no changes
  • Decommissions the old gateway before the remainder is empty

context