When would you keep a working stitched GraphQL gateway instead of moving to composition?
answer
- Ask whose queue the change enters
- Say the invoice out loud
- Some schemas are not yours to change
- Transforms have nowhere to live
- Trigger on lead time, not headcount
basics
~20 sKeep it when the mapping is edited rarely by the people who own the services, when some merged schemas are not yours to change, or when you rely on gateway transforms composition cannot express. Move when it becomes a cross-team queue.
solid answer
~50 sThe decision is about **who edits the join**, not about which technique is newer. Composition buys team autonomy — the owning team declares its contribution and ships on its own cadence — and charges for it with a real platform: a composition step in the build, somewhere to publish and version schemas, a router to operate and carry a pager for, and a new class of failure where two teams' declarations cannot merge and nobody's release goes out. With three or four contributing teams, that platform costs more than the queue it removes. Two hard blockers keep systems stitched regardless of size: schemas you do not control cannot declare their own participation, and gateway-side transforms — renames, hidden fields, synthesised shapes — have no specified equivalent. The signal to move is lead time, not headcount: when a cross-service field change waits on one team's review, the configuration is the bottleneck.
code
pseudocode · 8 lines# measure these before proposing a migration
count(delegation_rules) # 41 now, 12 a year ago
count(rules_no_current_engineer_can_explain)
count(teams_contributing) / count(teams_that_may_merge_gateway_config)
median_days(cross_service_field_request -> in production)
days_between(remote_schema_change, gateway_snapshot_refresh)
# migrate when the join is a queue, not when the technique is oldergo deeper
Understand that both approaches look identical to a client, so the choice is about internal ownership. Knowing that composition needs a build step and a router of its own is enough at this level.
Be ready to list what composition costs as well as what it buys, and to say why a small number of contributing teams often does not clear that bar. Avoid framing stitching as simply outdated.
Show the blockers that override team size — schemas you cannot change, and gateway transforms with no composed equivalent — and describe the staleness and drift failures a stitched gateway actually produces in operation.
Own the decision with a measurable trigger rather than a preference: lead time and approver concentration for cross-boundary change, rules nobody owns, and a costed inventory of the transforms before anyone commits a quarter to the move.
## Decide on the queue, not on the diagram Both models present clients with one endpoint and one schema. Nothing a client does distinguishes them. So the entire decision lives on the inside, and one question captures it: **when a client needs a field that crosses a service boundary, whose queue does that work enter?** Stitched: the gateway team's, always, no matter which team owns the data. Composed: the owning team's, and only theirs. That is the trade being purchased, and it is worth buying exactly when the queue is long enough to hurt. ## What composition actually costs Candidates who have only read about federation describe the benefits and skip the invoice. Say the invoice out loud: - **A composition step in the build**, which can fail — and when it fails, it fails closed. Two teams' declarations that cannot be merged mean no new graph is published, so an unrelated team's release is blocked by a schema conflict it did not cause. - **Schema publishing and versioning** — somewhere services register what they contribute, with history, checks and a promotion path per environment. - **A router to operate**: capacity, latency budget, deploys, and a pager. It is now the single component every client request passes through. - **Real changes in every backing service**: identifying fields declared, resolution by those values implemented, ownership of every shared field decided explicitly rather than left to a gateway rule. - **Ownership arguments made explicit.** Stitching lets a gateway rule quietly decide that the fares service supplies `Seat.price`. Composition makes some team declare it, and that conversation is sometimes months long. With nine contributing teams that invoice is obviously worth paying. With three teams in one department who ship together anyway, it usually is not, and saying so is a stronger answer than reciting the migration plan. ## The two blockers that are not about size First, **you cannot compose a schema you do not control.** Composition merges what each service declares, so a third-party schema, a vendor endpoint, or a frozen legacy service that nobody may redeploy cannot join a composed graph on its own. You can front it with a thin service you own that declares participation on its behalf — but notice you have just written the delegation rule again, in code, and you now operate the wrapper too. Second, **gateway-side transforms have no specified equivalent.** If your merged schema exists partly because the gateway renames types, hides fields and synthesises wrapper shapes over an old backend, composition offers no central place to do that. Every transform must become someone's real schema or someone's real service. That is often the true cost of the migration, and it is invisible on an architecture diagram. ## The signals that it is time Measure rather than assert. The useful indicators are all about the configuration behaving like a shared resource: - **Lead time for a cross-service field**, and how many people can approve it. If the median is measured in weeks and the approver set is one team, the gateway is the bottleneck. - **Rules nobody can explain.** In one seat-map graph, a third of the 41 delegation rules predated everyone on the platform team; that is drift with no owner. - **Staleness incidents.** A stitched gateway that snapshots remote schemas at start-up carries an ageing copy of every service's type system. When the seat-inventory team added a crew-hold value to a seat-status enum, the gateway kept serving its stale copy for eleven days, and the new value surfaced as an unattributable field error — the gateway's view of that enum did not contain it. Nobody owned the refresh because the refresh belonged to a configuration, not to a team. - **Growth rate of the configuration versus growth of the graph.** If rules grow faster than types, the central file has become the design. None of those is fixed by choosing a newer technique — they are fixed by moving the declarations to the people who own the data, which is what composition is for. ## What a strong answer sounds like Name the axis (who edits the join), price both sides honestly, name the two blockers that override team count, and then commit to a measurable trigger rather than a preference. Something like: "Below roughly a handful of contributing teams with no autonomy pain, I keep the stitched gateway and spend the effort on schema checks and a refresh policy instead. I move when cross-service lead time is dominated by waiting on one team, or when the configuration has rules nobody owns. And I would not attempt it at all until I know what the transforms in that configuration are going to cost, because that is the part that never appears on the plan." The weakest answer, and it is common, is that composition is the modern approach and stitching is legacy. That skips the invoice, ignores the schemas you do not control, and mistakes a technology preference for an organisational decision.
- What does the composed path cost that a stitched gateway does not?A composition step that fails closed and can block an uninvolved team's release, a place to publish and version service schemas, a router to operate and be paged for, and real changes inside every backing service to declare identity and ownership. Those costs are fixed; the benefit scales with the number of contributing teams.
- Can a schema you do not control join a composed graph?Not on its own. Composition merges what each service declares, so a vendor endpoint or a frozen legacy service that cannot be redeployed has to be fronted by a thin service you own which declares participation on its behalf. That wrapper is a delegation rule written in code, plus another service to operate.
- If you had one metric to decide, what would it be?The lead time for a field that crosses a service boundary, together with how many people may approve that change. Long lead times concentrated on a single approving team mean the central configuration is a queue, and that is precisely the cost composition removes. Team count alone predicts it poorly.
- What would you do for a team that must stay stitched for now?Spend the effort on the failure modes rather than on the topology: a scheduled or triggered refresh of the remote schema snapshots with an owner, checks that fail the build when a service change would alter the merged schema, and a periodic audit deleting delegation rules nobody can explain. Most stitching incidents are staleness and drift, not stitching itself.
Central configuration is a shared filing cabinet: fine while four people share a room, unbearable once nine teams need something out of it every week.
saying these in an interview costs you the question
- Migrates because the technique is newer
- Ignores the operational cost of running a router
- Assumes third-party schemas can join a composed graph
- Counts services instead of counting owning teams
- Forgets gateway transforms have nowhere to go
- Treats an organisational choice as purely technical