skip to content

When does splitting a GraphQL graph across services pay off, and when does one server still win?

level: seniorimportance: must knowfreq 59%

answer

  1. Ask who is waiting for whom
  2. Autonomy, not throughput
  3. Count contributing teams, not services
  4. Implicit couplings become explicit or broken
  5. Measure schema lead time first

basics

~20 s

Team autonomy is the payoff: each team ships its slice of the schema without a shared release train. It is not a performance win — composition adds hops and failure modes. With few contributors on one cadence, one server still wins.

solid answer

~40 s

The driver is organisational, not technical. Splitting a graph converts a coordination problem — every schema change queueing behind one repository and one deploy — into an operational one: a composition step in CI, somewhere to publish and version schemas, a router to run, distributed tracing to debug with, and a class of build failure where two teams' declarations will not merge. That trade is worth it once enough teams contribute that the queue is the bottleneck. It is a bad trade when one small team owns the schema and ships daily, when the graph is essentially one bounded context, or when the product is still being discovered. Two anti-reasons come up constantly: splitting to reduce latency, which usually worsens it, and splitting because the backends are already microservices, which is a separate decision.

code

pseudocode · 9 lines
pseudocode
# both fields are on Vehicle, in one aggregating server
resolve Vehicle.lastTrip(vehicle, ctx):
    trip = ctx.tripStore.mostRecent(vehicle.vin)
    ctx.perRequest["odometerKm"] = trip.endOdometerKm   # side effect nobody declared
    return trip

resolve Vehicle.serviceDueIn(vehicle, ctx):
    odometer = ctx.perRequest["odometerKm"]             # assumes lastTrip already ran
    return vehicle.serviceIntervalKm - (odometer % vehicle.serviceIntervalKm)

go deeper

for a junior

Know that one GraphQL server calling several backends is a normal, respectable design, and that splitting the graph across services is something teams do for organisational reasons rather than for speed.

for a middle

Explain the trade in both directions: what a shared schema repository costs in coordination, and what a composed graph costs in platform work, extra hops and new failure modes. Be able to say why latency is not the justification.

for a senior

Demonstrate you have measured it. Talk about schema lead time and blocked reviews as the signal, name the implicit in-process couplings a split exposes, and describe what debugging looks like afterwards.

for a principal

Own the sequencing and the reversibility. Decide what capacity has to exist before a split is safe, which halfway steps buy autonomy cheaply, and what evidence would make you stop the migration rather than finish it.

## The answer people give, and the answer that holds up Asked why they split a graph, teams say scalability, or decoupling, or that their backends are microservices. None of those survives contact with the actual bill. The honest answer is narrower and more useful: **splitting the graph buys independent change**. Each team edits and deploys its own slice of the schema without waiting on anyone else's release. Everything else about the arrangement gets harder. Notice what that implies. The signal to split is a property of the *organisation*, not of traffic. Nine services behind one graph owned by one small team is a perfectly healthy shape. One service behind a graph that thirteen teams contribute to is not. ## The bill A composed graph is a platform, and someone has to run it: * a composition step wired into every contributing team's CI, so a breaking declaration is caught before publish; * somewhere to publish and version schemas, and a policy for what happens when composition fails; * a router in the request path — one more hop, one more thing to size, one more thing that pages someone; * distributed tracing, because "why was this field slow" stops being answerable from one process's profile; * governance: naming, deprecation and an arbitration path for two teams that want the same type name. In a single aggregating server none of that exists. There is one deployable, one stack trace, one place to put a breakpoint, and the join between two backends is a function call you can read. ## What you give up, concretely Here is a failure worth carrying into an interview, because it is the shape of the loss rather than a detail of it. A fleet telematics platform ran one aggregating server. Its `Vehicle.serviceDueIn` resolver read an odometer reading that the `Vehicle.lastTrip` resolver had already put on the per-request context. That worked for two years. It was never *guaranteed* to work: GraphQL does not promise that sibling fields of an object resolve in any particular order, and the coupling was invisible — it lived in shared mutable per-request state, not in any declaration. When the graph was split, trips and maintenance became separate services and were fetched independently. The odometer was simply not there. `serviceDueIn` came back wrong for 4 of the platform's 19 vehicle classes and nobody noticed for 17 days, because a plausible-looking number is not an error. The general lesson: a single server gives you total control over ordering, sharing and in-process caching, and teams quietly build on that control without writing it down. Splitting the graph withdraws it. Every implicit coupling either becomes an explicit declaration or becomes a bug. ## The signal to split, measured Do not argue this from architecture diagrams. Measure the queue: * lead time from a schema pull request being opened to the field being live; * how many of those pull requests are blocked on a reviewer from another team; * merge-conflict rate on the schema files; * how often rolling back one team's change reverts another's, because they shipped together. If a schema change reaches production in a day, autonomy is not your bottleneck and composition will make you slower. If it takes three weeks and half of that is waiting for a shared release window, you have found the thing splitting actually fixes. ## When one plain schema still wins * **A small number of contributors on one cadence.** A 4-person platform team owning the whole schema is faster with one server, full stop. * **One bounded context.** If the graph is mostly one domain, the boundaries you would draw are arbitrary and the composition machinery buys nothing. * **An unsettled product.** Early on the schema is reshaped weekly; cross-team declarations calcify exactly the decisions you most want to keep cheap. * **No platform capacity.** A composed graph without an owner degrades into a graph nobody can change safely, which is worse than the release train you replaced. ## The two anti-reasons **Performance.** Splitting adds a router hop and turns in-process joins into network calls. It can improve *isolation* — one slow owner no longer blocks a thread pool shared with everyone — but as a latency play it goes the wrong way, and a candidate who claims otherwise is guessing. **"Our backends are already microservices."** The shape of the graph layer is a separate decision from the shape of the services beneath it. One aggregating server over thirty microservices is a completely legitimate design; whether the *graph* splits depends on how many teams need to change the *schema*, not on how many processes hold the data. ## Middle grounds worth naming Before committing, there is room in between: enforced code ownership over sections of one server's schema, one server per client audience, or composing just two willing teams as a pilot while everyone else stays in the aggregating server. Each gets you part of the autonomy for a fraction of the platform cost, and each is reversible.

  • Our backends are already thirty microservices. Does that settle whether the graph should be split?
    No. The graph layer's shape is an independent decision. One aggregating server over thirty services is entirely legitimate, and it is the right answer whenever a single team owns the schema and can ship it quickly. The question that settles it is how many teams need to change the schema itself, and how long they currently wait to do so.
  • How would you demonstrate that a shared schema release train is genuinely costing you?
    With numbers rather than anecdotes: lead time from schema pull request to live field, the share of those pull requests blocked on another team's reviewer, merge-conflict rate on the schema files, and how often rolling back one team's change reverts another's. If a field ships in a day, autonomy is not the bottleneck and splitting will make things slower.
  • What gets harder about debugging once the graph is split?
    A single process's stack trace and profile stop answering the question. Diagnosing a slow or failing field needs distributed tracing correlated across the router and every owner it consulted, plus a way to look up which team owns a field. Budget that tooling as part of the split rather than discovering it during the first incident.
  • Is there a halfway step before committing to a split graph?
    Several, and all are reversible. Enforce code ownership over sections of the one server's schema so teams review only their own. Run one server per client audience. Or compose just two willing teams as a pilot while everyone else stays in the aggregating server, and compare their schema lead time against the rest after a quarter.

Splitting the graph is like giving each team its own printing press. Nothing prints faster, and you now own several presses — but nobody waits in the queue.

saying these in an interview costs you the question

  • Says splitting the graph makes queries faster
  • Splits the graph to fix backend N+1 calls
  • Assumes microservice backends force a split graph
  • Ignores the platform team a router and registry need
  • Claims composition removes cross-team coordination entirely
  • Says one aggregating server cannot serve many backends

context