skip to content

How do you decide whether a company runs one GraphQL graph or several?

level: principalimportance: should knowfreq 43%

answer

  1. Ask what the merge makes possible
  2. Do client operations cross the boundary?
  3. Name the team that operates it
  4. Audiences differ; so should graphs
  5. Merging later is cheaper than splitting later

basics

~20 s

Decide on audience and traversal, not tidiness. One graph earns its cost only where real client operations cross domains. Separate graphs are right when audiences, cadences or trust boundaries differ, or nobody can staff the platform a shared graph needs.

solid answer

~40 s

Four tests, in order. **Traversal**: if no client operation ever selects from domain A into domain B, joining them buys coupling and no capability — two graphs are cheaper and equally useful. **Staffing**: a shared graph needs a standing owner for composition, schema publication, the router and conflict arbitration; a 4-person platform team can carry that for a dozen contributing teams only if contribution is genuinely self-service. **Blast radius**: one graph means one composition that can fail for everyone and one router whose incidents are everyone's incidents. **Audience**: an internal graph and a partner-facing graph have different security postures, evolution rates and deprecation promises, and forcing them into one type system means the slowest promise governs the fastest team. Most organisations land on a small number of graphs aligned to audiences, each composed internally.

go deeper

for a junior

You will not make this call, but know that a company can run more than one GraphQL graph and that 'the graph' is not always singular. Which endpoint your client points at is a deliberate decision somebody made.

for a middle

Be able to say what a shared graph costs beyond the code: a composition step, a place to publish schemas, a router to run, and a naming policy. Recognising that these need an owner is what is expected at this level.

for a senior

Argue the traversal test with evidence — recorded client operations that cross a boundary — and describe the failure isolation you give up when two domains share one composition and one router.

for a principal

Own the whole frame: traversal, staffing, blast radius and audience, plus the reversibility asymmetry that makes merging later cheaper than splitting later. Name the metrics you would hold the decision to a year on.

## "One graph" is a slogan before it is a decision The pitch is genuinely attractive: every client, one endpoint, every relationship traversable. It is also the kind of goal that survives on aesthetics long after the evidence stops supporting it. A principal engineer's job here is to convert it into tests that can come back negative. ## Test one: traversal The only capability a combined graph adds over two separate ones is the ability to select **across** the boundary in a single operation. So ask for the evidence: are there real client operations that would traverse from the fleet telematics domain into, say, the billing domain? Not "could imagine", not "someday" — recorded operations, or a screen someone is actually building. If the answer is no, joining the two graphs adds a shared namespace, a shared composition failure mode and a shared release surface, in exchange for nothing a client can use. Two graphs, cleanly separated, are the cheaper and more honest answer. Domains that never traverse are the strongest signal for separate graphs there is, and it is the test teams most often skip because the answer is embarrassing. ## Test two: who owns it on a Tuesday afternoon A shared graph is a product with an operator. Concretely, someone must own the composition step, the place schemas are published and versioned, the router in the request path, the breaking-change checks that run against recorded operations, the naming and deprecation policy, and the arbitration when two teams want the same type name. Say the candidate arrangement is 13 contributing teams and a 4-person platform team. That is workable — but only if contribution is self-service: composition checks run inside each team's own CI, publishing is automated, and the platform team is consulted for conflicts rather than for changes. The moment the platform team becomes a review gate, the shared graph has reproduced the release train it was built to remove, with more moving parts. If you cannot staff that team at all, one graph is not a plan; it is a promise you will be unable to keep. ## Test three: blast radius One graph is one failure domain in more ways than teams expect. Composition fails closed, which is the correct behaviour and also means one team's bad publish blocks everyone else's until it is resolved. The router is a single component every client depends on; its capacity, its rollout and its incidents are shared. A shared graph makes every team's worst day slightly more likely and considerably more crowded. Separate graphs give you failure isolation for free, which for a domain with a different availability target — a partner-facing graph with contractual uptime, say — may matter more than any traversal you would gain. ## Test four: audience and trust Internal and external consumers are not the same consumer. A partner-facing graph is a curated subset with explicit deprecation guarantees, a slower evolution rate, tighter authorization and a much smaller surface. An internal graph changes weekly and expects callers to keep up. Fusing them into one type system means the partner promise governs the internal pace, and the internal surface has to be actively hidden rather than simply not existing. The same logic applies to trust boundaries generally: if two parts of the business cannot share a namespace without a security review of every field, they probably should not share a graph. ## Reversibility, and which mistake is cheaper Ask which error you would rather make. Merging two graphs later is mostly a naming exercise plus a composition run — unpleasant but bounded. Splitting one graph that clients have been traversing for two years means breaking their operations, because a document that crosses the boundary simply cannot be expressed any more. That asymmetry argues for starting with more graphs than the slogan wants and merging on evidence. ## What to measure once you have chosen Whoever owns the decision should hold numbers on it: schema lead time per contributing team, how often composition fails and how long it stays failed, the share of recorded client operations that actually traverse a domain boundary, and how much of the platform team's week goes to arbitration rather than to the platform. If the traversal share stays near zero after a year, the graph was merged for tidiness and can be split back before more clients depend on it. ## The usual landing place Organisations that think this through rarely end at one graph or at one-per-team. They end with a small number of graphs aligned to **audiences** — an internal graph, a partner graph, sometimes one for a distinct product line — each composed internally from the services that genuinely traverse each other. That shape gets most of the traversal benefit, keeps failure domains and evolution rates separable, and is small enough for a platform team that actually exists to operate.

  • What is the cheapest experiment before committing a company to one shared graph?
    Compose two willing teams whose data genuinely traverses, leave everyone else in place, and run it for a quarter. The platform cost shows up in week one and is easy to measure; the autonomy benefit takes a couple of release cycles to appear. Compare those two teams' schema lead time and incident count against the rest before extending the arrangement.
  • How do you stop a shared graph from turning the platform team into the new bottleneck?
    Make contribution self-service. Composition and breaking-change checks run inside each contributing team's own CI against recorded operations, publishing is automated on merge, and ownership of every field is discoverable without asking. The platform team's queue should contain conflicts and policy questions only. If it contains routine schema changes, the release train has been recreated.
  • When would you deliberately run separate internal and partner-facing graphs?
    Whenever the two audiences have different promises. A partner graph is a curated subset with contractual deprecation guarantees, tighter authorization and a slow evolution rate; the internal graph changes weekly. Merging them makes the partner promise govern internal pace and turns internal surface into something you must actively hide rather than simply not publish.
  • Which is the more expensive mistake: too many graphs or too few?
    Too few. Merging two graphs later is largely a naming reconciliation plus a composition run. Splitting one that clients have traversed for years breaks their operations outright, because a document crossing the boundary can no longer be expressed. That asymmetry argues for starting with more graphs and merging only once real traversal appears.

Deciding how many graphs to run is closer to deciding how many front doors a campus needs than to deciding how many buildings it has.

saying these in an interview costs you the question

  • Treats one company-wide graph as an automatic goal
  • Cannot name who operates the router and registry
  • Merges domains no client operation traverses
  • Forgets a failed composition blocks every contributing team
  • Gives internal and partner audiences the same evolution promises
  • Plans a shared graph with no naming or deprecation policy

context