How do you decide whether a GraphQL server resolves fields with batch loaders or compiles documents into joins?
answer
- Not a performance question first
- How many stores back the graph
- Uniform statements versus generated ones
- Who debugs it at three in the morning
- Invisible to clients, so reversible
basics
~20 sDecide by how many stores back the graph, how predictable the traffic must be, and who can own a generated statement. Loaders give uniform, debuggable traffic per edge; a compiler buys round trips and costs ownership.
solid answer
~60 sTreat it as an ownership decision, not a performance one. A **loader layer** issues one bulk lookup per edge per request: traffic is uniform, every statement is one a person wrote and can read in a trace, and the same layer works whether an edge is backed by the same store, another store or another service. A **document compiler** folds a whole selection into a handful of joined statements, buying round trips and, on deeply nested reads, real latency — but it emits a different statement per document shape, so the store sees enormous statement variety, a slow operation is a machine-generated statement nobody wrote, and the component becomes load-bearing for a team that may be four people. In a graph split across 11 services, the compiler is only even available inside the one service that owns both sides of an edge, which caps its reach. The defensible default is loaders everywhere, joins inside a single aggregate, and a compiler only for a measured, hot, deeply nested read path.
code
pseudocode · 10 lines# Loader baseline: one bulk lookup per edge, shape known in advance
sites = load("site.byTrialId", trialIds) # same store
arms = load("arm.byTrialId", trialIds) # same store
sponsors = load("sponsor.byId", sponsorIds) # another service
adverse = load("adverse.byTrialId", trialIds) # another store
# Compiler: one statement per document shape, only for edges one store owns
plan = compile(document.selectionSet)
# -> select ... from trial t join trial_status st on ... where t.id = any(:ids)
# sponsor and adverse edges cross a service boundary: still loadedgo deeper
You are not expected to make this call, but know that the same document can be served by many small keyed lookups or by fewer joined statements, and that clients cannot tell the difference.
Be able to explain why a loader layer produces a small, predictable set of statements while a generated one produces a statement per document shape, and why that matters when something is slow.
Argue the case with evidence: round trips on the deepest common path, the share of traffic it carries, and what changes if an edge crosses a service boundary. Be able to say what would make you revisit the choice.
Own it as a staffing and blast-radius decision as much as a latency one. Set the default, name the edges that may deviate, keep the strategy invisible to clients so it stays reversible, and say who maintains a compiler two years from now.
## What is actually being decided Both strategies solve the same problem — bounding the backend calls one document causes — and both can be made fast. What differs is where complexity lives, who can operate the result at 03:00, and how far the approach reaches across a graph that is not one database. ## The criteria that decide it **How many stores back the graph.** A join exists only where one component can address both sides of an edge in one statement. In a registry graph served by 11 services, the edges inside one service are joinable and the edges between services are not, at any price. A compiler therefore covers a fraction of the graph and something else has to cover the rest — which usually means you are running both strategies regardless, and the question becomes how much of the graph justifies the second one. **Predictability of the traffic.** A loader layer produces a small, fixed set of statement shapes: one bulk lookup per edge, parameterised by a key list. You can enumerate them, review them, and reason about the load a new operation adds. A compiler produces a statement per document shape, and clients invent document shapes. The store sees far more distinct statements as a result; how a particular engine copes with that variety is its own subject, but the operational fact is that you can no longer hold the query set in your head. **Who owns a slow statement.** With loaders, a slow statement is one a person wrote, in a file, with a name. With a compiler, it is generated, its shape depends on a client's selection set, and diagnosing it means reasoning about the generator. That is fine with a team that has the depth to own a query compiler and dangerous with a team of four who inherited one. **Blast radius.** A regression in one loader degrades one edge. A regression in the compiler degrades every operation at once, and rolling it back is rolling back the fetch layer for the whole graph. **What the edges look like.** Wide parents with large fan-out, per-parent limits, per-caller predicates and cross-store edges all argue for loaders. Narrow, low-fan-out, one-to-one or one-to-few edges inside one store are exactly where a join is free and a loader is a wasted round trip. Deeply nested read paths — five levels, each adding a sequential round trip — are where a compiler's advantage is largest, because loader latency compounds per level while a join does not. **Latency budget and topology.** If the store is a sub-millisecond hop away, the round trips a compiler saves are close to free anyway and its case is weak. If the store is remote, or connection acquisition is expensive, the same saving is the dominant term. ## A defensible default 1. **Loaders as the baseline for every one-to-many edge.** Uniform, reviewable, portable across stores, and understood by anyone who has worked on a GraphQL server before. 2. **Plain joins inside one aggregate.** One-to-one and one-to-few edges owned by the same service — a trial and its status, a site and its address — join with no duplication and no extra round trip. This is a code convention, not a framework. 3. **A compiler only where measured evidence demands it**, scoped to a named read path rather than adopted graph-wide, and only if the team can staff its ownership. "We might need it later" is not evidence; a p99 dominated by five sequential loader hops on a path that serves most of your traffic is. ## How to make the call reviewable Write down the numbers the decision turns on, so it can be revisited without re-arguing taste: rows read per object returned, statements per operation, round trips on the deepest common path, and the share of traffic that path carries. Then set the review trigger in advance — the fan-out or nesting depth at which you would reconsider — rather than waiting for an incident to force it. ## What makes this reversible The strategy is invisible to clients. Both approaches produce identical responses for the same document, because the response shape comes from the schema and the selection set, not from how the data was fetched. That is the single most useful property in this decision: it means you can start with the simpler option, prove the harder one on a slice of traffic behind a flag, and revert without a client migration. Any design that leaks the fetch strategy into the schema — a field that exists only because the compiler can serve it, a nullability choice made to accommodate a join — has thrown that property away, and is worth rejecting on those grounds alone. ## The organisational reading A query compiler is a platform component with a permanent owner, an upgrade path and an on-call story. Adopting one is a staffing commitment as much as a technical choice, and the honest question in a review is not "is it faster" but "who maintains this in two years, and what happens to the graph if that person leaves".
- What evidence would actually justify adopting a document compiler?A measured latency budget dominated by sequential round trips, not by the store. Concretely: a common read path five levels deep where loader hops account for most of p99, that path carries a large share of traffic, and a prototype on a slice of that traffic shows the improvement holding under real document variety. If the win only appears on a synthetic deep document nobody sends, or the store dominates the budget anyway, the evidence does not support it.
- How do you keep two fetch strategies in one codebase from becoming incoherent?By drawing the boundary on a property of the edge rather than on who wrote the resolver. A workable rule is: one-to-one and one-to-few edges inside a single store may join; everything else loads. That is checkable in review, explains itself to a newcomer, and does not depend on taste. What goes wrong is per-resolver improvisation, where the same edge is fetched two ways in two code paths and neither is obviously wrong.
- Does either strategy constrain how the schema is designed?It should not, and preserving that is part of the decision. The response shape comes from the schema and the selection set, so any given document must return the same data under either strategy. If a field exists only because the compiler can serve it cheaply, or a list is nullable to accommodate a join's null rows, the fetch layer has leaked into the contract — which makes the strategy irreversible and forces a client migration to undo it.
saying these in an interview costs you the question
- Picks the strategy on benchmarks alone, ignoring ownership
- Assumes a compiler can join across service boundaries
- Ignores the statement variety a generator creates
- Lets the fetch strategy shape the schema contract
- Adopts a query compiler before measuring round-trip cost
- Treats it as an all-or-nothing choice for the whole graph