skip to content

questions

4

In GraphQL, how does a batched child lookup differ from fetching parents and children in one join?

level: juniorimportance: must knowfreq 64%

answer

  1. Two ways to stop one call per parent
  2. Trade round trips against bytes
  3. One statement duplicates the parent row
  4. Two statements, parent sent once
  5. Neither shape is in the specification

basics

~20 s

A join fetches parents and children in one database round trip and repeats each parent's columns on every child row. A batched loader makes two round trips: the parents, then one bulk lookup keyed by the parent ids collected during execution.

solid answer

~50 s

Both shapes exist to stop a nested list from firing one backend call per parent, and they spend different resources to do it. A **join** hands the store a single statement, so parents and children arrive together in one round trip — but every parent column is repeated once per child row, so a trial with 61 sites ships that trial's columns 61 times. A **batched second lookup**, the DataLoader pattern, resolves the parent list first, collects the parent keys the executor actually asked about, then issues one bulk statement such as `select ... from site where trial_id in (...)`. That is two round trips instead of one, but each row travels once and the child statement can carry its own filter, ordering and per-parent limit. Joins win when a round trip is expensive and children are few and narrow; loaders win when children are wide, numerous, or live in a different store. Neither is specified by GraphQL — both are server-side fetch strategies.

code

graphql · 11 lines
graphql
query TrialSites {
  trials(first: 25, phase: PHASE_3) {
    id
    protocolTitle
    sites {
      id
      city
      principalInvestigator
    }
  }
}

go deeper

for a junior

Be ready to state both shapes plainly: one join returning parents and children together, or two statements where the second looks up all collected parent keys at once. Naming the round-trip versus duplication trade is enough at this level.

for a middle

An interviewer expects the arithmetic. Show why 25 parents with 61 children each means 1,525 rows and 1,525 copies of the parent columns, and explain what the join cannot express — a per-parent limit or ordering.

for a senior

Demonstrate that you measure rather than assume: rows read versus distinct objects returned, bytes transferred, and where the round trip actually sits in the latency budget. Be ready to say when you would mix both strategies in one schema.

for a principal

Own the default. Decide which edges the platform joins and which it loads, make that the reviewed convention rather than a per-resolver taste call, and be clear that the choice is invisible to clients so it can be revised on evidence.

## Where the choice comes from GraphQL executes field by field: when a document selects a list of parents and then a field on each one, the child field is resolved once per parent. A page of 25 trials that each select their sites calls the `sites` resolver 25 times, and if each call hits the store on its own you get 26 statements for one document. Every serious server answers that with one of two fetch shapes, and interviewers ask about the difference because candidates usually know only one of them. ## Shape A — one join The server issues a single statement that carries parent and child together and then regroups the flat rows in memory: ``` select t.*, s.* from trial t join site s on s.trial_id = t.id where t.id in (page of 25 ids) ``` The store returns one row per (parent, child) pair. The server walks the result, groups by the trial id, and hands each `Trial` object its list of `Site` objects. One round trip, no coordination, no second component. What it costs is **duplication**. If those 25 trials have 61 sites each, the result set is 1,525 rows, and every one of them carries the trial's columns again. With a 2.1 KB trial row — protocol title, eligibility text, sponsor block — that is roughly 3.2 MB of parent columns transferred to deliver 25 distinct parents. The duplication is invisible in a one-trial test and grows with the child fan-out. A join also constrains what the child fetch can be. `limit` applies to the whole result set, not per parent, so "the 5 most recent sites for each trial" needs a window function rather than a plain join. And a `left join` invents a row of nulls for a parent with no children, which the grouping code has to recognise rather than emit a one-element list containing a null site. ## Shape B — a batched second lookup The server resolves the parent list normally, and each child-field resolution registers its parent key with a request-scoped loader instead of querying. When the executor has nothing left to run, the loader fires once with all the keys it collected: ``` select * from trial where ... limit 25 -- round trip 1 select * from site where trial_id in (...) -- round trip 2, 25 keys ``` Two round trips, and the parent row is transferred exactly once. The child statement is a statement in its own right, so it can filter, sort and apply a per-parent limit, and it can even hit a different store than the parents did. If the same trial id is reached twice in one document — once through `trials` and once through `sponsor { trials }` — the loader collapses it to one key. The price is latency: two sequential round trips instead of one, and one more per nesting level. Over a 0.4 ms in-datacentre hop that is noise; over a 38 ms cross-region hop on a five-level document it is the whole budget. ## How to choose Reason about the two costs separately. - **Round trips** favour the join. If the store is far away, or the connection is expensive to acquire, collapsing two statements into one is real latency saved. - **Bytes and rows** favour the loader. The wider the parent row and the larger the child fan-out, the more the join pays to ship the same parent columns repeatedly. - **Shape of the child fetch** usually favours the loader. Per-parent limits, per-parent ordering, and children that live in a different service or store are awkward or impossible inside one join. - **Selection sensitivity** favours the loader too, in a weaker way. A join's column list is fixed when the query is authored, so a document that selects two parent fields still ships all fourteen; a loader at least keeps that waste to one copy. (Narrowing the projection to match what the document selected is a distinct technique with its own owner.) The two are not exclusive. Most real servers join within one aggregate — a trial and its status row, which is one-to-one and duplicates nothing — and use loaders across one-to-many edges and across store boundaries. ## What the specification says Nothing. The GraphQL specification defines how fields are collected, resolved and completed into a response; it says nothing about databases, joins, statements or batching. DataLoader is the widely used name for the batching pattern, not a specified construct, and "one join or two statements" is entirely a server implementation choice that no client can observe except through latency.

  • If the join is one round trip and the loader is two, why is the loader the more common default?
    Because the round trip is usually the cheaper of the two costs. In-datacentre, a second statement adds under a millisecond, while the join's duplication grows with the child fan-out and the parent row width. The loader also keeps the child fetch independent — its own filter, ordering, per-parent limit and even its own store — and it deduplicates a key reached twice in the same document. The join's advantage becomes decisive mainly when round trips are expensive or the fan-out is tiny.
  • Which one-to-many cases are still a good fit for a join?
    Ones where the fan-out is bounded and small and the parent row is narrow — a trial and its three or four amendments, say. Also any edge that is genuinely one-to-one or one-to-few, such as a trial and its current status row, where the join duplicates nothing. The other strong case is filtering parents by a child predicate: "trials that have at least one site in this country" is a join or a subquery, because the loader can only fetch children after the parents are already chosen.
  • Does the choice between a join and a batched lookup change the response the client receives?
    No. Both produce the same JSON, because GraphQL's response shape is determined by the document's selection set and the schema, not by how the server fetched the data. That is exactly why the choice can be revisited without a client change. The only observable differences are latency, and any behavioural gap the implementation introduces by accident — for example a `left join` grouping bug that emits a one-element list containing null where an empty list belongs.

A join is going to the warehouse once and carrying back a crate that repeats the pallet label on every item; a batched loader is two trips, one for the pallet list and one for everything on those pallets, with each label written once.

saying these in an interview costs you the question

  • Claims the GraphQL specification prescribes one join per document
  • Thinks a join always means fewer bytes because it is one query
  • Says batching removes a round trip rather than adding one
  • Believes a join can apply a per-parent limit with plain LIMIT
  • Assumes the client can tell which fetch strategy the server used
  • Treats DataLoader as a specified GraphQL construct

context

open as a page

A GraphQL server replaced its batch loaders with one wide join and got slower — how do you diagnose it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Compare rows the store read against objects the response contained. A large gap means rows were multiplied by sibling joins or discarded after the fetch — typically by a filter that runs in a field resolver instead of inside the statement.

open as a page

How do you decide whether a GraphQL server resolves fields with batch loaders or compiles documents into joins?

level: principalimportance: should knowfreq 37%

basics

~20 s

Decide by how many stores back the graph, how predictable the traffic must be, and who can own a generated statement. Loaders give uniform, debuggable traffic per edge; a compiler buys round trips and costs ownership.

open as a page

Why does one join fetching a GraphQL parent list and two of its child lists return too many rows?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Because the two child tables each match the parent independently, the store emits one row per combination. A trial with 61 sites and 23 arms returns 1,403 rows, not the 84 you expected. Fetch each child list with its own statement.

open as a page