A DataLoader batch function gets all 87 keys in one call yet issues 87 backend queries — what key design causes that?
answer
- The batch collected; the lookup did not
- Ask what varies across the keys
- An argument riding inside every key
- Constant arguments cost nothing
- Five per account resists one predicate
basics
~10 sThe keys carry an input that one bulk lookup cannot vary per key — a per-key filter, sort or row limit. Collection worked; the lookup cannot express the variation, so the batch function loops.
solid answer
~50 sBatching has two halves and only the first worked. The loader gathered all 87 keys, so collection is fine; the failure is that the key list does not collapse into one backend call. Two causes. Either the key carries a **field argument that varies across the keys** — a filter, a sort, a page size — which the bulk lookup cannot apply differently per key; or the key shape is inherently un-mergeable, most often a **per-key row limit** ("the five most recent transactions for each account"), which a flat set-membership predicate cannot express. Diagnose by counting *distinct argument tuples* against the key count. Fixes, in order: group keys by argument tuple and issue one call per group, hoist a request-constant argument onto the loader instance, or use a lookup form that returns top-N per key. An argument identical across the batch never fragments anything.
code
graphql · 6 linesquery Overview($ids: [ID!]!) {
accounts(ids: $ids) {
id
transactions(first: 5) { id postedAt amount }
}
}go deeper
Take away the core idea: a loader can collect all the keys and still not save any work, because the keys have to be answerable by one backend call. Field arguments are what usually makes them not.
Explain the mechanics of the two causes — a varying argument that the lookup cannot apply per key, and a per-key row limit that a flat set predicate cannot express — and be able to describe grouping keys by argument tuple.
Show the diagnosis and the ordered remedies: measure key count against distinct argument tuples, count calls per batch, then group, hoist a constant argument onto the loader instance, or change the lookup form. Say plainly when a single call is unachievable.
Own the surface that creates the fan-out. Deciding which arguments a list field accepts, whether page sizes are constrained, and where loaders are constructed determines how much of this a team can ever fix downstream of the schema.
## The symptom, and why it is confusing A four-person platform team instruments a bank statements graph and finds a document that resolves 87 accounts and their recent transactions. The trace looks contradictory: the loader dispatched **once** with all 87 keys, exactly as intended, and the database still logged 87 queries. Everyone's first instinct is that the batch broke up. It did not — batching is two halves, *collecting the keys* and *answering them in one call*, and only the second half failed. So the question is never "why did dispatch split" here. It is: **why can these 87 keys not be answered together?** ## Cause one: an argument in the key that varies When a field takes arguments that change its value, those arguments have to be in the key. That is a correctness requirement and it is not negotiable. But it has a consequence: the keys now differ along a dimension the bulk lookup must handle. The crucial distinction — and the one interviewers are listening for — is between an argument that is **constant across the batch** and one that **varies**: * Every selection passes `period: LAST_QUARTER`. All 87 keys carry the same period. The lookup is one query with the account ids in a set and the period as a constant. Nothing fragmented. * The document selects statements for several different periods across different accounts. Now the key list spans, say, three distinct periods, and one flat set-membership predicate cannot say "these 40 accounts for Q1 and those 47 for Q2". So the mistake is not putting the argument in the key. The mistake is assuming a key list that varies along an argument is automatically answerable in one call. ## Cause two: a key that no single lookup can satisfy Some keys are perfectly well-formed and still cannot be merged. The canonical one is the **per-key row limit**: ```graphql query Overview($ids: [ID!]!) { accounts(ids: $ids) { id transactions(first: 5) { id postedAt amount } } } ``` Every key is `(accountId, first: 5)`. The argument does not even vary. But "the five most recent per account" is a top-N-per-group problem, and a plain predicate over a set of account ids returns *all* the rows for those accounts, not five each. The naive batch function does the only thing available to it and loops. Per-key sort orders and per-key free-text filters land the same way. ## Diagnosing it deliberately Do not guess between the two causes — measure inside the batch function: 1. Log the key count and the count of **distinct argument tuples** for each batch. If distinct-tuples is near 1, the arguments are effectively constant and cause two is your problem. If distinct-tuples tracks the key count, you have genuine fan-out. 2. Count backend calls per batch, not per request. A loader that halves the call count still looks like a win in a request-level metric while hiding a loop. 3. Test with a realistic key set. A fixture with one account produces one key, one call, and a green test under every one of these defects — which is exactly why the defect survives to production. ## Fixes, roughly in order of preference **Group inside the batch function.** The most broadly useful move and the one most people miss: partition the keys by argument tuple and issue one call per *group*. Eighty-seven keys spanning three periods becomes three queries, not 87 and not 1. It requires nothing to be constant and it degrades gracefully as the argument space widens. **Hoist a request-constant argument out of the key.** If an argument never varies within a request, the key can be the plain id and the argument can live on the loader instance the request built — effectively one loader per argument shape, created lazily. This keeps keys small and lookups flat. The hazard is loader proliferation: if the argument space is large, you are minting a loader per distinct value, so cap it and fall back to grouping. **Use a lookup form that expresses per-key limits.** For top-N-per-group, one round trip is achievable with a query form built for it — a windowed or per-key-limited query — rather than a flat set predicate. This is a data-access design change, not a key change, and it is the honest fix for cause two. **Over-fetch and trim in memory.** Fetch all rows for the key set in one call and slice per key in the batch function. Safe only when the per-key row count is genuinely bounded: with 87 accounts and 1,433 transactions between them it is trivial; against accounts with six-figure transaction histories it converts a query-count problem into a memory incident. **Shrink the argument surface.** If page sizes are arbitrary, every caller invents a new key shape. Constraining the field to a small set of accepted page sizes turns unbounded fan-out into a handful of groups — a schema-side decision, but often the one with the highest leverage. ## The judgement to voice Sometimes the argument really is per-caller unique — a free-text search term, say. Then the loader is buying deduplication and nothing else, and the mature answer is to say so, count that as the expected behaviour, and put the effort into bounding the fan-out instead of pretending a loader fixed it. Claiming a loader guarantees one query per batch is the answer that ends the conversation badly.
- Does putting a field argument in the batch key always fragment the batch?No, and this is the distinction that matters. An argument that is identical across every key in the batch simply becomes a constant in the single lookup — the ids still go into one set predicate. Fragmentation happens only when the argument *varies* across keys and the lookup cannot apply a different value per key. Omitting a value-affecting argument to protect batching trades a throughput problem for a wrong-value bug.
- When would you deliberately accept one backend call per key?When the varying input is genuinely per-caller unique — a free-text filter, an arbitrary page size — so grouping collapses nothing. At that point the loader is buying deduplication only. The right response is to say so explicitly, bound the fan-out at the schema or limiter level, and stop treating the loader as the mitigation.
- What would you measure to confirm the fix worked?Backend calls per batch, not per request — a request-level count can improve while a loop survives inside one batch function. Log key count and distinct argument tuples per batch alongside it, and run it against a realistic key set: a single-account fixture produces one key and one call, and looks identical before and after.
saying these in an interview costs you the question
- Blames dispatch timing instead of the key shape
- Says any argument in a key breaks batching
- Claims a loader guarantees one query per batch
- Fetches every row per key and trims without bounding it
- Mints an unbounded loader per distinct argument value
- Declares it fixed without counting calls per batch