skip to content

questions

3

What must a DataLoader batch key satisfy for one bulk lookup to answer every key?

level: juniorimportance: must knowfreq 60%

answer

  1. Think about what the batch function is told
  2. The resolver knows things the key omits
  3. The loader must compare two keys
  4. Many keys, one backend call
  5. Field arguments change the value too

basics

~20 s

A batch key must fully determine the value on its own, be comparable by value so the loader can deduplicate it, and differ from the other keys only along a dimension one bulk lookup can vary.

solid answer

~50 s

The key is the entire contract between a resolver and the batch function of the DataLoader pattern — a widespread convention, not specification text. Three things must hold. First, **sufficiency**: whatever the resolver knows and the key omits, the batch function cannot see, so a field whose value depends on a parent id *and* an argument needs both in the key. Second, **comparability**: the loader deduplicates and memoizes per key inside one request, which requires deciding whether two keys are the same, so a key must compare by value rather than by object identity. Third, **one-lookup shape**: the whole list of keys must collapse into a single request to the backend — typically one predicate over a set of values. A key that varies along something a single lookup cannot express, such as a per-key row limit, forces the batch function to loop, which is an N+1 wearing a loader's clothes.

code

graphql · 10 lines
graphql
type Account {
  id: ID!
  displayName: String!
  balance(asOf: Date!): Money!
  statements(period: StatementPeriod!): [Statement!]!
}

type Query {
  accounts(ids: [ID!]!): [Account!]!
}

go deeper

for a junior

Be ready to state the three requirements plainly: the key fully determines the value, the loader can compare two keys, and all the keys in a batch answer with one backend call. Naming a field argument as part of the key is the detail interviewers listen for.

for a middle

Explain the mechanics behind each requirement — why a missing argument produces a wrong value rather than an error, why object keys break deduplication, and what a single bulk lookup can and cannot vary across keys.

for a senior

Show how you catch bad keys before production: reviewing a key against the field's full input surface, and reading a trace for a loader that batched cleanly while the backend still saw one call per key. Note that single-row fixtures hide all of it.

for a principal

Own the convention. Decide what the house rule is for constructing keys, where loader instances are built and scoped, and how key design is reviewed, so a small team is not re-deriving the sufficiency argument on each new loader.

## What a batch key is actually doing The DataLoader pattern sits between a field resolver and a data source. It is a **convention** — nothing in the GraphQL specification or the GraphQL over HTTP specification mentions it — but it is close to universal, because the execution algorithm resolves a list's child fields item by item and something has to put those calls back together. The shape is always the same. Resolvers call `load(key)`. The loader accumulates the keys reached during one stretch of execution and hands the whole list to a **batch function**, which is expected to return one value per key. That makes the key load-bearing in a way people underestimate: **the key is everything the batch function is told.** It is not a hint, not an id it can enrich from context — it is the complete description of the work. Three properties follow from that, and a key that misses any one of them fails in a different way. ## 1. Sufficiency: the key must fully determine the value Take a bank statements graph: ```graphql type Account { id: ID! displayName: String! balance(asOf: Date!): Money! statements(period: StatementPeriod!): [Statement!]! } ``` A loader for `Account.balance` keyed by account id alone is already wrong, and it is wrong in the quietest possible way. Two selections in the same document — one asking the balance as of the quarter end, one as of today — produce the same key. The loader dedupes them, calls the backend once, and hands the same value to both fields. Nothing throws. One of the two fields is simply reporting a number for a date nobody asked about. The rule is mechanical: enumerate every input that can change the answer — the parent's identifier, each field argument the backend actually honours — and every one of them belongs in the key. The converse rule matters too. Anything **constant for the whole request** — the viewer, a trace id, the connection — does *not* belong in the key. It belongs to the loader instance the request built. Folding it into every key adds bytes and buys nothing, because it never varies within the batch. ## 2. Comparability: the loader must be able to tell two keys apart Inside one request the loader memoizes by key, so repeated `load` calls for the same key hit the backend once. That requires an equality test. Primitive keys — an id string, an integer — compare by value and behave. A key expressed as an object or a tuple usually compares by **reference** in the host runtime, so two structurally identical keys look like two different keys: the deduplication silently stops working, the batch grows, and the same row is fetched several times in one request. The standard remedy is a cache-key function that projects a non-scalar key onto a stable primitive, which is a design exercise of its own. ## 3. One-lookup shape: all keys in a batch must collapse into a single request This is the property candidates skip, and it is the one that decides whether batching pays at all. Batching is only a win if the whole key list becomes **one** call to the backend — in the common case, one predicate over a set of values: *return every row whose account id is in this set*. So the keys must differ only along a dimension that a single lookup can vary. Ids: fine, they go into the set. A statement period that is the same for every key: fine, it rides along as a constant in the query. A **per-key row limit** — "the five most recent transactions for each of these accounts" — is not fine, because a flat set-membership predicate has no way to say *five per account*. Neither is a per-key sort order or a per-key free-text filter. When the keys cannot be merged, the batch function has only one honest move left: loop over the keys and issue one call each. The trace then shows a loader that batched perfectly and a database that was asked eighty-seven separate questions. ## A quick check you can apply in an interview Before writing the loader, say the key out loud as a sentence, then write the single backend call that answers a list of them. * "The balance of account `A-4471` as of `2026-03-31`" — sentence is complete, and a list of them is one lookup over pairs. Good key. * "The statements of account `A-4471`" — incomplete; the field takes a period. * "The five most recent transactions of account `A-4471`" — complete, but a list of them is not one flat lookup. The key is honest and the *lookup* is the problem, which is a different fix. A four-person platform team can carry that check in their heads during review, which is roughly the point: bad keys do not announce themselves in tests, because a test with one account in the fixture batches one key and looks perfect.

  • Should the viewer's identity be part of the batch key?
    No. Within one request the viewer never varies, so putting it in the key cannot prevent a mix-up and only makes every key longer. Viewer scoping is a property of the loader instance the request builds, not of the key. Put in the key only what varies *between keys in the same batch*.
  • A field takes an argument that every selection in the document passes identically. Does it still belong in the key?
    Yes, if it changes the value. Sufficiency is about correctness, not about whether the argument happens to vary today: the next document may select the field twice with different arguments, and a key that omits it would serve one selection's value to both. An argument that is identical across the batch costs nothing at lookup time — it becomes a constant in the single query.
  • How would you notice, in review, that a key is missing an input?
    Compare the key's fields against the field's full input surface: the parent identifier plus every argument the resolver actually passes downstream. If the resolver reads something it does not forward into `load`, that is the gap. The runtime symptom is the giveaway too — wrong-but-plausible values that only appear when one document selects the same field twice with different arguments.

A batch key is the order slip handed through to the kitchen: whatever is not written on it, the cook cannot make, and two slips only count as one order if someone compares what they say rather than which piece of paper they are.

saying these in an interview costs you the question

  • Says any value at all can serve as a batch key
  • Keys by the parent id and ignores field arguments
  • Assumes the loader compares object keys by content
  • Claims the DataLoader pattern is part of the GraphQL specification
  • Thinks batching still pays when the lookup runs per key
  • Puts the viewer in every key instead of scoping the loader

context

open as a page

How do you key a DataLoader batch when the value depends on more than one input?

level: middleimportance: should knowfreq 48%

basics

~10 s

Use a composite key holding every input that changes the value, and give the loader a cache-key function that projects that object or tuple onto a stable primitive string so structurally identical keys deduplicate.

open as a page

A DataLoader batch function gets all 87 keys in one call yet issues 87 backend queries — what key design causes that?

level: seniorimportance: should knowfreq 44%

basics

~10 s

The keys carry an input that one bulk lookup cannot vary per key — a per-key filter, sort or row limit. Collection worked; the lookup cannot express the variation, so the batch function loops.

open as a page