What contract must a DataLoader batch function honour between the keys it receives and the values it returns?
answer
- One call, many items
- What goes in and what must come out
- Two lists that must match
- Same length, and index means key
- Null and errors still occupy a slot
basics
~20 sA DataLoader batch function takes a list of keys and returns a list of values of the same length, with each value at its key's index. Absent keys get null; failed keys get an error value.
solid answer
~50 sThe loader collects many single-item requests and calls your batch function **once**, with an ordered list of keys. Your function must return a list of the same length, positionally aligned: the value at index *i* is the value for the key at index *i*. Position is the only correlation mechanism — the loader does no id matching — so a shorter, longer or reordered list silently attaches values to the wrong keys. A key with no matching record still occupies a slot, holding `null` for a to-one field or an empty list for a to-many field. A key whose lookup failed holds an error value, which rejects that key's deferred value alone. None of this is in the GraphQL specification: it is a convention of the DataLoader pattern, reimplemented with the same shape across languages.
code
graphql · 25 linestype Order {
id: ID!
lines: [OrderLine!]!
}
type OrderLine {
quantity: Int!
menuItem: MenuItem
}
type MenuItem {
id: ID!
name: String!
priceCents: Int!
}
query OrderPage($first: Int!) {
orders(first: $first) {
id
lines {
quantity
menuItem { name priceCents }
}
}
}go deeper
Recall the two lists and the rule that ties them: keys in, values out, same length, same order. Be able to say what fills the slot for a key that had no record, and to write the four-line batch function that guarantees it.
Explain why correlation is positional rather than by id, and what that buys and costs. Be ready to state what a to-one slot versus a to-many slot holds when nothing was found, and to say plainly that the contract is a convention rather than specification text.
Show that you treat the batch function as the place correctness is enforced: it is written against the key list, it is tested with several keys including a missing one, and its failure modes are deliberate. An interviewer expects you to name misalignment as a silent-wrong-data class, not a crash class.
Own it as a house rule rather than a habit: one reviewed helper that turns a bulk lookup plus a key list into an aligned result, applied across every loader in the estate, so no team re-derives the mapping step and no service ships a batch function that trusts row order.
## What the pattern actually asks of you The DataLoader pattern has two halves. The half you *call* is a `load(key)` function, invoked once per item from wherever a field needs a related object; it returns a deferred value (a promise, future or task, depending on the language) rather than the object itself. The half you *write* is the **batch function**: a single function the loader hands many keys to at once, which produces the values for all of them. Every benefit of the pattern rests on that function honouring one narrow contract, and almost every bug people attribute to "the loader being broken" is that contract being violated. The contract has three clauses. **One list of keys goes in.** The loader hands your function an ordered list — say 214 menu-item ids collected while resolving the lines of an order page. Your function is called once for that list. When the loader decides to make the call is a separate concern; your side of the contract is only about what to do once it does. **One list of values comes back, of exactly the same length.** 214 keys in, 214 values out. Not 211 because three ids no longer exist. Not 218 because a join fanned out. The list must be the same length even when the underlying store had nothing to say about some of the keys. **Position is identity.** The value at index *i* belongs to the key at index *i*. There is no id matching, no name matching, no "the loader will figure it out". The loader resolves the deferred value it handed back for the *i*-th key with the *i*-th element of your list, and that is the whole mechanism. ## Why positional and not a keyed map Positional alignment is the cheapest possible correlation: it needs no equality function for keys, works for keys that are tuples or objects rather than scalars, and works for a key type the loader knows nothing about. The cost is that alignment is invisible — nothing about a list of values announces which key each one is for, so a mistake produces wrong data rather than a type error. That is the tradeoff the pattern deliberately took, and it is why the discipline of rebuilding the result *from the key list* matters so much. ## Null and errors are values too A key with no matching record is not an omission; it is a slot holding `null`. A key whose lookup failed is not an omission either; it is a slot holding an error value, which the loader will use to reject just that key's deferred value. Both keep the list length intact. For a field that returns a **list** — the lines of an order rather than the menu item of a line — the value in the slot is itself a list, and a key with no children gets an empty list, not `null`, so the field's own type stays honest. ## A worked example A restaurant ordering graph has `Order`, `OrderLine` and `MenuItem`. A page of orders resolves 1,137 lines, each of which needs its `menuItem`. Without batching that is 1,137 lookups. With a loader, the executor collects the ids, the loader calls your batch function once with 214 distinct ids, and your function issues one bulk lookup. Three of those ids are archived items that were purged from the catalogue table. The correct return is a 214-element list with `null` in exactly those three positions — the three lines get `"menuItem": null` and every other line gets its item. ```pseudocode function loadMenuItems(keys): # keys = [id_0, id_1, ... id_213] rows = catalogue.findAllByIds(keys) # may return fewer, in any order byId = index rows by row.id return [ byId.get(k) or null for k in keys ] # always len(keys) long ``` The last line is the entire contract in one expression: iterate **the keys**, in their order, and produce one value each. ## This is a convention, not the specification Nothing in the GraphQL specification mentions batch loading, batch functions, or key lists. The specification describes a type system and an execution algorithm; it deliberately says nothing about how a field obtains its data. The keys-in/values-out contract is a convention popularised by the DataLoader pattern and reimplemented, with the same shape, across ecosystems. Say so in an interview: it signals you can tell specified behaviour from community practice, which is a real distinction on this technology. ## What the contract buys Because the batch function is a pure adapter from a key list to a value list, it composes with everything else. The field's resolver stays a one-liner that asks for a single object; the number of backend round trips drops to one per loader per batch; and the fan-in is invisible to the schema — nothing about the SDL changes when you introduce a loader behind a field. That property is exactly why the pattern survived: it is a data-access concern that never leaks into the type system or the document.
- The batch function is asked for 214 keys and the store has records for 211 of them. What exactly do you return?A 214-element list, with `null` in the three positions whose keys had no record and the found value everywhere else. Returning 211 values breaks the contract: the loader aligns by position, so every value after the first gap would be handed to the wrong key. If the field were a to-many field instead, the three empty slots would hold empty lists rather than `null`, so that a non-null list type stays satisfiable.
- Why does the pattern correlate values to keys by position instead of by an id on the value?Because position needs nothing from the key or the value: no equality function, no id field, no assumption that a key is a scalar. It works for tuple keys and for value types the loader knows nothing about. The cost is that misalignment is invisible — it produces wrong data rather than a type error — which is why you rebuild the result by walking the key list rather than the rows.
- Is any of this defined by the GraphQL specification?No. The specification defines a type system and an execution algorithm and says nothing about how a field obtains its data — no batching, no loaders, no key lists. The keys-in/values-out batch contract is a convention of the DataLoader pattern that independent implementations converged on. Being able to draw that line between specified behaviour and community practice is itself something interviewers on this technology listen for.
It is a coat check that takes a stack of numbered tickets and hands back a stack of coats: nobody reads the labels, they just count down the pile, so one missing coat mis-dresses everyone behind it.
saying these in an interview costs you the question
- Returning only the records that were found
- Assuming the loader matches values to keys by id
- Calling the batch function once per key
- Claiming the GraphQL specification defines batch loading
- Returning null for a to-many key instead of an empty list
- Treating a missing record as an error value