skip to content

How should a DataLoader batch function report a failure that affects only one of its keys?

level: seniorimportance: should knowfreq 41%

answer

  1. One key's problem, many keys' consequence
  2. Two very different blast radii
  3. Throwing has no key to attach to
  4. An error can sit in a single slot
  5. Null already means no such record

basics

~20 s

Put an error value in that key's slot and real values in the rest. The loader rejects only that key's deferred value, so the field that asked for it gets a field error while the rest of the response resolves normally.

solid answer

~50 s

The batch function has two blast radii, and you choose deliberately. Throwing out of the function fails **every** key in the batch, because the loader has one outcome and no way to attribute it. Placing an **error value in a single slot** fails that key alone: the loader rejects just that key's deferred value, and the engine turns it into an ordinary field error carrying the path of the field that awaited it, leaving the rest of the data intact. Use per-key errors for failures genuinely about one key — a record failing a downstream check, a per-item error in a bulk response. Let it throw for failures about the batch — pool exhaustion, timeout, a 503 — where no honest per-key answer exists. Never substitute `null`: it already means *no such record*, and collapsing the two hides the failure from the response envelope.

code

pseudocode · 17 lines
pseudocode
function loadMenuItems(keys):
    try:
        result = catalogue.findAllByIds(keys)      # pool exhausted / timeout / 503
    catch e:
        raise e                                    # no honest per-key answer exists

    byId = {}
    for r in result.rows:
        byId[r.id] = r

    out = []
    for k in keys:
        if result.perItemErrors.has(k):
            out.append(Error("menu item unavailable"))   # this key only
        else:
            out.append(byId.get(k, null))                # null == no such item
    return out

go deeper

for a junior

Know that the batch function has a per-key way to report failure at all: an error value in that key's slot, distinct from a null meaning no such record. Recall that throwing out of the function affects every key in the batch.

for a middle

Explain the two mechanisms and their blast radii, and how a per-key error surfaces to the caller — a field error with the path of the field that asked, and null in the data at that path. Be ready to say why null is not an acceptable substitute.

for a senior

Show judgement about which failures are per-key and which are batch-wide, and demonstrate that you have thought about what the client sees in each case, including non-null propagation turning one bad key into a missing parent. Mention keeping the two failure classes apart in metrics.

for a principal

Own the policy: which failure classes are isolatable, what a client-visible error message may contain, how the two are counted and alerted on differently, and how nullability on loader-backed fields is chosen so that an isolatable failure stays isolated rather than bubbling into a wide hole.

## Two failure modes, and only one of them is right A batch function does bulk work on behalf of many independent field resolutions. That creates a question no single-item fetch has: when the work fails for *one* key, what happens to the other 213? The DataLoader pattern gives you two mechanisms, and they have very different blast radii. **Throwing (or returning a rejected deferred value) fails the whole batch.** The loader has one outcome to distribute and no way to tell which key it belonged to, so every deferred value it handed out for that batch is rejected with that error. In a restaurant ordering graph, one archived menu item that makes the catalogue call blow up takes down `menuItem` on every line of the page — a response that arrives half empty for a reason that concerned one row. **Placing an error value in a key's slot fails exactly that key.** The returned list still has one element per key; the element for the failed key is an error object rather than a value. The loader inspects each slot, and where it finds an error it rejects that key's deferred value only. Every other slot resolves normally. The execution engine then treats that one rejection as an ordinary field error: an entry in the response's `errors` array with the path of the field that asked for it, and `null` in the response at that path, subject to the usual non-null propagation rules of the field's type. ## When each is actually correct The choice is not stylistic; it follows the shape of the failure. * **Per-key errors** for failures that are genuinely about one key: a record that fails a downstream permission check, a row that violates an invariant the mapper cannot handle, an upstream bulk response that returned a per-item error entry for that identifier. Isolate it, let the rest of the page render. * **Whole-batch failure** for failures that are genuinely about the batch: the connection pool is exhausted, the bulk call timed out, the upstream returned 503. There is no honest per-key answer, and inventing 214 identical errors just makes the response noisier. Let it throw. Getting that backwards is the mistake interviewers listen for. Turning a pool exhaustion into 214 per-key errors hides an outage inside a partially successful response; turning one bad row into a batch-wide throw converts a one-row problem into a page-wide one. ## Why `null` is not a substitute Returning `null` in the slot for a failed key "works" in the sense that the list stays the right length and nothing throws. It is still wrong, because `null` already means something: **no such record**. Collapsing "absent" and "broken" into the same value destroys the distinction the client and your own dashboards need. A client can reasonably render "item unavailable" for an absent item; it should not render that for a catalogue outage. And the response carries no `errors` entry at all, so the failure is invisible to anyone reading the envelope — a silent hole is the worst of the three outcomes. ```pseudocode function loadMenuItems(keys): try: result = catalogue.findAllByIds(keys) # may raise: pool, timeout, 503 catch e: raise e # batch-wide: no per-key answer exists byId = index result.rows by id out = [] for k in keys: if result.perItemErrors.has(k): out.append(Error("menu item " + k + " unavailable")) # this key only else: out.append(byId.get(k) or null) # null == no such item return out ``` ## What the client sees The two mechanisms are indistinguishable in the schema and very distinguishable in the response. A per-key error produces a partial response: `data` still holds the order page, `errors` holds one entry whose `path` points at the exact line whose item failed, and that line's `menuItem` is `null`. A batch-wide throw produces an `errors` array with one entry per field that awaited the batch, all carrying the same message, and a much emptier `data`. If the field was declared non-null, the null bubbles up to the nearest nullable ancestor, which can turn one failed key into a missing order — a reason to think hard about nullability on fields fed by a loader. ## Operational notes worth saying out loud Give the per-key error a message that names the key's *class* rather than pasting an upstream exception into the response; the envelope is client-visible. Count per-key errors separately from batch failures in your metrics — they mean different things, and a rising per-key rate with a flat batch-failure rate is a data problem, not an infrastructure one. And a key that fails may be cached as a failure by the loader for the remainder of the request, so a retry within the same request will not re-attempt it; if a transient failure should be retryable, handle the retry inside the batch function where you still control the attempt.

  • Why not just return null for the key that failed? The list stays the right length and nothing throws.
    Because null already carries a meaning: no such record. Collapsing 'absent' into 'broken' destroys a distinction the client and your dashboards both need — a client may reasonably render 'item unavailable' for an absent item and must not do so for a catalogue outage. It also produces no entry in the response's errors array, so the failure is invisible to anyone reading the envelope. A silent hole is the worst of the available outcomes.
  • The batch function's bulk call times out. Per-key errors or throw?
    Throw. A timeout is a property of the batch, not of any key, and there is no honest per-key answer to synthesise. Emitting 214 identical per-key errors dresses an outage up as a partially successful response and makes the errors array useless. Keep the two paths separate in metrics as well: a rising per-key error rate against a flat batch-failure rate is a data problem, while the reverse is an infrastructure one.
  • What does a per-key failure do to the response when the field is declared non-null?
    The rejection becomes a field error, and because the field cannot hold null the null propagates to the nearest nullable ancestor — so one failed key can remove a whole order, or in a bad case a whole list, from data. That makes nullability on loader-backed fields a real design decision: declaring such a field non-null converts an isolatable per-key failure back into a wide one.

A kitchen that burns one dish sends that plate back with an apology; a kitchen whose gas is cut sends the whole table's order back. Reporting the burnt dish as a gas outage, or the gas outage as 214 burnt dishes, both misinform the room.

saying these in an interview costs you the question

  • Throwing from the batch function for one bad key
  • Returning null to mean the lookup failed
  • Assuming a throw fails only the offending key
  • Emitting one per-key error each for a timeout
  • Pasting the raw upstream exception into the response
  • Expecting a within-request retry to re-attempt a failed key

context