A GraphQL request hangs with DataLoader keys queued and never dispatched — what causes it?
answer
- Dispatch is an observation, not a promise
- Something is waiting where the engine cannot see
- Two fields coordinating outside the loader
- Keys loaded up, batch invocations at zero
- Nothing re-arms the pass that never ran
basics
~20 sDispatch fires only when the engine sees execution can make no further progress. A field parked on something the engine cannot recognise as a loader wait, or keys queued after the last pass, leaves nothing to re-arm dispatch.
solid answer
~50 sDispatch is triggered by an observation — every started field is parked on a loader, so nothing else can run — and that observation can be defeated. Three shapes cause the hang. A field waits on something the engine does not count as a loader wait, typically a future or lock completed by another field, so the engine believes work is progressing and never dispatches the keys that would have unblocked it. A resolver blocks a worker thread on a pending loader result — possibly the worker that would have run the dispatch pass. Or keys are queued *after* the last pass, from inside a batch function, and the implementation does not re-arm. The tell: keys loaded above zero with batch invocations at zero. Fix the coupling, never block on a pending result, and treat a scheduled dispatch or a timeout as containment rather than a cure.
code
pseudocode · 11 linesresolve Locker.compartments(locker, ctx):
rows = await ctx.compartmentLoader.load(locker.id) // engine sees: parked on a loader
ctx.compartmentsReady.complete(rows)
return rows
resolve Locker.parcels(locker, ctx):
rows = await ctx.compartmentsReady // engine sees: a field still running
return ctx.parcelLoader.loadMany(rows.map(r -> r.id))
// no stall observed -> no dispatch -> compartments never completes
// -> compartmentsReady never completes -> parcels waits forevergo deeper
Know the shape rather than the cure: a loader sends its keys only when nothing else can run, so a resolver that sits waiting on something else can stop the send from ever happening. Ask a senior before wiring two resolvers together.
Explain why the engine cannot tell a loader wait from any other wait, and why that makes a shared future between two fields dangerous. Be able to state the correct alternative: each field asks the loader itself and lets deduplication share the fetch.
Diagnose it live — keys loaded against batch invocations, parked tasks or thread dumps showing what a resolver frame waits on, then bisect the document. Reject a longer timeout as a fix, and name the ordering assumption about query field resolution as the real defect.
Own the invariant across teams: resolvers coordinate only through loaders and request-scoped setup, blocking calls on pending results are banned, and the dispatch behaviour a service depends on is verified rather than folklore. Decide whether a delay-based dispatch is an accepted platform-wide cost.
## Dispatch rests on an observation, and observations can be wrong A loader's queued keys leave when execution can make no further progress. In an engine that resolves fields on tasks or threads, that judgement is made by counting: fields started, fields completed, fields parked on a loader's pending result. When everything reachable is parked on a loader, the engine dispatches. The failure mode falls straight out of the mechanism. If a field is stuck on something the engine does **not** count as parked-on-a-loader, the engine believes work is still progressing, so it never dispatches — and the thing that field is stuck on may be exactly what a dispatch would have unblocked. Both sides wait, forever, and the request produces no response, no error and no log line until a timeout somewhere upstream gives up. ## Shape one: coordinating two resolvers by hand The one that actually happens, and it usually starts as an ordering assumption. On a parcel-locker graph, a team noticed that `Locker.parcels` needed the compartment rows that `Locker.compartments` had already fetched. `compartments` is written first in every document the client sends, so — the reasoning went — it resolves first, and `parcels` can just wait for what it produced: ```pseudocode resolve Locker.compartments(locker, ctx): rows = await ctx.compartmentLoader.load(locker.id) // engine sees: parked on a loader ctx.compartmentsReady.complete(rows) // hand-off return rows resolve Locker.parcels(locker, ctx): rows = await ctx.compartmentsReady // engine sees: a field still running return ctx.parcelLoader.loadMany(rows.map(r -> r.id)) ``` The assumption is wrong twice over. Field resolution order within a query selection set is not specified, and engines commonly resolve those fields concurrently — only a mutation's root fields are required to run serially. So both fields start. `compartments` parks on its loader, which the engine understands. `parcels` parks on a plain future, which the engine does not: to the engine that field is simply still executing, so execution has not stalled, so no dispatch happens, so the compartment keys sit in the queue, so `compartments` never completes, so `compartmentsReady` never completes. Deadlock. It shipped because the test document selected only `compartments`. The hang appeared the first time a client asked for both fields on the same locker. The fix is not a lock or a timeout. It is to let the loader be the coordination point: have `parcels` load the compartment rows itself. The loader's per-key memoization means the second ask costs nothing and shares the first fetch, and now both fields are parked on a loader, which is a state the engine can act on. ## Shape two: blocking a thread on a pending result Calling something like `.get()` or `.join()` on a loader's pending result inside a resolver looks harmless in a synchronous codebase and is the other reliable way to hang. The thread you have blocked is a worker the engine may need in order to run the remaining fields or the dispatch pass itself; you are holding it hostage against work that only that pass can start. With a small pool this can also strand a request that would otherwise have been fine, which is why it surfaces under load and not in a test. ## Shape three: keys queued after the last pass The subtler variant. If a batch function or a completion callback issues a `load` of its own, those keys arrive **after** the dispatch pass that triggered it. Whether they ever go out depends on whether the implementation re-arms — re-checks for queued keys once a batch completes, or lets the event loop schedule another dispatch naturally. Some do, some require you to opt into a strategy that does, and a chained load is exactly where you discover which. This is implementation behaviour, not specified behaviour, so verify it rather than assume it. ## Diagnosing a live one * Read the loader's counters for the stuck request: **keys loaded greater than zero with batch invocations at zero** is the whole diagnosis in one line. * Dump threads or inspect parked tasks and look at what the frames are waiting on. A resolver frame waiting on a lock, a queue, a countdown or a bare future — anything that is not the loader — is your culprit. * Bisect the document. Remove fields until the hang stops; the field you removed last and the field it was waiting for are the pair. * Reproduce with two selected fields on one object. Width does not matter here; the coupling does. ## Fixing, and the fallback that is not a fix Remove cross-field coordination from resolvers: every field should ask the loaders for what it needs and let dedup do the sharing. Never block a resolver on a pending loader result. Where a value truly must be computed once per request, compute it during request setup, before execution starts, not by having one field wait on another. A scheduled dispatch — send after a short delay whether or not execution has stalled — will unstick a deadlock, and it is worth knowing as an escape hatch, but be honest about the bill: that delay is paid on every dependency level of every request. A request timeout is not a fix either; it converts a hang into a 30-second error. Both are worth having in place while you remove the coupling that caused it.
- Why did this pass every test and only hang in production?Because the test document selected one of the two coupled fields. With only `compartments` in the selection set nothing waits on the hand-off, the engine reaches a clean stall, and the batch dispatches. The deadlock needs both fields selected on the same object — width and load are irrelevant, which is why a single-item smoke test is exactly the wrong shape to catch it.
- If a batch function itself calls load, do those keys ever go out?It depends on the implementation, and that is the honest answer. Keys queued from inside a batch function or a completion callback arrive after the pass that triggered them, so they go out only if something re-arms — a re-check once a batch completes, or an event loop scheduling another dispatch. Some implementations do this by default, some need a strategy chosen explicitly. Verify it for your engine rather than assuming.
- Would adding a delay-based dispatch be an acceptable fix?As containment, not as a fix. Dispatching after a short delay regardless of quiescence will unstick the deadlock, but the delay is paid on every dependency level of every request for every user. Pair it with a request timeout so a future occurrence fails loudly, then remove the cross-field coupling that caused the hang and take the delay back out.
The shuttle leaves only when the driver sees the platform empty; one passenger standing in a doorway the driver cannot see keeps the bus parked, and that passenger is waiting for the bus.
saying these in an interview costs you the question
- Assumes sibling query fields resolve in the order written
- Coordinates two resolvers through a shared future or lock
- Blocks a resolver thread on a pending loader result
- Raises the request timeout instead of finding the stall
- Believes a queued key always dispatches eventually
- Reproduces with one item, where the coupling never bites