skip to content

Why is a DataLoader instance created per request rather than shared across requests?

level: middleimportance: must knowfreq 61%

answer

  1. Ask what makes the map correct at all
  2. No TTL, no eviction, no invalidation
  3. The key records the id, not the asker
  4. Two callers, same key, different question
  5. Build them where the request context is built

basics

~20 s

Because the loader's memo map has no expiry, no size bound and no notion of who asked. Confined to one request it cannot go stale or reach another caller; shared between requests it serves old values, leaks one viewer's data to the next, and grows without limit.

solid answer

~50 s

A loader's memoization is correct only because it is short-lived. It has no TTL, no eviction and no invalidation hook, so the only thing keeping its entries honest is that a request is too brief for the underlying rows to change. Widen that window and three problems appear at once. Values go stale, with no mechanism to notice. Entries fetched while executing one caller's request get handed to the next caller, whose visibility of that data may be different — the key is usually a bare id, so nothing in the map records whose request populated it. And the map grows for the life of the process, since nothing ever evicts. The standard shape is therefore to construct a fresh set of loaders when the per-request execution context is built, and to reach them only through that context. This is convention rather than anything the GraphQL specification says, but it is close to universal.

code

pseudocode · 16 lines
pseudocode
// WRONG: one instance for the life of the process
sharedDriverLoader = new Loader(batchLoadDrivers)

// RIGHT: a fresh set per request, owned by the execution context
function buildContext(httpRequest):
    principal = authenticate(httpRequest)
    return {
        principal: principal,
        loaders: {
            drivers: new Loader(keys -> batchLoadDrivers(keys, principal)),
            devices: new Loader(keys -> batchLoadDevices(keys, principal))
        }
    }

resolve Trip.driver(trip, args, ctx):
    return ctx.loaders.drivers.load(trip.driverId)   // only ever via ctx

go deeper

for a junior

Remember the rule and one reason for it: a fresh set of loaders per request, because the memo map never expires and would otherwise hand old values to the next caller.

for a middle

Explain all three consequences of a long life — staleness with nothing to invalidate it, one caller receiving values fetched for another, and a map that only grows — and show where the instances should be constructed instead.

for a senior

Demonstrate that you would find this from symptoms: intermittent stale reads reported separately from a slow heap, neither obvious in a single-user process. Be ready to defend the structural fix over a per-field rule.

for a principal

Own the boundary between a request-scoped memo map and a real shared cache: what an organisation may reuse across callers, what must key on the asker, and who is accountable for expiry and invalidation on the shared tier.

## The lifetime is the design Every property that makes a loader's memo map safe is a consequence of it being thrown away at the end of a request. It is worth stating that as a claim, because it explains why the rule has no exceptions in the default case: the map has no TTL, no maximum size, no eviction policy and no way to hear that a row changed. A cache with none of those things is only correct if it does not live long enough to be wrong. Confine it to one request and it cannot be. ## The three failures of a shared instance **Staleness with no way to notice.** Once an entry is in the map it is final. If `vehicle-4417` is memoized at 09:00 and the process holds the loader all day, every request for that vehicle for the rest of the day gets the 09:00 row — including requests that ran *after* a mutation the same process handled. The symptom is worse than a plain cache's, because there is no configured freshness anyone can point at and reason about. Someone reading the code sees "DataLoader", assumes request scope, and never suspects the value. **The viewer problem.** This is the one interviewers push on. Loader keys are normally the bare identity of the thing being fetched — an id. If the batch function's lookup is at all sensitive to who is asking, then two callers with the same key are not asking the same question, but the map cannot tell them apart. The result is a caller receiving a value assembled during someone else's request. Two shapes of this are common: a batch function that filters rows by the requesting principal, and one that returns a row whose contents differ by caller. Both are silent — nothing errors, nothing logs, and the wrong data simply appears in a normal-looking response. Keeping the loader per request removes the class outright, which is why the fix is structural rather than a rule someone must remember to apply per field. **Unbounded growth.** No eviction plus a long life means the map only ever gets bigger. Every distinct key the service has ever been asked for stays resident. ## A failure that only shows in production A fleet telematics service built its loaders once, at startup, and stored them on a module-level object rather than on the per-request context. Nothing looked wrong for months. In development the process handled one operator, restarted on every file save, and the map never exceeded a few thousand entries. In staging a nightly redeploy did the same job. In production the same process served every depot continuously; by the end of a deploy cycle the driver and device maps together held around 1.8 million entries. The visible symptom was a depot-week document that timed out only in production, intermittently, with no slow backend calls in the trace — the time was going into a heap that had grown enough to make collections long and frequent. The staleness and cross-caller reads had been happening the whole time and had been reported as "the dashboard sometimes shows yesterday's driver", filed separately, and never connected to the timeouts. The fix was one line in the wrong place moved to the right one: build the loaders in the function that builds the execution context, so each request gets its own set and drops them when it finishes. ## Where the instances live The conventional structure is that the per-request context — the object execution threads through every resolver — owns the loaders, and resolvers only ever reach a loader through that context. Two rules keep it honest: * nothing constructs a loader at module scope or in a singleton; * nothing reaches a loader other than through the context argument, so a resolver physically cannot get hold of a longer-lived one. Both are cheap to enforce with a lint rule or a review habit, and both are easier than diagnosing the failure above. ## The deliberate exception, and its price People do sometimes want cross-request reuse, and it can be done — but as a *second tier*, not by extending the loader's life. The shape is a per-request loader whose batch function consults a shared store, so the shared layer is an explicit cache with the properties the memo map lacks: a bounded size, an expiry, an invalidation path, and keys that include whatever dimension makes two callers' questions different. Making that a separate layer keeps the distinction visible; a candidate who proposes it should be able to say what the expiry is and what changes the key, because those two answers are precisely what the request-scoped map does not have. ## Specified or not None of this is in the GraphQL specification, which says nothing about loaders. What the specification does supply is the per-request execution context that the convention hangs on, and the freedom for a server to resolve fields in any order within a query selection set. The per-request rule is an ecosystem convention that arrived with the original pattern and never seriously changed.

  • If the batch function already filters by the caller, is a shared loader safe?
    No — that is the dangerous case, not the safe one. The filtering happens once, when the key is first fetched, and the filtered outcome is what the map keeps. The next caller supplies the same id, hits the entry, and receives a value computed for someone else's principal. Either the identity has to be part of the key, which mostly defeats sharing, or the loader stays per request.
  • You do want cross-request reuse for reference data that barely changes. How do you build it?
    As two tiers. Keep the per-request loader for deduplication, and have its batch function consult a separate shared cache that has the properties the memo map lacks: a bounded size, an explicit expiry, an invalidation path, and a key that includes anything making two callers' questions different. Keeping them separate keeps the distinction reviewable, rather than quietly extending a memo map's life.
  • How would you catch a wrongly-scoped loader before it reaches production?
    Structurally rather than by testing for the symptom: forbid constructing a loader anywhere but the context factory, and forbid resolvers reaching one other than through the context argument. A test can also assert that two sequential requests in one process do not share loader identity. Symptom-level tests are unreliable because a single-user development process almost never reproduces it.

saying these in an interview costs you the question

  • Creates one loader at startup and reuses it
  • Thinks the memo map expires on its own
  • Assumes filtering in the batch function makes sharing safe
  • Treats cross-request reuse as a free performance win
  • Cannot say what would evict an entry
  • Calls per-request scoping a rule from the specification

context