A GraphQL server's request-scoped context is corrupted only under load — how do concurrent resolvers cause that?
answer
- Request-scoped bounds lifetime, not sharing
- Dozens of writers, one context object
- Width of the selection widens the window
- Ordering between siblings was never promised
- Freeze it, and pass values down instead
basics
~20 sSibling fields may resolve on different threads, so any mutable object in the per-request context is written concurrently: a plain list or map, a lazily built client, a scratch field one resolver sets for another. Lost updates and stale reads follow.
solid answer
~50 s"Request-scoped" bounds the *lifetime* of that state, not the number of things touching it. One request can have dozens of resolvers in flight at once, so a per-request object is shared mutable state. Four hazards recur: a plain list or map used as an accumulator, where concurrent appends lose entries or corrupt the structure; a scratch field one resolver sets and another reads, which is a race on ordering the executor never promised; thread-local values — a principal, a tracing scope, a correlation id — that vanish when work resumes on another thread; and lazy initialization that runs twice. It is load-dependent because the window is small and widens with the selection: a 37-field type selected in full gives 37 overlapping writers. Fix it by freezing the context after construction, accumulating per field and joining at the end, and passing values down explicitly rather than sideways.
code
pseudocode · 7 lines# UNSAFE: three hazards in five lines
function resolveStringField(parent, args, ctx):
ctx.auditTrail.append(fieldName) # plain list, many writers
if ctx.meterClient == null:
ctx.meterClient = buildClient() # lazy init runs twice
ctx.currentStringId = parent.id # a sibling may read this
return ctx.meterClient.read(parent.id)go deeper
Know the headline risk: the object handed to every resolver in a request can be touched by several of them at once, so writing to it during execution is dangerous even though it belongs to one request.
Explain the concrete hazards — a non-thread-safe accumulator, a scratch field read by a sibling, a thread-local lost on an async hop, a doubled lazy init — and why a wider selection makes each of them more likely.
Demonstrate the diagnosis: force serial resolution to confirm timing is the variable, widen rather than narrow the reproducing document, instrument the shared object, then remove the sharing instead of guarding it with a lock.
Set the rule that prevents the whole class: contexts are built before execution and read-only during it, values flow parent-to-child, and any deliberately concurrent shared component is documented as such and reviewed on those terms.
## Why the per-request context is the object that breaks Every resolver in a request is handed the same context object, and that is the point of it — it is where the request's authenticated principal, the correlation id, the tenant, the configured downstream clients and the request's batch loaders live. It is also, therefore, the one object in the system that dozens of concurrently executing resolvers all hold a reference to. The phrase "request-scoped" does a lot of quiet damage here. It sounds like isolation, and it is — *between requests*. It says nothing at all about how many resolvers inside one request touch it at once. On a wide document, that number is large: a `SolarString` type with 37 fields, all selected, produces up to 37 resolvers that may run at the same instant, each holding the same context. ## The four hazards, in the order they actually show up **1. A mutable collection used as an accumulator.** Someone adds an audit trail: each resolver appends the field it served to a list on the context, and a plugin writes the list out at the end. With a plain, non-thread-safe list this loses entries under concurrent appends, and in some runtimes corrupts the backing array or its size counter outright — the classic symptom being entries that are null, duplicated, or a length that does not match what was appended. **2. A scratch field used as a channel between siblings.** "The `entitlements` resolver sets `ctx.allowedSiteIds`, and the `readings` resolver reads it." This works in the two-field document the developer tested and fails the moment both fields are in flight together, because the executor never promised that one runs first. The failure is not a crash; it is a resolver reading `null` or reading last request's shape of the data, and returning a plausible wrong answer. That is the worst class of bug in this family, because nothing logs. **3. State attached to the current thread.** A principal in a thread-local, a tracing scope, a logging context: all of them assume the code that reads is running on the thread that wrote. A resolver may begin on one thread and resume on another after awaiting, and a child field may be dispatched to a worker entirely. The value is then simply absent — or, worse on a pooled thread, is the *previous* task's value, which is a cross-request leak rather than a null. **4. Lazy initialization.** `if (ctx.client == null) ctx.client = build()` runs twice under two concurrent resolvers, and now two of your fields are using two different clients, one of which nobody will ever close. There is also a subtler variant of all four: even without a torn structure, a write by one thread is not guaranteed to be *visible* to another without a proper happens-before relationship. Code that looks correct because the interleaving is benign can still read a stale value. ## Why it is load-shaped Under light load a wide document may still resolve mostly sequentially — resolvers whose data is already in memory return before the next one starts — so the window where two writers overlap is tiny. Add real backend latency, real concurrency, and a client that selects everything, and the window opens on every request. A useful worked shape: at 180 requests per second against a document selecting all 37 fields on each of 214 strings, an audit list that drops entries on roughly 0.4% of requests is invisible in a dashboard and completely reproducible in aggregate. ## Confirming it rather than guessing The cheapest discriminator is to **force serial execution** — most servers can be configured or wrapped to resolve fields one at a time — and replay the same document. If the symptom disappears while the data and the code stay identical, it is a race, not a data bug. From there: replay with the widest document you can construct rather than the narrow one from the bug report, since width is the variable; wrap the shared object in a checking decorator that records the thread of every mutation and screams when it sees a second one; and grep for every write to the context after construction, because the population of writers is small and enumerable. ## The fixes, cheapest first **Freeze the context.** Build it fully during request setup, then treat it as read-only. This removes hazards 2 and 4 by construction, and it is almost always achievable — what resolvers actually need from the context is configuration, identity and loaders, all of which are known before execution starts. **Accumulate per field, join at the end.** If you genuinely need to collect something across fields, either use a collection built for concurrent writers or have each resolver produce its own value that a single collector merges once, after execution. **Pass values down, never sideways.** A value a child field needs should come from the parent's *resolved value* — the executor already carries that to the child, in the right order, with no sharing. This is the structural answer to hazard 2 and the one to say out loud in an interview. **Propagate context explicitly.** Prefer passing the context object through calls over relying on the thread to carry it; where the runtime offers a context-propagation mechanism across async hops, use it deliberately rather than assuming it. One shared object is a deliberate exception: the request's batch loaders are designed to be called from many fields at once, and their coordination and cache semantics are their own subject. That is a designed concurrent component, not an accumulator someone added to the context on a Friday.
- Making every access to the shared context synchronized would fix it — why is that a poor answer?Because it trades a correctness bug for a throughput bug. A single lock around the object every resolver touches serializes the concurrency the executor was created to exploit, so a wide document's latency collapses toward the sum of its fields rather than the maximum. It also leaves the ordering hazard untouched: locking makes each read atomic, not correctly timed. Removing the sharing is strictly better than guarding it.
- How would you prove a suspected race rather than infer it from the symptom?Replay the exact document with the executor forced to resolve fields one at a time. Identical inputs and code, symptom gone, means timing is the variable. Then widen the document rather than narrowing it, and instrument the shared object with a decorator that records the mutating thread — a second thread appearing is the confirmation, and it names the offending field for you.
- Which context values are safe to share across concurrently resolving fields?Anything immutable after construction — identity, tenant, correlation id, configuration — and any component explicitly designed for concurrent callers, such as a downstream client that is documented as thread-safe or the request's batch loaders. The unsafe residue is always the same shape: something a resolver writes during execution that another resolver can observe.
A per-request context is a shared whiteboard, not a private notebook: it is wiped between customers, but during one order every clerk is writing on it at once.
saying these in an interview costs you the question
- Assumes all resolvers in one request share a thread
- Uses the context to pass values between sibling fields
- Says request-scoped means nothing can race
- Relies on thread-locals surviving an async hop
- Adds a retry and calls it a flaky test
- Locks the whole context and ignores the latency cost