skip to content

How do you bound and isolate the concurrency that one GraphQL document's resolvers can create?

level: principalimportance: should knowfreq 40%

answer

  1. The multiplier has no built-in ceiling
  2. CPU is rarely the resource that runs out
  3. Bound it where the scarcity actually is
  4. Lanes per dependency, not one shared pool
  5. Blocking on same-pool work deadlocks

basics

~20 s

Put the bound where the scarce resource is: a cap on in-flight calls per downstream dependency, separate pools so one slow source cannot starve the others, and admission control on requests — not one global pool sized by guesswork.

solid answer

~50 s

Concurrent field resolution multiplies: a 37-field type selected in full across 214 items is nearly eight thousand potentially simultaneous calls from one document. CPU is rarely what saturates — connection pools, downstream capacity and thread count are. So decide **where the ceiling lives**. A single shared pool for all resolvers couples every dependency, and if a resolver blocks on work queued behind it in that same pool you get exhaustion and deadlock rather than slowness. Prefer bulkheads: a bounded permit count per dependency, so a degraded source fails fast in its own lane instead of consuming every thread; admission control at the edge to bound concurrent requests; and deadline propagation so abandoned work stops rather than piling up. Size each lane against the *weakest* dependency's real capacity, measure in-flight and queue depth per dependency, and accept the tradeoff: more overlap improves one request's latency and converts a spike into downstream overload.

code

pseudocode · 14 lines
pseudocode
lanes = {
    "meter-store":   semaphore(permits = 48),   # sized to its connection pool
    "weather-api":   semaphore(permits = 12),   # third party, strict rate limit
    "alarm-service": semaphore(permits = 30)
}

function callDependency(name, deadline, work):
    if not lanes[name].tryAcquire(timeout = 40ms):
        raise LaneSaturated(name)          # fail fast in one lane only
    try:
        if now() > deadline: raise Expired(name)
        return work(remaining(deadline))
    finally:
        lanes[name].release()

go deeper

for a junior

Take away the mechanism: one document can start a very large number of backend calls at once, and somebody has to decide the limit — it is not something the protocol or the executor provides for free.

for a middle

Be able to compute the fan-out from a document's shape and name what runs out first: connection pools and downstream capacity long before CPU. Know why blocking on work queued in the same pool can deadlock.

for a senior

Show the operational instinct: isolate dependencies into separate bounded lanes, propagate deadlines so abandoned work stops, and instrument per-lane in-flight and wait time rather than reading aggregate latency.

for a principal

Own the policy. Say where the ceiling lives, that it is set against the weakest dependency's measured capacity rather than a latency target, that different lanes get different numbers, and that a predictable low-concurrency mode is a defensible choice for fragile backends.

## Where the concurrency comes from, quantitatively The executor is permitted to resolve the fields of a selection set at the same time, and every list multiplies that permission. Take a solar-array telemetry graph: a site has 214 strings, each string's type declares 37 fields, and a monitoring client selects all of them. If a meaningful fraction of those fields calls a backend, one HTTP request has authorised on the order of seven or eight thousand concurrent operations. Nothing in the protocol bounds that number, and nothing in the executor does either — the permission to overlap comes with no ceiling attached. This is what makes it a leadership question rather than an implementation detail. The ceiling is a choice somebody makes on purpose, and if nobody makes it, the answer is "whatever the runtime happens to allow", discovered during an incident. ## What actually saturates Almost never CPU. The resources that run out, roughly in the order they do: * **Connections to a downstream** — a pool of, say, 40 database connections is the real ceiling on concurrent reads, no matter how many resolvers are ready to run. * **The downstream's own capacity** — you can be perfectly healthy while making a dependency unhealthy, which is worse, because the failure is somebody else's dashboard. * **Threads or task slots**, if resolvers block rather than yielding. * **Memory**, when thousands of in-flight operations each retain a buffer. So "tune the resolver pool" is usually the wrong frame. The pool is not the resource; it is one gate in front of several different resources with different capacities. ## The blocking-pool deadlock, which is worth naming The sharpest failure mode in this area: a resolver runs on a pool thread, submits dependent work to *the same* fixed pool, and blocks waiting for it. Under enough concurrency every thread in the pool is occupied by a resolver waiting for a task that is sitting in that pool's queue behind it. Throughput does not degrade — it goes to zero, and the service looks alive while serving nothing. Any design where a task can wait on another task in the same bounded pool has this property; the fixes are to never block on same-pool work, to give nested work its own pool, or to make the resolvers non-blocking. ## Four places the bound can live, and what each one buys **Admission control at the edge.** Cap concurrent in-flight requests. Coarse, cheap, and the only control that protects you from "many small documents" as well as "one huge document". It bounds the multiplier's outermost factor. **A per-request ceiling.** A permit count each request's resolvers draw from, so no single document can occupy the whole server. Preserves fairness between a monitoring client's wide document and everyone else's small ones. The cost is that a legitimately wide request now takes longer, which is usually the correct trade. **Per-dependency bulkheads.** A separate bounded lane for each backing source. This is the highest-value control, because the failure it prevents is the expensive one: a single degraded dependency absorbing every unit of concurrency in the process and taking down fields that had nothing to do with it. When its lane is full, calls to that source queue briefly and then fail fast, and the rest of the graph keeps serving. **The downstream's own limits.** Connection pools and server-side rate limits are a real bound, but they are a backstop, not a plan — reaching them usually means queueing somewhere you cannot see. A related control is coalescing repeated fetches so the concurrency never gets created in the first place; that is the batching pattern's job and its own subject. Static limits on how expensive a document may be before execution starts are a third, different lever that belongs with abuse controls. Concurrency bounding is what remains once you have taken both: the load a *legitimate, already-admitted* document is allowed to generate. ## Deadlines and cancellation A bound that only queues work is half a design. If a client has already given up, or the request's deadline has passed, every still-queued call is pure waste that is also holding a permit. Propagate a deadline into resolvers and their downstream calls, and cancel outstanding sibling work when the response can no longer be delivered. This is also where you discover whether your executor abandons in-flight resolvers when a field error nulls their parent — some do, some do not, and "some do not" means abandoned work keeps consuming the very capacity you are trying to protect. ## Measuring the right things Request latency alone will not tell you where the ceiling is. Instrument **per-dependency in-flight count and queue wait**, **permit acquisition time** for each bulkhead, and **fan-out per request** — how many downstream calls one document produced. That last metric is the one that turns "the graph got slow" into "this one client's document changed shape last Tuesday". Alert on saturation of a lane, not on the aggregate. ## The tradeoff to own out loud More concurrency lowers the latency of a single request and raises the peak load a burst of requests puts on shared dependencies. Those pull in opposite directions, and there is no number that is right for both. The honest position is that the ceiling is set against the weakest dependency's measured capacity, not against a latency target; that different lanes get different ceilings because the sources have different capacities; and that a mostly-sequential mode is a legitimate, predictable configuration for a graph whose backends are fragile. The failure to avoid is the one that looks like tuning: raising the pool size until the symptom moves, which usually just relocates the queue to the database.

  • Why does raising the resolver thread pool size often fail to fix a slow, wide GraphQL document?
    Because the pool is rarely the binding constraint. More threads simply deliver more concurrent calls to a dependency whose connection pool or capacity has not changed, so the queue moves from your process to theirs — where it is harder to see and where it now degrades every other consumer of that source. The fix is to find which lane is saturated and size against it.
  • How would you decide the permit count for one dependency's lane?
    From measurement, not arithmetic. Start from that source's known concurrency limit — connection pool, published rate limit, or a load test to its knee — and set the lane below it, leaving headroom for other consumers of the same source. Then watch permit-acquisition time and the dependency's own latency curve: rising wait with flat downstream latency means the lane is too small, rising downstream latency means it is too large.
  • When is deliberately reducing concurrency the right call for a graph?
    When the backends are fragile or shared and predictability matters more than the latency of the widest document. A near-sequential mode gives a load profile you can reason about and a tail that does not depend on which client shows up. It is a poor default for an interactive graph, but a reasonable one for a reporting or batch-facing endpoint.
  • What single metric best predicts this class of incident before it happens?
    Fan-out per request — the number of downstream calls one document produced — segmented by client and operation. It turns a diffuse "the graph got slower" into a specific document whose shape changed, and it is the only measure that connects a client-side selection change to a backend saturation event. Pair it with per-lane saturation to see where the fan-out lands.

One shared pool is a single doorway for every errand in the building; bulkheads are separate doors per destination, so the queue for the slow one does not block everybody else.

saying these in an interview costs you the question

  • Tunes one global thread pool as the whole answer
  • Assumes CPU is what limits resolver concurrency
  • Blocks on work queued in the same bounded pool
  • Bounds requests but never bounds per-dependency calls
  • Adds capacity limits with no deadline or cancellation
  • Measures only request latency, never per-lane saturation

context