skip to content

When a GraphQL operation deadline fires mid-execution, what happens to the in-flight resolvers?

level: middleimportance: must knowfreq 52%

answer

  1. A timer is not a stop button
  2. Three layers cancel independently
  3. Cancellation must be cooperative all the way down
  4. Backend work can outlive the response
  5. Partial data, null field, errors entry

basics

~10 s

Nothing, unless cancellation is propagated. A deadline is a timer, not a stop button: the server can abandon the response while resolvers, driver calls and backend statements keep running and keep holding connections.

solid answer

~50 s

A deadline expiring only means a timer fired. Whether in-flight work stops depends on whether cancellation is threaded from the operation context into every resolver and onward into the client libraries and backend statements they invoked. If it is not, the classic failure is orphaned work: the response goes out, and a database statement keeps running for another nine seconds holding a pooled connection — which is worse than no timeout at all, because the caller retries and multiplies the load. On the response side, the GraphQL specification defines no timeout behaviour, so the server chooses. The common shape is partial `data` with the unfinished fields null, plus entries in `errors` carrying the field `path` and a server-chosen code. If a cut-short field is non-null, that null bubbles up to the nearest nullable parent and can empty a whole branch.

code

pseudocode · 6 lines
pseudocode
function resolveBinStock(parent, args, context):
    if context.deadline.expired():
        raise Cancelled("operation deadline exceeded")
    remaining = context.deadline.remainingMillis()
    return inventoryDb.query(SELECT_BIN_STOCK, args.binId,
                             timeoutMillis = min(remaining, 120))

go deeper

for a junior

Be ready to say that a deadline only means a timer fired, and that the resolvers and backend calls keep going unless something explicitly cancels them.

for a middle

Explain the three layers that cancel independently, how an absolute deadline in the per-request context reaches concurrent resolvers, and what the response looks like when some fields finished and others did not.

for a senior

Demonstrate the amplification story: unpropagated cancellation plus client retries drains a connection pool and reads as a database outage. Show how you would measure orphaned work and set the retry policy alongside the deadline.

for a principal

Own the standard: which layers must accept a deadline, what error shape all services return, and whether partial responses or whole-response failures are the house rule so clients can be written once against a consistent contract.

## A deadline is a timer, not a stop button The single most common misconception here is that "the operation timed out" means the work stopped. It does not. A deadline expiring is an event on a clock. What happens next depends entirely on whether the server was built to *act* on that event, and on whether every layer beneath the resolver cooperates. There are three layers of work behind a single field, and they cancel independently: 1. **The resolver function itself** — your code, waiting on something. 2. **The client library call it made** — an HTTP client, a database driver, a cache client. 3. **The work the remote system is doing** — the SQL statement executing, the request the downstream service is serving. Abandoning the outermost future stops only the first. Layer two stops if the driver supports cancellation and you passed it a timeout or a cancellation token. Layer three stops only if the protocol has a way to say "abort this" and the driver actually sends it — many do, some do not, and a fire-and-forget HTTP call to a downstream service almost never does. ## Threading the deadline through execution GraphQL execution makes this harder than in a plain request handler, because sibling fields in a selection set may be resolved concurrently. A deadline that lives only in the outermost handler cannot reach thirty resolvers running in parallel. The working pattern is to put an absolute deadline — a point in time, not a duration — into the per-request context that execution already passes to every resolver, and to have each resolver do two things: short-circuit immediately if the deadline has already passed, and pass the *remaining* budget down into whatever call it makes. ``` function resolveBinStock(parent, args, context): if context.deadline.expired(): raise Cancelled("operation deadline exceeded") remaining = context.deadline.remainingMillis() return inventoryDb.query(SELECT_BIN_STOCK, args.binId, timeoutMillis = min(remaining, 120)) ``` The cheap entry check matters more than it looks. Once the deadline has passed, every field the execution engine has not yet started is pure waste; refusing them instantly keeps a doomed operation from launching another two hundred backend calls on its way out. ## Orphaned work, and why a bad timeout is worse than none Suppose a warehouse inventory service holds a 340 ms p99 budget and sets a 2-second operation deadline. A valuation query goes slow; at 2 seconds the server gives up and answers. If cancellation did not reach the database, that statement keeps running — perhaps another nine seconds — and keeps its pooled connection checked out the whole time. Now add the caller's retry. The client sees a failure and re-sends. The server starts a second copy of the same expensive work while the first is still running. Under sustained load the pool drains, unrelated operations queue behind it, and the incident looks like a database outage rather than a timeout misconfiguration. This is precisely the failure that only appears in production, where the data volume makes the query slow in the first place. The lesson is blunt: a deadline you do not propagate converts a slow endpoint into an amplifying one. ## What the client sees The GraphQL specification says nothing about deadlines, so there is no specified error, no reserved code, and no required shape. Servers choose between two behaviours. **Partial response.** Keep whatever resolved in time, put `null` where a field was cut short, and add entries to `errors` with the field's `path`. The caller gets something useful and can see exactly which branch failed. ```json { "data": { "warehouse": { "name": "Zaragoza DC-3", "stockValuation": null } }, "errors": [ { "message": "Operation deadline exceeded", "path": ["warehouse", "stockValuation"], "extensions": { "code": "DEADLINE_EXCEEDED", "correlationId": "7f3ac1" } } ] } ``` The `code` there is a server's own choice, not a specified value. **Whole-response failure.** Abandon the operation and return a single error with no data. Simpler to reason about, and appropriate when a partial answer would mislead, but it throws away work that already succeeded. One consequence to be ready for: if the cut-short field was declared non-null, the null cannot stay where it is. It propagates up to the nearest nullable ancestor, which can wipe out a large, otherwise-complete branch of the response. A schema with aggressive non-null annotations turns one slow field into a mostly-empty answer. ## Mutations and subscriptions Root mutation fields execute serially, so a deadline that fires in the middle of a multi-field mutation can leave the first fields applied and the rest not. A deadline is not a rollback; if atomicity matters, it has to come from a transaction, not from the timer. For subscriptions, a per-operation deadline is the wrong tool — the stream is meant to outlive any single budget. What a deadline can sensibly bound there is the execution of each individual delivered payload; how long the stream itself may live is governed by separate connection and authorization limits.

  • Why can a timeout that does not propagate be worse than having no timeout at all?
    Because it hides the cost while multiplying it. The caller sees a fast failure and retries, so a second copy of the same expensive work starts while the first still runs and still holds its pooled connection. Load compounds until the pool drains and unrelated operations queue behind it. Without the timeout you would at least have felt the true latency and stopped sending more.
  • How does a schema's nullability affect what a client receives when a deadline cuts a field short?
    A cut-short nullable field is simply null with an entry in errors, and the rest of the response survives. If the field is non-null, the null cannot be placed there and propagates up to the nearest nullable ancestor, so one slow field can erase a large, otherwise-complete branch. Aggressive non-null annotations therefore make partial responses far less useful under timeouts.
  • What is a sensible way to observe whether deadlines are firing in production?
    Count deadline hits per operation name, not just in aggregate, and record where in the field path they occur. Also track the gap between operation end and backend statement end, which is the direct measure of orphaned work. A rising deadline rate on one operation name is a regression signal long before it becomes an availability incident.

Hanging up on a call does not make the person on the other end stop talking; something has to tell them the conversation is over.

saying these in an interview costs you the question

  • Assumes the deadline firing automatically stops the database query
  • Says returning the response cancels the backend work
  • Thinks the specification defines a timeout error code
  • Adds client retries without bounding the work already in flight
  • Believes a deadline rolls back a half-applied mutation
  • Puts the deadline in the handler only, never in resolver context

context