skip to content

What does a per-operation timeout on a GraphQL server bound that a cost limit cannot?

level: juniorimportance: should knowfreq 42%

answer

  1. Two different questions about one operation
  2. Shape is not duration
  3. A static score sees no clock
  4. Same document, different data, different runtime
  5. Wall-clock budget enforced during execution

basics

~20 s

A cost limit scores the document's shape before anything runs, so it never sees a clock. A per-operation deadline bounds wall-clock time — the slow backend call, the lock wait, the cold cache — which no static score can predict.

solid answer

~50 s

Depth caps, field-count caps and cost scores are all computed from the document text and the schema **before execution**. They answer "how big is this request?" — how deeply it nests, how many fields it selects, how many list elements it could multiply out. None of them answers "how long will this take?", because duration depends on data the limiter has never seen: how many rows sit behind the field, whether a cache is warm, whether a downstream service is degraded. A two-field document scoring two points can run for eleven seconds. A per-operation deadline is a wall-clock budget attached to the operation and enforced *during* execution: when it expires the server stops waiting and answers. Neither control is in the GraphQL specification — both are server conventions — and they are complements, not alternatives: shape limits reject the pathological document, the deadline bounds the ordinary one that turns out to be slow.

code

graphql · 6 lines
graphql
query WarehouseValuation($warehouseId: ID!) {
  warehouse(id: $warehouseId) {
    name
    stockValuation
  }
}

go deeper

for a junior

Be ready to say in one sentence that cost and depth limits are computed from the document before execution, so they bound size, not time, and that a timeout is what bounds time.

for a middle

Explain the three reasons a cheap-scoring document runs slowly — data volume, variable values and environment — and describe how transport, operation and per-call timeouts nest so the inner budget is always smaller.

for a senior

Show how you would choose the numbers: measure per-operation p99, set the deadline above legitimate traffic, alert on the deadline-hit rate, and explain why a staging dataset hides the failures that only appear in production.

for a principal

Own the argument that a latency budget is an API contract. Decide which operations get their own budgets, what the organisation promises callers, and how deadline breaches feed capacity planning rather than being silently retried.

## Two different questions about one operation Every abuse control on a GraphQL endpoint answers exactly one of two questions, and confusing them is the mistake this topic exists to correct. The first question is **how big is this request?** Depth caps, field-count caps, body-size caps and complexity scores all answer it. They are computed from two artefacts the server holds before it runs anything: the executable document the client sent, and the schema. Walk the selection set, add up per-field weights, multiply where a list-slice argument such as `first` tells you a subtree repeats, compare the total against a budget, reject if it is over. This is *static analysis of a document*. It is fast, deterministic, and it happens before a single resolver is called. The second question is **how long did this request take?** Nothing in the first category can answer it, because the answer does not live in the document. It lives in the data, in the caches, and in whatever the downstream systems are doing at that moment. ## Why a small document can be slow Consider a warehouse inventory graph with a valuation field: ```graphql query WarehouseValuation($warehouseId: ID!) { warehouse(id: $warehouseId) { name stockValuation } } ``` Two fields, one nesting level. Any depth cap passes it. A cost limiter that weights object fields at 1 scores it at 2 or 3. And `stockValuation` may sum across 8,700 bins, join a pricing table, and take four seconds. The score was not *wrong* — the document really is small. It simply measured the wrong dimension. Three independent sources of that gap: **Data volume the document does not mention.** The same document run against a warehouse holding 12 bins and one holding 8,700 has the same score and wildly different runtimes. **Variables.** Static scoring uses the variable *values* only when they are literal slice arguments the limiter knows how to read. An identifier, a date range, a filter object — all of these change the work enormously and none of them change the score. **Environment.** This is why the classic symptom is a query that times out only in production. A staging dataset seeded with a few dozen rows makes every resolver instant; production has the real cardinality, a cold cache after a deploy, a lock held by a batch job, or a downstream service answering at three times its usual latency. The document did not change. The clock did. ## What a deadline actually is A per-operation deadline is a wall-clock budget attached to the operation at the moment it starts executing, carried through the execution as part of the per-request context, and honoured by the code that waits on anything. When it expires the server stops waiting for the outstanding work and produces whatever answer it can. It is worth being precise that this is **not in the GraphQL specification**. The specification defines parsing, validation and an execution algorithm; it has no notion of elapsed time, no timeout error, no cancellation. Deadlines, like cost limits, are entirely a server-side convention. That matters in an interview because candidates routinely say "the spec requires" about controls the spec has never mentioned. ## Deadlines come in layers A real service usually has three, and they must nest: * A **transport-level timeout** on the HTTP request or the socket, the outermost backstop. * A **per-operation deadline** inside execution, the one this topic is about, which knows it is running a GraphQL operation and can therefore produce a GraphQL-shaped answer. * A **per-call timeout** on each downstream call a resolver makes, so one slow dependency cannot consume the whole operation budget. The inner budgets must be strictly smaller than the outer ones. If a downstream call is allowed 5 seconds while the operation deadline is 2 seconds, the operation deadline fires first and the downstream timeout never does anything useful. A sensible starting point is to derive the deadline from the latency budget you have already committed to. If a page has a 340 ms p99 budget, an operation still running at two or three seconds is not going to produce a useful answer for anyone — the user has navigated away, the caller's own client has probably given up. The deadline is not the target; it is the point past which the answer is worthless and holding resources for it is pure loss. ## What a deadline does not do It does not stop the backend work by itself — that requires cancellation to propagate into the resolvers and their drivers, which is a separate problem. It does not stop a caller from sending the same expensive operation a thousand times; that is what rate limiting and per-caller cost budgets are for. And it is the wrong tool for a subscription, which is *designed* to run indefinitely: a subscription needs limits on how many streams exist and how long an authorization lasts, not a deadline on the stream itself.

  • Does the GraphQL specification say anything about operation timeouts?
    No. The specification covers document syntax, validation rules and an execution algorithm that produces a response; it has no concept of elapsed time, no timeout error and no cancellation. Deadlines are purely a server-side convention, as are depth caps and cost limits. Be careful never to attribute any of these controls to the specification.
  • If you already time out each downstream call, why also set a deadline on the whole operation?
    Because an operation makes many calls. Twenty resolvers each honouring a 500 ms per-call timeout can still total ten seconds, and a document that fans out further multiplies that again. The per-operation deadline bounds the sum; the per-call timeout stops any single dependency from eating the whole budget. You want both, with the per-call value strictly smaller.
  • How would you pick the value of a per-operation deadline?
    Work back from the latency budget the caller has already committed to, and from measured p99 execution time per operation rather than a single global guess. Set the deadline above your legitimate p99 with headroom, so it trims the pathological tail rather than failing normal traffic, and alert on the rate of deadline hits so a regression shows up as errors you notice.

A parcel's dimensions tell you whether it fits through the door; they tell you nothing about how long the delivery will take once traffic, the lift and the recipient are involved.

saying these in an interview costs you the question

  • Says a cost limit already bounds how long a query runs
  • Claims the GraphQL specification defines an operation timeout
  • Thinks a shallow document is necessarily a fast one
  • Sets the downstream call timeout larger than the operation deadline
  • Treats the deadline as the latency target rather than a cutoff
  • Applies a per-operation deadline to a subscription stream

context