skip to content

Subgraph Failure & Partial Data

What a client receives when one subgraph errors or times out mid-plan: partial data, an error entry carrying the failed path, and the non-null bubbling that can erase a whole branch.

part ofGraphQLoverview, primer and where to startread it →
on this pageshow

questions

4

In a federated graph, what does a client receive when one subgraph fetch fails?

level: juniorimportance: must knowfreq 57%

answer

  1. One response, not an outage
  2. The healthy services already answered
  3. A hole, plus an entry explaining it
  4. path walks to the missing branch
  5. Nullability decides how big the hole is

basics

~20 s

One ordinary response carrying both halves: data holding everything the healthy subgraphs returned, and an errors entry whose path names the field the failed fetch should have filled. Partial data is the normal outcome, not an outage.

solid answer

~50 s

A federation router turns one client document into several subgraph fetches. If one of them fails or times out, the others have already answered, so the router returns what it has: `data` with the failed branch set to `null`, plus an entry in `errors` whose `path` locates that branch. The transport status is still a success — nothing stopped the request from executing, so this is a field-level failure, not a request error, and `data` is present rather than absent. How much disappears is decided by nullability: a nullable field leaves a single hole, a Non-Null one erases its parent chain. Worth being precise about ownership: neither the GraphQL specification nor Apollo Federation's composition specification says what a router does with a broken fetch — partial data plus an error is convergent implementation behaviour, and it is what clients are built to expect.

code

graphql · 8 lines
graphql
query CaseOverview {
  case(id: "CASE-2019-4471") {
    caseNumber
    court { name }
    parties { role displayName }
    billing { unbilledHours lastInvoicedOn }
  }
}

go deeper

for a junior

Recall the shape of the response: one body, a data key still present, the failed branch null, and an errors entry naming that branch. Be able to say that partial data is the expected outcome and not a bug.

for a middle

Explain why it is partial: a broken fetch is a field error, not a request error, so execution had already begun and data survives. Know that nullability, not the router, decides how much of the response disappears.

for a senior

Show what this does to operations: a subgraph outage arrives as a 200, so availability has to be measured on response completeness and error paths rather than status codes, and null-with-no-error must be kept distinct from null-with-error.

for a principal

Own the contract question — whether partial responses are a first-class part of the graph's client contract, what each team may put in an error entry, and how completeness rather than uptime becomes the number the organisation is judged on.

## The setting A federated graph is one client-facing schema composed from many independently deployed subgraph schemas. A **router** parses the client's document, works out which service owns which field, and issues its own requests — **subgraph fetches** — to those services, splicing the results back into one response shape. Take an 11-service legal case-file graph. A **Cases** subgraph owns `Case` and its `caseNumber`, `title` and `court`. **Parties** contributes `Case.parties`. **Billing** contributes `Case.billing`. Eight more services cover documents, deadlines, conflicts, audit and so on. A caseworker opens one screen: ```graphql query CaseOverview { case(id: "CASE-2019-4471") { caseNumber court { name } parties { role displayName } billing { unbilledHours lastInvoicedOn } } } ``` The plan is three fetches: Cases first, then Parties and Billing in parallel against the case key. Billing is having a bad afternoon and its fetch expires after 1,850 ms. ## What the client actually gets A response that arrives half-empty — one body, one status, two halves: ```json { "data": { "case": { "caseNumber": "2019/4471", "court": { "name": "Central District" }, "parties": [{ "role": "PLAINTIFF", "displayName": "Adeyemi Holdings" }], "billing": null } }, "errors": [ { "message": "Subgraph request failed", "path": ["case", "billing"] } ] } ``` Three things to read off it. The `data` key is **present** — execution began and mostly succeeded. The failed branch is `null`. And the `errors` list carries an entry whose `path` walks from the response root to the exact field that could not be filled, which is how a client tells *which* of eleven services let it down without knowing the topology. ## Why it is partial rather than total GraphQL distinguishes two failure classes. A **request error** — malformed syntax, a validation failure, a variable that cannot be coerced — means nothing executed; the response carries no `data` key at all, and over HTTP that is the case that earns a 4xx. A **field error** happens after execution started, when one field could not produce a value. A dead subgraph is the second kind: the document was valid, the plan ran, most of it worked. So the status stays a success and the body carries both halves. A client that treats any non-empty `errors` list as a total failure is throwing away a screen it could have drawn. ## Nullability decides the size of the hole The router cannot invent a `Billing` object; GraphQL has no default value for an output field. So the field becomes null, and the schema decides how far that null travels. `billing: Billing` leaves exactly the hole above. `billing: Billing!` cannot hold null, so the null climbs to the nearest nullable ancestor — possibly erasing `case` entirely, including the `caseNumber` and `parties` the healthy subgraphs already returned. Same outage, wildly different response, and the difference was written in the SDL months earlier. ## The silent variant Not every empty branch comes with an error. When the router resolves an entity by key, the subgraph returns one slot per key, and a slot may legitimately be null — an **unresolvable reference**: the service is healthy and simply has no such record. The client then sees `billing: null` with **no** entry in `errors`. That asymmetry is the practical diagnostic: a hole *with* an error entry is a failure, a hole *without* one is an absence. Any monitoring or client logic that reads "null means broken" will be wrong in one direction or the other. ## Who specifies this Be careful with attribution in an interview. The GraphQL specification defines the `data` / `errors` envelope, field errors and null propagation — it knows nothing about routers or subgraphs. Apollo Federation is a **composition** specification: it defines directives, entity keys and the reserved fields a subgraph exposes, and it does not define what a router does when a fetch breaks. Partial data plus a `path`-carrying error is a strong convention that mainstream routers converge on and that clients depend on — not a specified rule. Saying "the spec requires it" is the kind of confident wrong answer interviewers remember. ## What this changes downstream Because a service outage arrives as a normal 200 with a hole in it, ordinary HTTP monitoring sees nothing. Availability for a federated graph has to be measured on the *completeness* of responses — error entries per operation, grouped by the `path` prefix that identifies the failing branch — not on status codes. And clients must render per branch: draw the case, draw the parties, show a small unavailable marker where billing should be.

  • Does the caller get a failing HTTP status when one subgraph fetch dies?
    Normally no. The document parsed and validated and execution began, so this is a field error, and the GraphQL over HTTP working draft ties a failing status to the request being unexecutable rather than to a field going wrong. The response is a success status carrying `data` and `errors` together. That is exactly why status-code monitoring in front of a federated graph will happily report 100% availability through a subgraph outage.
  • Is a null branch always accompanied by an error entry?
    No, and the difference matters. When the router resolves an entity by key, a healthy subgraph can return a null slot for a key it simply has no record for — an unresolvable reference. The client then sees a null field with an empty `errors` list. A hole with an error is a failure; a hole without one is an absence. Client code and alerting both need to keep those apart.
  • How does a client work out which service failed from the response alone?
    It should not have to, and by design it cannot. The `path` names a field in the client-facing schema, not a service; the topology stays behind the router. Operationally, teams map `path` prefixes to owning subgraphs on the server side and key the detail by a correlation id. Putting the failing service's name or host into the client's error message hands out an internal map for free.

A case bundle handed over with one folder missing and a slip in its place naming what is absent — you still get the other ten folders, and you know exactly which one to chase.

saying these in an interview costs you the question

  • Says the whole request fails when any subgraph is down
  • Expects a 500 status for a failed subgraph fetch
  • Assumes data and errors cannot both be present
  • Cannot say which field the error's path names
  • Claims the GraphQL specification defines router failure behaviour
  • Reads every null branch as evidence of an outage

context

open as a page

In a federated graph, why can one subgraph's failure erase data other subgraphs returned?

level: middleimportance: should knowfreq 51%

basics

~20 s

Because the failed field was declared Non-Null. A null cannot sit there, so it climbs to the nearest nullable ancestor, discarding sibling fields the router had already collected from healthy services. Nullability at ownership seams sets the blast radius.

open as a page

How do you set per-fetch timeouts in a federated router so one slow subgraph cannot stall a request?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Give the whole operation one deadline, derive each fetch's timeout from the budget left rather than a fixed constant, size it from that subgraph's own latency distribution, and propagate the deadline downstream so an abandoned fetch stops working.

open as a page

In a federated graph with many services, which subgraph failures should be allowed to fail a whole request?

level: principalimportance: should knowfreq 38%

basics

~20 s

Only the ones whose absence would make the rest of the response wrong or unsafe to act on. Everything else degrades to a hole. Tier the subgraphs by that test, then encode each tier in nullability, budgets and the client contract.

open as a page