skip to content

In a federated graph, why can one subgraph's failure erase data other subgraphs returned?

level: middleimportance: should knowfreq 51%

answer

  1. Two services, one exclamation mark
  2. The null has nowhere to sit
  3. It climbs until something may be null
  4. Healthy data thrown away on the way up
  5. Reliability of the owner sets the wrapper

basics

~20 s

Because the failed field was declared Non-Null. A null cannot sit there, so it climbs to the nearest nullable ancestor, discarding sibling fields the router had already collected from healthy services. Nullability at ownership seams sets the blast radius.

solid answer

~50 s

The router cannot manufacture a value for a field whose subgraph is down, and GraphQL has no default for an output field. If that field is nullable, the response keeps a single hole. If it is Non-Null, the null has nowhere to live and the failure propagates upward until it reaches a nullable position — which may be a parent object a *different*, perfectly healthy subgraph resolved seconds earlier. The router genuinely holds those bytes and must throw them away. The federation-specific hazard is that the person who wrote `!` on a contributed field sits in one team while the type it hangs off is owned by another, and composition does not warn anyone that they just coupled the parent type's availability to their service's. The practical rule: at every ownership seam, a contributed field should be nullable unless its owner is as reliable as the parent.

code

graphql · 10 lines
graphql
# Conflicts subgraph schema
type Case @key(fields: "id") {
  id: ID! @external
  conflictCheck: ConflictCheck!
}

type ConflictCheck {
  clearedOn: String!
  reviewedBy: String!
}

go deeper

for a junior

Recall that an exclamation mark means the field can never be null in the response, so when its value cannot be fetched the parent object is what disappears instead. Know that this is decided in the schema, not by the router.

for a middle

Explain the interaction concretely: the failed field is Non-Null, the null climbs to the nearest nullable ancestor, and that ancestor may hold sibling fields another service already returned successfully. Be able to walk an error path and say what was erased.

for a senior

Demonstrate the availability arithmetic — every Non-Null contributed field multiplies its owner's reliability into the parent type's — and show where you would put the nullable boundary and how you would catch violations in CI rather than in an incident.

for a principal

Own nullability at ownership seams as a cross-team contract: who may add a Non-Null field to a type they do not own, what review gate enforces it, and how you migrate an already-shipped ! off a flaky service without breaking generated clients.

## Two independent mechanisms meeting The first mechanism is plain GraphQL: a field error nulls its field, and a null in a Non-Null position cannot stand, so it propagates to the nearest nullable ancestor; if every ancestor up to the root is Non-Null, the whole `data` value becomes null. The second is federation: one field's value comes from one process and its sibling's from another, and those processes fail independently. Each mechanism is unremarkable alone. Put them together and you get the response that surprises people: an 11-service legal case-file graph where the least important service takes the most important data down with it. ## The concrete shape **Cases** owns the type: ```graphql type Case @key(fields: "id") { id: ID! caseNumber: String! parties: [Party!]! } ``` Months later the **Conflicts** team — a different team, a different deployment, a different on-call rota — contributes one field to the same type in its own subgraph schema: ```graphql type Case @key(fields: "id") { id: ID! @external conflictCheck: ConflictCheck! } ``` They wrote `ConflictCheck!` for an honest reason: every case really does have a conflict-check record, so the field is never legitimately null. But `!` in GraphQL is not a statement about the data. It is a promise about the *response*, and the only way to keep it when the value cannot be fetched is to delete the parent. Conflicts times out. The client asked for `caseNumber`, `parties` and `conflictCheck`. Cases and Parties both answered in under 200 ms. The router still has to return: ```json { "data": { "case": null }, "errors": [ { "message": "Subgraph request failed", "path": ["case", "conflictCheck"] } ] } ``` A response that arrives half-empty is the good outcome. This one arrives empty, and the router was holding the caseNumber the whole time. ## Why the router cannot paper over it Three escape routes people propose, and why none exists. There is no default value for an output field — GraphQL defines defaults only for input arguments and variables, so the router has nothing to substitute. It cannot demote `ConflictCheck!` to nullable at runtime, because the client-facing schema is the contract every generated client type was built from; a null there would break callers that were told it could not happen. And it cannot omit the key, because a Non-Null field that was selected must appear with a value. ## Lists make it sharper `deadlines: [Deadline!]!` on `Case`, contributed by a **Deadlines** subgraph, has two `!` marks doing different jobs. The inner one says no element may be null: one bad element nulls the whole list. The outer one says the list itself may not be null: so a nulled list climbs into `Case`. A single unresolvable deadline in a list of 47 therefore erases the case node — an effect nobody intends when they type the wrapper out of habit. `[Deadline]` and `[Deadline!]` and `[Deadline]!` all fail differently, and on a cross-service field the choice is an availability decision, not a style one. ## Why this is a federation problem specifically In a single server, a Non-Null field and its parent usually fail together anyway: same process, same database, same outage. The `!` costs little. Across a composed graph the assumption breaks. The failure domains are genuinely separate — that separation is the entire reason for splitting the graph — and a `!` written in one subgraph silently re-couples them. Composition checks the field's *type* against other subgraphs; it does not and cannot tell you that Conflicts at 99.2% availability just capped `Case` at 99.2%. With eleven services contributing to a handful of core types, availability multiplies. Ten enrichment fields, each Non-Null, each on a service that is up 99.5% of the time, give a core type worse availability than any single service in the graph. ## The design rule A field is only as Non-Null as the reliability of the service that fills it. Keep `!` for fields the owning subgraph resolves locally from its own store — an entity's own identifier, its own name. At an ownership seam, make the contributed field nullable, and let its *inner* structure be strict: `conflictCheck: ConflictCheck` where `ConflictCheck` itself has non-null fields is both honest and safe, because a failure now leaves one hole instead of a crater. Enforce it at review time rather than by exhortation: a schema check that flags a Non-Null field added to a type the subgraph does not own catches this in CI, where fixing it is free. ## Diagnosing it after the fact When `data` comes back null with a single error, read the error's `path`. Its last segment is the field that failed; every segment before it is a level that got erased on the way up. That is the whole diagnosis: the depth of the path tells you how much collateral there was, and the field at the end tells you which team to talk to.

  • Composition succeeded on that schema. Why did no check catch the coupling?
    Composition validates that the merged schema is well formed and that every field is reachable by some plan — type mismatches, enum drift, satisfiability. Availability is invisible to it. `ConflictCheck!` is a perfectly valid declaration; nothing in the composed artefact records that the service behind it is less reliable than the type's owner. Catching it needs a schema lint rule of your own: flag a Non-Null field added to a type this subgraph does not own.
  • Would marking that branch with @defer stop it erasing the case?
    No. @defer changes *when* a piece of the response arrives, not what happens when it fails. The initial payload is delivered without the deferred branch, which does help a slow subgraph stop blocking the rest of the screen, but if the deferred field is Non-Null and its fetch fails, the null still has nowhere to sit inside that incremental payload. Deferring is a latency tool; nullability remains the availability tool.
  • Which fields on a federated type should still be Non-Null?
    The ones the owning subgraph resolves from its own store, where the field and its parent share a failure domain: an entity's identifier, its number, its name. There the exclamation mark buys a cleaner client type at almost no availability cost. The rule bites only at ownership seams, where the field and its parent come from processes that can fail apart.

A case bundle where one missing folder is stamped mandatory, so the clerk refuses to hand over the whole bundle — including the ten folders sitting complete on the desk.

saying these in an interview costs you the question

  • Treats Non-Null as a statement about the data, not the response
  • Thinks the router substitutes a default when a fetch fails
  • Believes composition warns about availability coupling
  • Copies the parent type's Non-Null style onto contributed fields
  • Cannot explain why a healthy subgraph's data got discarded
  • Assumes the outer and inner list markers fail the same way

context