skip to content

Why can adding a member to a GraphQL union break a field at run time, and how do you prevent it?

level: seniorimportance: should knowfreq 41%

answer

  1. Schema half checked, code half not
  2. New member, unchanged mapping
  3. Rare paths deploy quietly
  4. A 200 response hides it from status monitoring
  5. Make an unhandled member fail the build

basics

~20 s

Abstract type resolution lives in server code, not the schema: a new member passes validation while the resolver has no branch for it, so values of the new shape raise a field error in production. An exhaustive mapping fails the build instead.

solid answer

~60 s

Adding a member to a union or a new implementer to an interface is a **schema** change, but the thing that decides which member a value is is **server code**. Nothing checks that the two agree: the SDL parses, validation passes, a registry accepts the change, and generated client types compile. The gap only shows up when a value of the new shape is actually produced, and then the type-resolution step names no possible type and raises a field error — the field goes null, propagating outward if it is non-null, with an entry in `errors` naming its path, inside an otherwise ordinary 200 response. If the new member only appears on a rare path, that can be weeks after deploy. The fixes are structural: make the mapping exhaustive over a closed set so an unhandled member fails the build, add a test that enumerates each abstract type's possible types and asserts every one resolves, alert on field errors by path, and never silence the gap with a fallback member.

code

graphql · 13 lines
graphql
union RecordDonationResult =
    DonationRecorded
  | DonationRejected
  | DuplicateDonation

type DuplicateDonation {
  originalDonationId: ID!
  suppressedAt: String!
}

type Mutation {
  recordDonation(input: RecordDonationInput!): RecordDonationResult!
}

go deeper

for a junior

Remember that the list of members in the schema and the code that decides which member a value is are two separate things, and that nothing automatically keeps them in step when a member is added.

for a middle

Be able to trace the failure end to end: unmatched value, field error at that path, null in data or propagation under a non-null field, and a 200 response envelope that carries the error alongside whatever else succeeded.

for a senior

Show the production instincts: why the bug survives staging, how you would have caught it from field-error-by-path telemetry, why the resolution branch ships before or with the schema member, and why a fallback member is worse than an error.

for a principal

Argue the general rule: any place where schema declarations and hand-written code must agree needs a mechanical check, not a convention. Decide whether that check is a compiler-enforced exhaustive mapping or a schema-enumerating test, and make it a platform default.

## The gap nothing checks Everything about an abstract type is declared except the one thing that matters at run time. The union lists its members in SDL. Validation checks that the client's fragments name possible types. A registry can check that the change is not breaking for existing operations. Generated client types recompile happily. And none of it inspects the function that decides, for a given value, which member it is — because that function is ordinary server code with no declared relationship to the schema. So the schema and the resolution logic drift, and the drift is invisible until a value of the new shape is produced. ## A concrete incident A charity donations service exposes `recordDonation`, returning a result union: * originally `DonationRecorded | DonationRejected`; * extended with `DuplicateDonation` when idempotency-key detection shipped, so that a client which retries after a gateway timeout is told plainly that its second write was suppressed rather than being handed a second receipt. The schema change and the detection logic ship. The type resolver still maps only two outcomes. Nothing fails in staging, because the duplicate branch requires a retry that arrives after the first write committed — a race nobody reproduces on purpose. In production it happens at a low rate, and every time it does, `recordDonation` comes back null with an error entry at path `["recordDonation"]`. The client's error handling reads that as "the donation failed", and the donor is invited to try again — the exact outcome the new member was added to prevent, and worse, because now the operator cannot tell from the response whether any write landed at all. ## Why it hides so well Four properties conspire: 1. **It is a field error, not a transport failure.** The response is a 200 with a partially null `data`. Monitoring built on status codes sees nothing. 2. **It is rare by construction.** New members are usually added for edge cases — duplicates, partial failures, newly-modelled states — which are precisely the paths integration tests do not walk. 3. **Nothing is red before deploy.** No compiler, linter, schema check or composition step relates the SDL member list to the resolver's branches. 4. **The error text points at the type system, not the caller,** so an on-call engineer without GraphQL background often files it as a client bug. The same failure arrives from the other direction too: a value whose discriminator column contains a state the mapping predates — written by an older service version, or backfilled — hits the identical dead end. ## Making it structural **Exhaustiveness that fails the build.** Write the mapping as a total function over a closed set of outcomes so that adding a case without extending the mapping does not compile. This converts a production field error into a build failure, and it is the single highest-value change. **A schema-driven coverage test.** Where the language cannot enforce totality, read each abstract type's possible types from the built schema and assert that resolution produces every one of them for a representative value. The test then fails the moment a member is added, and it keeps working as the schema grows because it enumerates rather than hardcodes. **Ship the two halves together.** The SDL member and its resolution branch belong in one change. If they must be separate, land the resolver branch first — a branch for a member no value yet has is inert, while a member with no branch is a live fault. **Alert on field errors by path.** Field errors grouped by path turn this class of bug from a customer report into a graph on a dashboard. A path that has never errored before and starts erroring after a deploy is the signature. **Never install a fallback member.** Returning "some member" when nothing matches trades a loud failure for a wrong answer: the client gets a value it will render as a different outcome. Duplicate-suppression reported as a successful recording is materially worse than an error. ## The half that is not the server's Even when resolution is correct, an added member is a client-visible change. A document whose inline fragments cover only the old members receives nothing type-specific for a value of the new one. That is not an error — it is a legitimately near-empty object — and it is why teams that add members regularly ask clients to keep a default branch in their handling from the first day the union exists.

  • The failing field is declared non-null. What does the client see?
    Not just that field. A field error under a non-null field cannot be represented as null there, so the failure propagates to the nearest nullable ancestor — for a non-null root mutation field, that means `data` itself is null and every other root field's work is discarded. A rare unresolved type therefore blanks the whole response rather than one branch of it.
  • How would you write a test that fails the moment someone adds a member?
    Enumerate, from the built schema, the possible types of every union and interface, and assert that type resolution returns each one for a representative value. Because it reads the possible-type set rather than a hardcoded list, adding a member with no branch makes it fail immediately, and it needs no edit when the schema legitimately grows.
  • Why not have the resolver fall back to a default member when nothing matches?
    Because it converts a visible failure into a wrong answer. The client receives a well-formed value it will render as the wrong outcome — a suppressed duplicate reported as a successful recording — and no error is emitted for anyone to alert on. Failing loudly at one field is cheaper than silently misreporting state.
  • Does adding a union member affect clients even when resolution is correct?
    Yes. An existing document whose inline fragments cover only the old members gets no type-specific fields back for a value of the new one, so its UI branch falls through. It is not an error and no validation rejects it, which is why a default branch in client handling should exist from the day the union ships.

saying these in an interview costs you the question

  • Says schema validation would catch an unhandled member
  • Thinks the failure surfaces as an HTTP 500
  • Silences it with a fallback member type
  • Blames the client for not sending the right fragment
  • Ships the schema member before the resolution branch
  • Assumes staging traffic exercises rare result members

context