In a large GraphQL schema, where is modelling errors as data worth its cost?
answer
- Treat each modelled failure as a spend
- The client's branch is the real cost
- Converting a field's type breaks documents
- Consistency across teams beats per-field elegance
- Nothing in the language enforces the convention
basics
~20 sOnly where a failure is a durable business outcome that clients render distinctly. Apply it through one house carrier and a small shared failure vocabulary, because a graph carrying three inconsistent styles is worse than either style applied everywhere.
solid answer
~50 sTreat it as a budget, not a principle. Each modelled failure costs a permanent contract entry, an inline fragment in every document that touches the field, a branch in every client, and — when it is introduced on an existing field — a breaking change, since documents selecting fields directly on an object type stop validating when that type becomes a union. So spend it where the failure is **durable** and **rendered differently**, and default to a field error everywhere else. At scale the consistency problem dominates the per-field one: if forty result types across nine teams each invent their own failure vocabulary, no client can share a fallback and the convention has bought nothing. Pick one carrier, publish a small failure taxonomy, give every failure a shared abstract type with one selectable message field, and enforce it in schema review and linting — the specification will not.
code
graphql · 22 linesinterface DomainFailure {
message: String!
code: FailureCode!
}
type OfferBelowReserve implements DomainFailure {
message: String!
code: FailureCode!
reservePrice: Int!
}
type ViewingSlotTaken implements DomainFailure {
message: String!
code: FailureCode!
nextAvailableAt: String!
}
type Mutation {
submitOffer(input: SubmitOfferInput!): SubmitOfferResult!
}
union SubmitOfferResult = OfferAccepted | OfferBelowReservego deeper
Take away the shape of the tradeoff: every modelled failure adds a branch that clients must write and maintain, so it is a choice with a price rather than an automatic improvement.
Be able to name the concrete costs — an inline fragment per branch in every document, a permanent contract entry, and a breaking change when an existing field's type becomes a union.
Show judgement per field: which failures are durable and rendered distinctly, and how you would roll the change out to existing clients without breaking their documents.
Own consistency and enforcement. Pick one carrier and a small shared failure vocabulary for the whole graph, publish the default, and put the gate in schema review and linting, because nothing in the language will do it for you.
## Why this is a budget and not a principle Errors as data reads like an unambiguous improvement: typed, documented, impossible to ignore. At the scale of one field it usually is. At the scale of a graph owned by many teams, each application of it draws on four accounts. **The contract account.** A failure type is permanent. Adding a member is easy; removing one breaks every client that branches on it. You are minting API surface that will outlive the reason you minted it. **The client account.** Every document touching that field grows an inline fragment per branch, every client grows a switch, and every branch is code somebody maintains. A team that models eleven failure kinds on a mutation has written eleven pieces of UI, or — far more often — one real branch and ten that fall through to a generic notice, which is the same outcome the errors list gives for free. **The migration account.** Converting an existing field from an object type to a union is a **breaking change**: documents selecting fields directly on that field stop validating, because a union selection set may contain only fragments and meta-fields. The honest rollout is to add a new field beside the old one, move clients using operation-usage data from the registry, and delete the old field when usage reaches zero. That is a quarter of coordination for a field, not an afternoon. **The consistency account.** This is the one that dominates, and the one candidates miss. Suppose 62 fields across nine teams adopt the convention independently. One team uses unions, another an interface, a third a failure list on the payload. One calls the human-readable field `message`, another `reason`, a third `detail`. Every client now writes bespoke handling per field, no shared fallback component is possible, and the graph is *harder* to consume than if nobody had bothered. **An inconsistently applied convention is worse than either choice applied uniformly.** ## Where the spend is clearly worth it Three patterns pay back reliably. * **Mutations whose refusals are the product.** Submitting an offer, booking a viewing, transferring a listing between agents: the refusals are business rules, they are stable, they are exactly what the screen has to explain, and they usually carry data — a reserve price, a next available slot. * **Failures carrying structured data.** If the client needs a date, a limit, a set of suggestions, that data has to live somewhere typed. A prose sentence in the errors list cannot carry it honestly. * **Failures with distinct, durable UI.** A withdrawn listing card is not a generic error notice, and it will still exist next year. ## Where it is not * **Faults of any kind.** Timeouts, unavailable dependencies, bugs. Catching those into data deletes them from the errors list, and with it from error-rate panels and alerting. * **"Not found" on a read, where null already says it.** A nullable field plus a documented meaning is often the whole answer, and it costs no fragments. * **Anything the client will render generically.** If every branch ends in the same notice, the convention has bought a switch statement. * **Authorization refusals, until you have decided what you are willing to disclose.** Making a refusal a typed, documented outcome is also making it a reliable signal about what exists. ## What you standardise, in order 1. **The carrier.** One choice for the graph. A shared **interface** for failures is the strongest default at scale, because it gives every failure a directly selectable message field and therefore lets every client write one fallback branch that renders members added after its document shipped. Unions are sharper per field and worse at evolution. 2. **The vocabulary.** A small, published set of failure classes with agreed names and an agreed machine-readable field, so a client can branch consistently across teams. Small is the operative word: a taxonomy of six that everyone uses beats a taxonomy of forty that nobody reads. 3. **The default.** Teams should not have to make this call per field. Publish "field error unless it meets the two tests" as the house default, and let modelling be the argued exception. 4. **The gate.** Nothing in the specification or in validation enforces any of this — to the server it is an ordinary schema. So it lives in schema review, a schema lint rule, and the registry's checks on every proposed change. A convention with no gate reverts to per-team taste within two quarters. ## How to close the answer Say what you would measure. Failure-member rates per class, so you learn which modelled failures anybody actually hits; field- and operation-usage data from the registry, so a member nobody branches on can be retired; and the count of distinct failure shapes in the graph, which is the real health metric for the convention. If that count is drifting up while the per-member usage is near zero, the budget is being spent on ceremony rather than on clients.
- What is the first thing you standardise when adopting this across many teams?The carrier and one shared abstract type with a selectable human-readable field. That single decision is what lets every client write one fallback branch that still renders correctly when a team adds a failure kind next quarter. The failure vocabulary comes second, and it should be small enough that people remember it without looking it up.
- How do you migrate a field that returns an object type today into one returning a result union?Treat it as a breaking change, because documents selecting fields directly on that type stop validating. Add a new field beside the old one, move clients across using operation-usage data from the schema registry, then remove the old field once its usage is zero. Without usage tracking that last step is a guess, which is why registries and this convention tend to arrive together.
- When does errors as data actively make a schema worse?When it is applied to faults, because the failure vanishes from the errors list and from monitoring. When every team invents its own failure vocabulary, because no client can share handling. And when a modelled failure carries nothing a client renders differently — then it is a permanent contract entry and an inline fragment bought in exchange for a switch statement.
- How do you know afterwards whether the investment paid off?Measure per-failure-member rates and client branching. A member nobody hits, or one every client routes to the same generic notice, is ceremony and should be retired into a field error. Track the number of distinct failure shapes in the graph too — drift there is the earliest sign the convention has become per-team taste again.
saying these in an interview costs you the question
- Applies result unions to every field on principle
- Calls the union conversion a backward-compatible change
- Lets each team invent its own failure vocabulary
- Models failures the client renders identically
- Expects the specification or validation to enforce it
- Ignores the migration cost for existing clients