Which failures belong in a GraphQL result payload, and which belong in the top-level errors list?
answer
- Ask whether the request itself was valid
- Ask whether the client renders it differently
- A caught fault leaves the errors list empty
- Dashboards count entries, not schema types
- Faults keep their path; outcomes get a type
basics
~20 sExpected outcomes of a valid request that a client renders differently belong in the payload as schema types. Faults — timeouts, bugs, unavailable dependencies — belong in the errors list, where operators and generic client handling can see them.
solid answer
~50 sTwo tests decide it. **Is it an expected outcome of a valid request?** An offer below the seller's reserve is; a read timeout to a geocoding dependency is not. **Does the client render it differently?** If the only handling is a generic failure notice, the errors list already provides that for free. Model the failure as data when both hold, because then it earns a documented, typed shape that clients must handle. Leave a fault as a field error, so it carries a `path` locating the field and, crucially, so it appears in the errors list at all: catching a dependency timeout and returning it as a union member produces a response with no errors entry, and anything counting errors — dashboards, alerts, a client's generic failure path — sees a completely healthy request. The reverse mistake is as bad: throwing an expected domain refusal leaves clients parsing a message string.
code
graphql · 15 linestype Query {
searchListings(area: String!, maxPrice: Int): ListingSearchResult!
}
union ListingSearchResult = ListingPage | AreaNotRecognised
type ListingPage {
listings: [Listing!]!
matchCount: Int!
}
type AreaNotRecognised {
submittedArea: String!
suggestions: [String!]!
}go deeper
Learn the two categories first: an expected answer the business defines, versus something that broke. Being able to sort a few examples into those buckets is what is expected here.
Explain what each choice costs the client — a typed branch it must write, versus a generic handler that already exists — and why a nulled field with a prose message is a poor contract for a domain outcome.
Bring the operational consequence: a caught fault returned as data produces a response with no errors entry, so error-rate panels, alerts and the client's generic failure path all see a healthy request.
Own the guardrail rather than the judgement call. Catch-all handlers in resolvers are where this rots, so decide how the boundary is reviewed, tested and instrumented before dozens of teams each make their own call.
## The two questions that settle it **Is this an expected outcome of a valid request?** In a real-estate listings graph, `OfferBelowReserve` is: the request was well-formed, the caller was allowed, the system worked perfectly, and the answer is no. A read timeout to the geocoding service is not — the system did not work, and the same request a second later might succeed. **Does a client render it differently?** A withdrawn listing gets its own card with a date and a reason. A timeout gets whatever generic notice the application already shows for anything that goes wrong. If the answer is "generic notice", the errors list already gives you that for free and a new union member buys nothing but selection-set weight. Model as data when both answers are yes. Leave it to be a field error otherwise. ## The failure that only shows up in production Here is the shape of a real incident, and the reason this question is asked at senior level rather than as a modelling curiosity. A listings search field calls a geocoding dependency to turn a free-text area into a bounding box. The team, having adopted errors as data enthusiastically, wrapped the resolver in a catch-all: any exception became a `SearchFailed` member of the result union, with a friendly message. Locally, where the dependency answers in 4 ms, nothing ever hit that branch. In production, at roughly 4,180 searches a minute, the dependency's p99 crossed a 2,340 ms client timeout and about 62 searches a minute began failing. Every one of those requests produced a **successful** response: `data` present, the union member set to `SearchFailed`, **no `errors` key at all**. The consequences compounded: * The service's error-rate panel, which counts responses carrying entries in the errors list, read 0.00%. The alert never fired. * No `path` entry existed to say which field broke, because no field error was ever raised. * The client's generic failure handling — retries, the error boundary around the results pane — never ran, because as far as it was concerned the query succeeded. * Users saw an empty results pane with a polite sentence, and the team learned about it from a support ticket rather than from monitoring. The bug was not the timeout. The bug was that a **fault had been reclassified as a domain outcome**, and that reclassification deleted it from every place operators look. ## Why the errors list is the right home for a fault A field error is not a failure of the API design; it is the specified mechanism for "this field could not be resolved". Raising one gives you three things you cannot get from a data-carried failure. It is **counted** — every server, proxy and dashboard in the path knows what an entry in the errors list means, without knowing your schema. It is **located** — the entry carries a path identifying the exact field, list indices included. And it is **generic to the client** — one handler covers every fault in every operation, present and future, with no per-field code. A data-carried failure has none of that by construction. It is a value, and values are indistinguishable from success to anything that does not read your schema. ## Why the errors list is the wrong home for a domain outcome The mirror-image mistake is just as common. Throwing on "offer below reserve" gives the client a null and a human sentence. To behave correctly it must string-match that sentence, or branch on an extensions code that nothing in the schema documents. Neither survives a copy edit. The failure is invisible in introspection, absent from generated client types, and untestable as a contract. Meanwhile the null it produced may propagate up a Non-Null chain and blank out unrelated parts of the response. ## The middle ground worth knowing Some failures are genuinely both — an upstream refusing because a rate limit was hit, say. Two workable answers. Keep it a field error and put a stable machine-readable marker in the error's extensions, which keeps it counted and located. Or model it as data *and* keep emitting a metric and a log at the point you catch it, so the operational signal survives even though the response looks clean. What is not defensible is a blanket catch-all that turns every exception into a data member: that is not a modelling decision, it is a monitoring outage waiting for the first slow dependency. ## How to say it in an interview State the two tests, give one example on each side, then name the failure mode: swallowing faults into data makes a broken system look healthy, and dressing domain outcomes as errors makes clients parse prose. The interviewer is checking whether you have operated a graph, not whether you like unions.
- How would you catch the flat error-rate problem before an incident does?Instrument the data-carried failures too: emit a counter per union member so a `SearchFailed` rate is graphable and alertable alongside the errors-list rate. Then assert the boundary in tests — a test that makes the dependency time out should observe an errors entry, not a data member. Reviewing catch-all handlers in resolvers is the third leg; they are where faults quietly become outcomes.
- If you must carry an operational fault in data, what do you owe the operator?A separate member, never one reused for a domain outcome, so the two can be counted apart. A metric and a log emitted at the point you catch it, carrying the same identity you put in the response. And an explicit note in the schema documentation that this member means the system failed, so the next reader does not treat it as a normal answer.
- What does a field error give a client that a data-carried failure does not?One handler for everything. A client can treat any entry in the errors list generically — log it, retry, show the fallback — with no per-field code and no knowledge of the schema. It also gets the failed field's location from the error's path. A data-carried failure requires the client to have selected that branch in advance and written code for it.
saying these in an interview costs you the question
- Wraps every resolver exception into a data payload
- Says a data-carried failure still shows in metrics
- Throws on ordinary domain refusals like a rejected offer
- Reuses one failure member for outcomes and faults
- Judges by exception type rather than expectedness
- Assumes the client retries what looks like a success