How do you design the extensions payload of a GraphQL error so clients can branch on it safely?
answer
- Never branch on human prose
- One map the specification leaves to you
- code is convention, not specification
- Additive vocabulary, default branch, namespaced keys
basics
~20 sPublish a small, stable, machine-readable code vocabulary inside extensions and have clients switch on that, never on the message string. Keep codes additive, require a default branch for unknown ones, and namespace anything beyond the code to avoid collisions.
solid answer
~50 sThe specification says only that `extensions` must be a map and constrains its contents not at all, so that map's schema is **your** contract, validated by nothing. Design it as a published API. Emit a short machine-readable `code` — a convention, not a specified key — and make that the only thing clients branch on, because `message` is prose that gets reworded, localised or masked without anyone calling it a breaking change. Treat the code set as **additive**: new codes are safe, redefining one is an undetectable break, and every client needs a default branch for codes it has never seen. Keep the set coarse — one code per client behaviour, not per throw site. Namespace richer detail under one owned key so a tracing layer or an intermediary writing into the same map cannot collide with you. Across several teams, uniformity comes from a shared mapping layer and one published list, not from the format.
code
json · 9 lines{
"message": "Bin is locked for cycle count",
"path": [ "warehouse", "bins", 4, "stockItems", 11, "reorderPoint" ],
"extensions": {
"code": "BIN_LOCKED",
"retryable": true,
"wms": { "binId": "WH-7-C14", "lockExpiresAt": "2026-04-02T05:31:00Z" }
}
}go deeper
Know that extensions is where a server puts machine-readable detail such as a code, and that your client should read that code rather than the message text when deciding what to do.
Explain that the specification constrains extensions only to being a map, that code is convention rather than a specified key, and why an unversioned message string is the wrong branch condition.
Demonstrate the failure mode you have lived through: a reworded message silently disabling recovery logic. Then give the discipline — additive codes, a documented meaning per code, a default branch, small namespaced payloads.
Own the vocabulary as an organisational contract: one published list, a shared mapping layer so every service emits the same shape, a rule that codes are added not redefined, and a view on what belongs in errors versus in the schema.
## The one map the specification hands to you A GraphQL error may carry an `extensions` entry. The specification requires only that its value be a **map**, and then explicitly declines to constrain what goes in it — it is reserved for implementors, with no further restrictions. In the same breath the specification asks services **not** to add other entries at the top level of an error map, since that namespace belongs to future editions of the specification. Put those two rules together and you get the design constraint: **every server-specific fact about a failure goes inside `extensions`, and its schema is entirely yours.** Nobody will validate it, nobody will version it, and no tool will warn you when you break it. That is why this is a design question rather than a recall question. ## Rule one: clients branch on a code, never on the message `message` is prose. Nothing in the format pins its wording, so it changes — a library upgrade rewords it, a localisation pass translates it, a hardening pass replaces the informative version with something deliberately bland. A worked failure, from a warehouse inventory graph. A live stock dashboard, fed by a subscription, appeared to keep working and quietly stopped updating one Tuesday: the tiles rendered but never changed. Nobody was paged, because nothing had thrown. What had happened was that each delivered payload had begun carrying a field error, and the client's reconnect-and-refetch logic was guarded by `err.message.startsWith("Bin is locked")`. A server release had reworded that message to mention the bin id. The guard stopped matching, the recovery path stopped running, and the failure became invisible — the worst class of outage, one with no signal at all. The fix is the convention almost every implementation converges on: an `extensions.code`, a short stable machine-readable string such as `BIN_LOCKED`. Worth stating plainly in an interview: **`code` is a convention, not a specified key.** The specification knows about `extensions` and nothing about its contents. ## Rule two: treat the code vocabulary as a published, additive contract Once callers switch on codes, the set of codes is an API surface with all that implies. * **Write it down**, in the same place the schema lives, with the meaning and the expected client reaction for each code. * **Add, never repurpose.** Introducing `BIN_QUARANTINED` is safe; silently changing what `BIN_LOCKED` means is a breaking change no consumer can detect. * **Require a default branch.** Every client must behave sanely on a code it has never seen — treat it as a generic failure — because a new code will always reach an old client. * **Keep the set small and coarse.** Codes exist so a caller can choose a behaviour: retry, re-authenticate, show a field-level message, give up. One code per behaviour, not one per throw site. A thousand-code vocabulary is one nobody switches on correctly. ## Rule three: namespace anything beyond the code `extensions` is a single flat map and more than one component along the response path may want to write into it — the field's own resolver, a tracing layer, an intermediary that composes several services. Two of them choosing `details` is a silent collision. Keep the conventional keys (`code` and whatever your organisation has standardised) at the top level and nest everything else under one owned key, so additions never collide: ```json { "message": "Bin is locked for cycle count", "path": ["warehouse", "bins", 4, "stockItems", 11, "reorderPoint"], "extensions": { "code": "BIN_LOCKED", "retryable": true, "wms": { "binId": "WH-7-C14", "lockExpiresAt": "2026-04-02T05:31:00Z" } } } ``` ## Rule four: uniformity is the hard part, not the format The format is trivial; getting one vocabulary across an organisation is not. A graph assembled from several teams' services will produce several code vocabularies — different casing, different granularity, some services emitting no code at all. Callers then write per-service special cases, which is exactly the coupling a single graph was supposed to remove. The fix is boring and organisational: one published code list, one shared error-mapping layer each service uses rather than reimplements, and a default mapping so an unmapped exception still emits a known generic code instead of nothing. ## Rule five: content, and what belongs elsewhere Two limits on payload. Keep it **small** — an error can be emitted per failing list element, so a fat `extensions` map multiplies. And keep it **audience-appropriate**: a correlation id is useful to everyone, whereas internal diagnostics are a separate judgement call about what an untrusted caller may see. Structured detail for a program goes in `extensions`; a sentence for a human goes in `message`; and anything a client is expected to render as normal product behaviour — a validation failure on a form, a business rule refusal — is often better modelled in the schema as data than as an error at all.
- Several teams own services behind one graph and each invented its own codes. How do you converge them?Publish one code list beside the schema, with the meaning and the expected client reaction per code, then ship a shared error-mapping layer every service uses instead of reimplementing. Give it a default mapping so an unmapped exception still emits a known generic code rather than nothing. Converge by adding the canonical codes first and retiring the old ones once consumers report they no longer read them — never by redefining a code in place.
- Should a failed business rule be an errors entry with a code, or modelled in the schema as data?If the client is expected to render it as ordinary product behaviour — a rejected quantity, a locked bin — modelling it in the schema is usually better, because it is typed, discoverable through introspection and cannot be missed by a caller that ignores `errors`. Reserve the `errors` list for faults: things that went wrong rather than outcomes the product anticipates. That boundary is a design choice worth stating explicitly rather than drifting into.
- What limits how much you put in an error's extensions map?Volume and audience. One failing field inside a large list can produce an error per element, so a fat map multiplies across the response and inflates every payload. And the map goes to the caller, so what belongs in it is a deliberate decision rather than a dump of internal state — a correlation id that lets an operator find the trace is usually the highest-value thing you can add per byte.
saying these in an interview costs you the question
- Matching on message substrings to drive client behaviour
- Calling extensions.code part of the GraphQL specification
- Repurposing an existing code's meaning instead of adding one
- Shipping a client with no default branch for unknown codes
- Inventing one code per throw site in the server
- Writing flat unnamespaced keys that other layers overwrite