How do you mask unexpected GraphQL errors without going blind in production?
answer
- One translation step, not many catches
- Fixed text out, full detail down
- An identifier joins the two halves
- Per error, not per request
- Count masked errors or go blind
basics
~20 sReplace an unexpected failure's message with a fixed string, attach a correlation identifier the caller can quote, and log the full exception under that same identifier. The client learns nothing internal, and an engineer can still find the exact failure.
solid answer
~50 sMask at the single place every field error passes through on its way out, not at each throw site — that way code nobody remembered to guard is masked too. For an unexpected failure, the outgoing entry carries a fixed message such as `Internal error`, a stable machine-readable code under `extensions`, and a freshly generated identifier; the same identifier goes into the server log beside the exception, the operation name and the failing field path. Both halves are mandatory: the identifier is worthless without the log line, and the log line is unfindable without the identifier. Be precise about what is specified — `extensions` is a free-form map on an error, but the keys inside it, `code` and any correlation key alike, are conventions your clients and support process agree on. Generate the identifier per error rather than per request so one request producing several failures stays separable.
code
pseudocode · 12 linestranslate(field_error):
cause = field_error.cause
if cause is ClientFacingError: # deliberate, reviewed, safe text
return entry(message = cause.message,
extensions = { code: cause.code })
id = random_id() # per error, not per request
log.error(id, operation_name, field_path, cause) # full detail stays here
return entry(message = "Internal error",
extensions = { code: "INTERNAL_ERROR", errorId: id })go deeper
Know the two halves. A fixed message plus an identifier goes to the caller; the real exception plus that same identifier goes to the server's log. Either half on its own is useless.
Explain why the masking lives in one translation step rather than in each resolver, and be able to say which keys under an error's extensions are conventions rather than specified names.
Show that masking creates a monitoring debt: you removed the caller's ability to report detail, so you owe a masked-error metric by operation and field path, and a log record a support engineer can actually retrieve.
Decide the contract once for the whole graph — the fixed text, the code vocabulary, the identifier format and the log schema — so support tooling behaves identically across services instead of being reinvented per team.
## Masking is one step, not a habit The wrong shape is a catch inside every resolver that rewrites the message. It works exactly as long as the last person to add a resolver remembered, and it is unreviewable: a diff adding a field without a guard looks identical to one adding a field whose backend is safe. The right shape is a single error-translation step that every field error passes through before the response is serialized. Put the decision there and the safe outcome becomes the default for code nobody thought about. ## What the outgoing entry carries Three things, each earning its place. A fixed message — the same text for every unexpected failure, so the text itself carries no signal at all. A machine-readable code under the error's `extensions`, so a client can tell "we do not know what happened" from a deliberate outcome without parsing English. And a correlation identifier, also under `extensions`, short enough that a user can read it off a screen and quote it to support. Be precise about what is specified here, because it is a cheap point to score and a cheap one to lose. The specification defines `extensions` on an error as a free-form map. It names nothing inside it. A `code` key is a widespread convention; a correlation key is purely a local agreement between your server, your clients and your support process. Treating `code` as a specified field is a common and easily caught mistake. ## The logging half is the one people forget Masking moves information; it does not destroy it. The same identifier must be written to the server's log alongside the full exception and its cause chain, the operation name, and the field path that failed. The common failure is shipping the identifier to the caller and logging the exception without it. The identifier is then decoration, and support can only correlate by timestamp across every service that was busy at that moment. ## The incident that makes the case An eleven-service music catalogue graph. A deploy changes `Album.releaseYear` from nullable to non-nullable in the catalogue service, but the backfill covering reissues has not run, so for a slice of the catalogue that field now receives a null it is not allowed to hold and the server raises a field error for each one. Masking is on and working perfectly: every affected caller sees `Internal error`. Over the next forty minutes the graph emits 3,412 masked errors, and the on-call engineer, looking at a dashboard of identical messages spread across eleven services, cannot tell which service, which field, or which records are involved. What ends the incident is one identifier a user quoted in a support ticket: `err_7f31c8`. A single log lookup names the service, the field path `album.releaseYear`, the affected album ids and the cause. That is the entire argument for the identifier. Masking deliberately destroys information on the client side; the identifier is what makes the destruction reversible for the people entitled to reverse it. ## Choosing the identifier Generate it per error, not per request: one request can produce several independent failures and you want them separable. Make it random and opaque. Never derive it from internal state — a database primary key, a row counter or an internal sequence tells the caller something about your storage and collides across unrelated failures. If you already run distributed tracing, reusing the trace identifier is tempting and often correct, but check two things first: that the identifier itself encodes nothing internal, and that handing it to a caller does not grant them a way into a trace viewer that displays internal service names and timings. ## Do not simply move the leak into the log Logs are read by more people, and retained longer, than most teams assume. Pushing the exception into the log is right. Copying the request's variables in wholesale is how personal data ends up in an aggregator with a different access policy from the database it came from. Always log the operation name and field path; log variables selectively or with redaction, and decide that once rather than per service. ## The monitoring debt masking creates The moment you mask, you have removed your callers' ability to tell you what broke, so you owe yourself the ability to notice. Emit a counter of masked errors keyed by operation name, field path and code, and alert on a change in rate rather than on a fixed threshold — a background level of masked errors is normal, a step change is not. That metric is the difference between the forty-minute version of the incident above and a five-minute one.
- Is the `code` key under an error's `extensions` defined by the GraphQL specification?No. The specification defines `extensions` on an error as a free-form map and says nothing about its contents. A `code` key is a widespread convention, and a correlation key is purely a local agreement. Treat both as a contract with your own clients — documented, versioned and reviewed like any other part of the API surface.
- Support has an identifier from a user, but the log lookup returns nothing. What went wrong?Almost always the identifier was generated at serialization time and never written to a log, or it was logged at a level the production configuration discards. Both halves have to be emitted in the same step, and the log level for that line has to be one production actually keeps. An identifier the server does not record is decoration.
- Can you reuse an existing distributed-tracing identifier as the correlation id?Often yes, and it saves a hop. Check two things first: that the identifier is opaque and encodes nothing internal, and that handing it to callers does not give them a route into a trace viewer showing internal service names and timings. If either check fails, generate a separate identifier and log both so they can be joined internally.
A hospital wristband number identifies your case precisely to anyone who can open the chart, and tells a stranger reading it over your shoulder absolutely nothing.
saying these in an interview costs you the question
- Masks inside each resolver instead of one place
- Returns an identifier the server never logs
- Uses a database key or counter as the identifier
- Thinks extensions.code is defined by the specification
- Logs the whole variables map, personal data included
- Masks without counting masked errors anywhere