skip to content

How do you keep deliberate, client-actionable GraphQL errors legible while masking the rest?

level: seniorimportance: should knowfreq 42%

answer

  1. Not every failure is a surprise
  2. Decide at the throw, not the wire
  3. Never classify by matching message text
  4. Unclassified means masked, by construction
  5. A legible message is public API

basics

~20 s

Classify at the throw site, never by inspecting message text. A failure explicitly marked client-facing keeps its reviewed message and stable code; everything else is masked. Masking must be the fallback, so an unclassified failure is never legible by accident.

solid answer

~50 s

Two classes of failure share one channel: deliberate outcomes a caller can act on — a rejected input, an exhausted quota, a track not licensed in the caller's region — and unexpected failures nobody planned. Make the classification where the failure is raised, using a distinct error type or explicit marker the translation step recognises, never by matching on message text; string matching breaks silently when a dependency rewords its exceptions. Make masking the fallback, so an unclassified failure is masked by construction — that is the only version of the rule that survives a new dependency. Then treat every client-facing message as public API: reviewed for internal detail, stable across releases, and paired with a machine-readable code, because clients should branch on the code rather than on English. A message quoting a raw database constraint name is a leak wearing a friendly label.

code

pseudocode · 12 lines
pseudocode
translate(exception):
    facing = find_in_cause_chain(exception, ClientFacingError)   # frameworks re-wrap

    if facing != null:
        return entry(message    = facing.message,                # reviewed, stable text
                     extensions = facing.details + { code: facing.code })

    id = random_id()                                             # fail closed
    log.error(id, operation_name, field_path, exception)

    return entry(message    = "Internal error",
                 extensions = { code: "INTERNAL_ERROR", errorId: id })

go deeper

for a junior

Know that not every error should be hidden: a rejected input or an exhausted quota is information the caller needs. Which class a failure belongs to is decided by the code that raises it, not by the caller.

for a middle

Explain the marker-type mechanism and why the fallback must be masking rather than disclosure. Be ready to say why matching on message text is fragile once a dependency rewords its exceptions.

for a senior

Demonstrate the failure you have actually lived through — a client-facing error wrapped by an intermediate layer and masked by accident. Show the cause-chain walk, the test that pins it, and a review habit for the legible messages themselves.

for a principal

Own the error-code vocabulary as an API contract across the graph: who may add a code, what review a client-facing message gets, and how a rename is treated, so clients can branch reliably instead of parsing prose.

## Why a blanket mask is the wrong end state Two very different failures travel in the same channel. One is a deliberate outcome the caller can do something about: an input rejected by a rule, an exhausted quota, a track that is not licensed for playback in the caller's region. The other is an unexpected failure nobody planned for. A mask that flattens both into `Internal error` is safe and useless — it converts every actionable rejection into a support ticket, and it trains client developers to retry things that will never succeed. The engineering problem is not "hide errors", it is "hide exactly the ones that were never meant to be read, and hide them by default". Note the boundary: whether an expected outcome belongs in the errors channel at all, or in the schema as part of a result payload, is a separate design question. What matters here is that whatever you do put in the errors channel stays both legible and safe. ## Classify where the failure is raised The classification must be made at the throw site, by a distinct error type or an explicit marker the translation step recognises. The tempting alternative — inspecting the message and deciding from its text — is a bad mechanism for a reason worth articulating: it couples your disclosure boundary to strings written by someone else. A dependency upgrade rephrases a message and the classification silently flips, in either direction. A leak becomes legible, or an error your clients depend on becomes an opaque `Internal error`. Neither breaks a test unless someone happened to pin that exact sentence. ## Fail closed The fallback must be masking. Concretely: the translation step looks for the client-facing marker, and if it does not find one, it masks. Stated the other way round — mask only what is marked sensitive — the rule fails the moment a new dependency raises a type nobody has seen, which is precisely the case you are defending against. Failing closed also makes the rule survive staff turnover: the person adding a resolver next month gets the safe behaviour without knowing this policy exists. ## The wrapping problem, which is what actually breaks in production The failure you will really see is the opposite of a leak. A deliberate, reviewed, client-facing error is raised deep in a service, and something between there and the translation step catches it and re-raises it inside its own type — a transaction helper, a retry wrapper, a mapping layer. The translation step inspects the top-level exception, finds no marker, and masks a message the client needed. Users report that a perfectly ordinary rejection now says `Internal error`, and the log shows the correct error sitting in a cause chain. The fix is mechanical: walk the cause chain looking for the marker rather than inspecting only the outermost exception. The discipline is a test that raises a client-facing error wrapped in two layers and asserts the original message and code survive. ## A legible message is public API Everything you decide to show is now part of the interface, and needs reviewing as such. On a music catalogue graph, a rejection reading `check constraint "playlist_track_uq_2019" violated` is legible, actionable-looking, and a leak wearing a friendly label — it names a table's constraint and the year someone added it. The version a caller deserves is `A track can only appear once in a playlist.` with a stable machine-readable code alongside it. Three properties make a client-facing message safe and useful. It is written in the caller's vocabulary, never the storage layer's. It is stable, because clients that branch on behaviour need something that does not change under them — and the thing they branch on should be the code, not the English, since prose gets reworded and localized. And anything structured belongs in `extensions` as data rather than being interpolated into the sentence, so clients do not parse it back out. There is a subtler review question too: the *set* of legible codes is itself a disclosure. A code that says an artist merge job is in progress tells callers about an internal batch pipeline they had no reason to know exists. Keep the vocabulary in domain terms. ## Keeping the two halves honest over time Two tests carry most of the weight. One throws an arbitrary, unclassified exception from a resolver and asserts the caller receives exactly the fixed message with no trace of the original — that pins the fail-closed default. The other, per client-facing code, asserts the exact message and code a client will see, which turns any reword into a deliberate, reviewed change rather than an accident. Organisationally, the code vocabulary needs an owner. Who may add a code, what review a client-facing message gets before it ships, and whether renaming a code is treated as a breaking change are decisions that should be made once for the whole graph rather than eleven times. Clients can only branch reliably on a vocabulary that behaves like a contract.

  • Why is classifying errors by matching the message string a poor mechanism?
    It couples your disclosure boundary to text written by someone else. A dependency upgrade that rephrases a message silently reclassifies the failure in either direction — a leak becomes legible, or an error clients depend on becomes opaque — and neither shows up as a test failure unless somebody pinned that exact sentence.
  • A deliberate, client-facing error is being masked in production. What is the usual cause?
    It was wrapped. Something between the throw site and the translation step caught it and re-raised it in its own type, so the marker survives only in the cause chain. The translation step must walk that chain rather than inspect the outermost exception, and a test should raise a doubly-wrapped client-facing error and assert the message and code survive.
  • How do you stop a legible error message from becoming its own leak?
    Review the text as public API. A rejection quoting a database constraint name, a file path or an internal service name is a leak with a friendly label. Write the message in the caller's vocabulary, keep the machine-readable code stable, and put anything structured under `extensions` as data rather than interpolating it into the sentence.
  • Should a client branch on the message or on the code?
    The code. Messages are human text: they get reworded, localized and shortened, and they are the part most likely to change. A stable machine-readable code is the contract, so renaming one is a breaking change to the API, and the vocabulary should stay small enough that clients can handle every value they might see.

A departure board tells you your flight is cancelled and that you should rebook; it does not tell you which hydraulic line failed on the inbound aircraft. One fact is yours to act on, the rest belongs in the maintenance log.

saying these in an interview costs you the question

  • Masks everything, including errors clients must act on
  • Classifies errors by matching on message text
  • Defaults to legible unless an error is marked sensitive
  • Leaves a constraint or table name in a friendly message
  • Inspects only the outermost exception, never the cause chain
  • Treats error codes as free text that can be renamed freely

context