In a GraphQL API, why should request errors and field errors drive different retry and alerting rules?
answer
- Determinism, not severity
- One can never succeed on retry
- One says the caller broke
- Baseline is zero versus never zero
- Alert on presence versus on rate
basics
~20 sA request error is deterministic — the same document fails identically forever — so retrying is waste and the signal is a client or schema release. A field error is a transient dependency fault where a retry can succeed.
solid answer
~50 sThe classes differ in determinism, and retry policy is built on determinism. A **request error** means the document did not parse, did not validate, or its variables did not coerce; nothing about waiting changes that, so a retry cannot succeed and only burns budget. It also tells you *who* broke: a caller and the schema disagree, which in practice is a client build or a schema publish. Because a healthy graph sits at essentially zero request errors, alert on their presence and dimension by operation and client version. A **field error** is one dependency failing inside execution — often transient, so a retry of an idempotent read is reasonable, while retrying a mutation needs an idempotency mechanism. Field errors are never zero across a large graph, so alert on rate and derivative per error path. Counting both in one metric hides the sensitive class behind the noisy one.
code
pseudocode · 19 linesfunction classify(responseBody):
if not hasKey(responseBody, "data"):
return REQUEST_ERROR # never executed; retry cannot help
if hasKey(responseBody, "errors"):
return FIELD_ERRORS # executed; data is partial
return OK
function handle(operation, responseBody):
switch classify(responseBody):
case REQUEST_ERROR:
countRequestError(operation.name, client.version)
fail(responseBody.errors) # no retry, ever
case FIELD_ERRORS:
for e in responseBody.errors:
countFieldError(rootFieldOf(e.path))
if operation.type == QUERY and transient(responseBody.errors):
retryWithBackoff(operation)
else:
render(responseBody.data, responseBody.errors)go deeper
Know the basic rule before the nuance: a request error will fail the same way every time, so retrying it is pointless, while a field error may be a passing dependency problem that a second attempt can get past.
Explain how you tell the classes apart in client code — the data entry, not the status code — and why a retry re-runs the entire operation rather than just the field that failed.
Demonstrate the operational split: separate metrics per class, alert on presence for request errors dimensioned by client version, alert on rate per error path for field errors, and retry only idempotent reads.
Own it as a contract across teams: request errors are a release-coordination signal that belongs to whoever shipped the client or the schema, field errors are a dependency-health signal, and one blended error metric hides both.
## Why the class, not the message, drives the policy Both error classes arrive over the same endpoint, in the same body shape, often with similar-looking messages. What separates them operationally is not severity but **determinism**, and determinism is what retry and alerting policy is actually built on. A **request error** is deterministic with respect to the request. The same document with the same variables, sent against the same schema, fails identically every time — it did not parse, or it did not validate, or its variables did not coerce. Nothing about the passage of time changes any of those. Retrying is guaranteed waste: it burns the caller's retry budget, doubles the load on parsing and validation, and cannot succeed until either the document or the schema changes. A **field error** is a runtime fault in the execution of one field. A downstream service timed out, a connection pool was exhausted, a permission lookup failed. Those causes are often transient, and a second attempt genuinely can succeed. ## What each class tells you about *who* broke This is the part interviewers are really probing. A request error says the caller and the schema disagree. In practice that means one of three things: a client shipped a document the current schema no longer accepts, someone rolled the schema forward past a deployed client, or a caller is sending malformed variables — and in a graph fronting an 11-service estate, the first two are release-timing problems rather than code problems. Steady-state request-error rate in a healthy graph is essentially zero, because every document a real client sends has been validated at build time or was working yesterday. So the signal is exquisitely sensitive: **any sustained non-zero request-error rate is a release event**, and its shape names the culprit. An operation that goes from 0% to 8.4% failure the moment a client version starts rolling out is a client build; an operation that fails for *all* client versions at the instant a schema is published is a schema change. A field error says one dependency inside the graph is unhealthy. It says nothing at all about the caller — the document was fine, the variables were fine, execution began, and something the graph depends on let it down. In an estate that size, the field-error rate is never zero and should not be expected to be: with enough services, some fraction of some field path is always degraded. ## Which means the two need different alarms - **Request errors**: alert on *presence and persistence*, not on rate thresholds. A handful of them is a bad build in canary; thousands is a schema change nobody coordinated. Dimension the metric by operation name and by the caller's client name and version, because the failure is deterministic and therefore always attributable to a specific build. This is the one class where paging a client team is the correct response. - **Field errors**: alert on *rate and derivative*, dimensioned by the error path's root field, because the baseline is non-zero. A jump from a 0.3% to a 6% error rate on one field path is a dependency incident; a flat 0.3% is a Tuesday. Paging on the mere presence of a field error in a graph this size produces alert fatigue within a day. Putting both in one counter is the anti-pattern the question exists to catch. It gives you a number that is dominated by the noisy class and blind to the sensitive one — precisely backwards, since the rare class is the more actionable of the two. ## Retry rules that follow - **Never retry a request error.** Fail fast, surface it, and treat it as a bug report. If a client library retries these by default, turn that off. - **Retry a field error only when the operation is a read and the cause looks transient.** A query is safe to re-issue. A mutation is not safe to blanket-retry, because execution may already have applied part of the work before the failing field — how partial writes behave is its own subject, and the honest interview answer is that retrying them needs an idempotency mechanism, not a policy switch. - **Retry the whole request, not the hole.** GraphQL has no protocol-level way to re-request just the failed field, so a retry re-runs everything, including the fields that succeeded. That cost is often the reason to render partial data instead of retrying at all. ## Classifying reliably in code The signal that survives every reverse proxy in front of the graph is the response body: **an absent `data` entry is a request error; a `data` entry that is present, including present-and-null, means execution began and the errors are field errors.** Do not classify on the transport status code — how statuses map to the two classes is a separate subject with its own subtleties, and a shared retry layer that guesses from a status will guess wrong. ## None of this is in the specification Worth saying plainly, because a strong candidate volunteers it. The specification defines the two classes and what the response looks like for each. It says nothing about retries, alert thresholds, dashboards, or whose pager fires. That is engineering policy built on top of a specified distinction — and the distinction is only worth having because the policies genuinely differ.
- Is it ever right to retry a field error on a mutation?Only with an idempotency mechanism the server honours, such as a caller-supplied key it deduplicates on. Root mutation fields execute serially, so a failure partway through can leave earlier work applied; a blind retry then risks applying it twice. Without that mechanism the correct move is to surface the failure and let the caller decide, not to re-issue the operation.
- Your request-error rate for one operation jumps from zero to 8% overnight. How do you find the cause quickly?Slice the metric by client name and version, and by operation name. Because the failure is deterministic, it maps to a specific build or a specific schema publish. A rate that tracks a client version's rollout curve is a bad client build; a rate that appears for every client version at the moment a schema shipped is a breaking schema change that got past the checks.
- Why should a shared retry layer not classify on the HTTP status code?Because status mapping is a separate concern with its own conventions and is not a reliable proxy for the two error classes — a reverse proxy in front of the graph can also rewrite it. The response body is the dependable signal: an absent data entry means the operation never executed, and a present data entry means it did and the errors are per-field.
One is a returned letter with the wrong address on it — posting it again changes nothing until you fix the address; the other is a letter lost at a sorting office, where a second attempt genuinely may get through.
saying these in an interview costs you the question
- Retries every response that carries an errors entry
- Counts both error classes in one metric
- Pages on any single field error in a large graph
- Assumes a request error might succeed later
- Classifies retryability from the transport status code
- Blanket-retries mutations after a field error