skip to content

How should GraphQL request metrics count errors returned in a successful response?

level: middleimportance: must knowfreq 57%

answer

  1. The transport says fine; the body disagrees
  2. Did execution begin at all?
  3. No data entry means request error
  4. Entries in the list are not requests
  5. Partial success deserves its own bucket

basics

~20 s

Meter the response body, not the transport outcome. Count request errors — parse or validation failures, where no data entry is produced — separately from field errors, where data is present with nulls and each entry carries a path.

solid answer

~50 s

A GraphQL response can describe failure in its body while the transport reports success, so an error-rate metric built from transport outcomes alone reads near zero during a real outage. Meter the envelope. A **request error** — a malformed document, a validation failure, a missing required variable — means execution never began: the response carries `errors` and no `data` entry at all, and the caller got nothing. A **field error** means execution ran and one or more resolvers failed: `data` is present, the failed positions are null, and each entry should carry a `path`. Those need separate counters, because one is a total failure and the other is often a partial success. Count *requests*, not error entries: a document whose 47 payslip rows each failed yields 47 entries and one failed request. Keep a third bucket — executed with at least one field error — so degradation is visible rather than rounded up to success.

code

json · 8 lines
json
{
  "errors": [
    {
      "message": "Cannot query field 'employerContributionCents' on type 'BenefitEnrollment'.",
      "locations": [{ "line": 4, "column": 7 }]
    }
  ]
}

go deeper

for a junior

Be ready to say that GraphQL reports failure in the response body, so a monitor watching only the transport outcome can read zero errors during an outage. Know that the body may carry both data and an errors list at once.

for a middle

Explain the structural split: no data entry means execution never began, while data present with a non-empty errors list means partial success with a path per failure. Be able to say why you count requests rather than error entries.

for a senior

Show you can diagnose from the split. Talk through what a rise in request errors versus field errors each implicates, why message text is a forbidden label, and how errors-as-data schemas silently move failures into your success bucket.

for a principal

Own the definition of failure for the organisation. Decide what counts as a failed request when a response is partial, make that predicate explicit and uniform across teams, and ensure it is what alerting and any objective are computed from.

## Why the usual error metric reads zero Every generic HTTP monitor computes error rate the same way: count responses whose status indicates failure, divide by total. On a GraphQL endpoint that computation can sit at zero through an outage, because GraphQL's failure channel is the response *body*. The execution result is a map with up to three entries — `data`, `errors` and `extensions` — and a response that carries a fully populated `errors` list is still, from the transport's point of view, a response that was produced successfully. Exactly which status accompanies which outcome is the transport layer's concern and varies with the media type in play; for metrics the rule is simpler and does not depend on it. **Parse the body.** ## The two failure classes, and why they are not one counter The envelope distinguishes them structurally, so your metric can too. **Request errors** happen before execution begins: the document does not parse, it fails a validation rule, the requested operation is not in the document, a required variable is absent or of the wrong type. The response contains `errors` and — this is the checkable signal — **no `data` entry at all**. The caller received nothing usable. Nothing your resolvers do is implicated; the document itself was rejected. **Field errors** happen during execution: a resolver raised, or returned something the field's type could not accept. Execution continues, the failed position is set to null (propagating up to the nearest nullable parent), and an entry is appended to `errors` describing what failed and, for field errors, a `path` locating the position in the response. `data` is present. The caller usually received most of what it asked for. One is "we could not run your request". The other is "we ran it and part of it broke". Rolling them into one `graphql.errors` counter destroys the only distinction that changes what you do next, and it makes the counter unreadable: a spike could mean a client shipped an invalid document or a downstream benefits provider is down, and the number cannot tell you which. ## Count requests, not entries The `errors` list is a list. A reconciliation document that touches an 8,400-row page of payslips can fail in thousands of positions from one downstream outage, and a counter incremented once per entry will report thousands of "errors" for a handful of requests. That metric has no denominator you can reason about: the ratio of error entries to requests is not a rate of anything. Count at the request grain. For each executed request, emit exactly one increment into exactly one outcome bucket: * `request_error` — no `data` entry; execution never began. * `field_error` — `data` present and `errors` non-empty; partial result. * `ok` — `data` present and no `errors` entry. Now error rate is a fraction of requests, it sums to one, and it can be dimensioned by the same operation identity and operation type as your duration metric. If you also want a sense of blast radius inside a request, record the count of error entries as a *separate* distribution — but never as the numerator of a rate. ## A worked incident: a client pinned to a removed field A payroll and benefits schema drops `BenefitEnrollment.employerContributionCents` in a release. A mobile build from eleven weeks ago still ships a document selecting it. From the moment the new schema is live, every request from that build fails **validation** — the field does not exist on the type — so the response has `errors` and no `data`, and the app shows an empty benefits screen. What the dashboards show depends entirely on what they meter. A transport-outcome error rate may show nothing at all. A single lumped `errors` counter shows a rise with no explanation. The split counters show it precisely: `request_error` climbs while `field_error` and `ok` are steady, which says the shape of what clients are sending changed, not that a backend broke — and dimensioned by operation, it names the operation. That is a one-minute diagnosis instead of an afternoon. ## Two wrinkles worth naming **Do not dimension by error message.** Messages are free text, frequently interpolate identifiers, and are therefore unbounded as a label. Servers commonly place a stable classification code under an error entry's `extensions` — that is a widespread convention, not something the specification defines — and if yours does, that code is the safe dimension. If it does not, dimension by the failure class you can determine structurally. **Errors-as-data undercounts you.** A schema that models expected failures as result types — a union of a success payload and a typed failure payload — produces no `errors` entry at all when the failure occurs. Those requests land in your `ok` bucket, correctly, because execution succeeded. If a meaningful share of real failures is modelled that way, this metric alone will not see them and you need domain counters emitted where those payloads are constructed.

  • A mobile build still selects a field the schema removed. Which counter moves, and why is that diagnostic?
    The request-error counter, alone. The document now fails validation, so execution never starts and the response has no `data` entry — nothing your resolvers or downstreams did is involved. Seeing `request_error` climb while `field_error` and `ok` hold steady tells you the documents arriving changed rather than the backends behind them, and the operation dimension names which document. A single lumped error counter would have shown a rise and told you nothing about where to look.
  • Should the error counter be dimensioned by the error message?
    No. Messages are free text and often carry identifiers, so the label set is unbounded — the classic way a metric pipeline gets flooded. Many servers attach a stable classification code under an error entry's `extensions`, which is a widespread convention rather than a rule of the specification; where that exists it is the right dimension. Otherwise dimension by the structurally determinable class — request error versus field error — and leave the message to logs.
  • Your schema models expected failures as typed result payloads instead of errors entries. What does that do to this metric?
    It undercounts. When a failure is returned as a member of a result union, execution succeeded and the response has `data` and no `errors`, so the request lands in the `ok` bucket — which is structurally correct and semantically misleading. If a real share of failures is modelled that way, emit domain counters at the point the failure payload is constructed, and be explicit that the envelope-level rate measures execution health, not business outcomes.

saying these in an interview costs you the question

  • Builds error rate from transport outcomes only
  • Counts entries in the errors list as failed requests
  • Treats any errors entry as a total outage
  • Dimensions a counter by raw error message text
  • Assumes a present data entry always means success
  • Says validation failures still return partial data

context