Why is one GraphQL endpoint's overall request latency metric not actionable?
answer
- One route, many workloads
- A path-based API got this free
- The number follows the traffic mix
- A rare heavy operation hides
- Dimension by operation, not route
basics
~20 sBecause every operation shares one route, so the metric blends unrelated workloads: a few-millisecond name lookup and a multi-second report land in the same series. The number tracks the traffic mix, not any operation's health.
solid answer
~50 sA path-per-resource HTTP API gets a dimension for free. `/employees/{id}` and `/pay-runs/{id}/reconciliation` are separate series, so latency per route means something. A GraphQL API usually exposes **one** route for every operation, so the route dimension collapses to a constant and the latency series becomes a mixture of every workload the schema permits. In a payroll and benefits graph, an `EmployeeBadge` operation reading two scalars off one row finishes in single-digit milliseconds, while a `PayRunReconciliation` operation walking an 8,400-row page of payslips into their deduction lines takes seconds. Mixed together, the endpoint p99 mostly reports what share of traffic was heavy. It moves when the mix shifts although nothing got slower, and it stays flat when the heavy operation doubles because that operation is rare. The fix is to stop metering the route and start metering the operation.
code
graphql · 22 linesquery EmployeeBadge($id: ID!) {
employee(id: $id) {
displayName
photoRef
}
}
query PayRunReconciliation($runId: ID!) {
payRun(id: $runId) {
payslips(first: 8400) {
edges {
node {
netPayCents
deductionLines {
amountCents
benefitPlan { code employerShareCents }
}
}
}
}
}
}go deeper
Be ready to say why a single POST route removes the dimension a path-based API gets free, and to name two operations of wildly different cost sharing it. Interviewers use this to check you understand what GraphQL changes about monitoring.
Explain the mechanics both ways: how a traffic-mix shift moves an aggregate percentile with no regression, and how a rare expensive operation can degrade badly while the aggregate stays flat. Name the dimensions that fix it.
Show operational judgement about what to page on. Talk about per-operation baselines, keeping subscriptions out of a request-duration series, splitting parse-and-validate from execution, and why per-field metrics belong to tracing rather than to your metric pipeline.
Own the cost and cardinality budget for the whole signal. Decide which grain the organisation meters at, what it refuses to meter, and how teams get per-operation baselines without every client release inventing new series.
## The dimension a single route takes away Every useful latency metric is really a latency metric *per something*. In a path-per-resource HTTP API, that something arrives free of charge: the request line already names the resource, so a proxy, a framework filter or a server-side timer can partition durations by path with no application knowledge at all. Two routes that do different amounts of work produce two series, and each series describes one workload. A GraphQL API deliberately throws that away. The whole design is that a caller composes what it wants and posts it to a single endpoint. The consequence for observability is immediate and often unremarked: the route dimension still exists, but it now has exactly one value, so partitioning by it partitions nothing. Whatever your timer wraps, it wraps *every* operation the schema permits. ## What that mixture actually looks like Take a payroll and benefits graph. Two of its operations sit at opposite ends of the cost range: * `EmployeeBadge` reads a display name and a photo reference off one row. Call it 6 ms at p50, and roughly 91% of the traffic, because it renders in a header on every screen. * `PayRunReconciliation` pages 8,400 payslips and, for each, expands the deduction lines and the benefit plan behind them. Call it 3.2 seconds at p50, and 0.3% of the traffic, because a payroll administrator runs it a few times per cycle. Now read the blended p99 of that endpoint. With the heavy operation at 0.3% of traffic, it sits below the 99th percentile entirely, so p99 is somewhere in the tens of milliseconds and describes the badge lookup's tail. The report is invisible. Two things follow, and both are the reason interviewers ask this question. **The number moves for reasons that are not regressions.** Move the reconciliation operation from 0.3% to 1.4% of requests — a business-calendar effect, month-end, nothing shipped — and it crosses into the top percentile. The endpoint p99 jumps from 40-odd milliseconds to seconds. The dashboard looks like an outage. No operation got slower; the mixture changed. This is the single most common false page on a GraphQL service. **The number stays still for reasons that are regressions.** Let the reconciliation operation degrade from 3.2 seconds to 9 seconds because a deduction-line lookup lost an index. At 0.3% of traffic it never reaches p99, so the endpoint metric does not twitch, while the administrators who depend on it are timing out. A rare, expensive, business-critical operation is exactly the kind the aggregate hides. ## What to measure instead Measure the **operation**, not the route. The minimum useful dimension set for a request-duration metric on a GraphQL endpoint is: * **operation type** — `query`, `mutation` or `subscription`. The server parses this out of the document itself, so it is authoritative and bounded to three values. It matters because subscriptions are long-lived and do not belong in the same duration series as request/response operations at all. * **operation identity** — which document ran. The client-declared operation name is the obvious candidate and the treacherous one; a server-computed hash of the document is the trustworthy one. That trade-off is a question of its own. With those, `EmployeeBadge` and `PayRunReconciliation` become separate series with separate baselines, and both failure modes above disappear: a mix shift changes the request *rate* of each series without touching either one's latency, and a regression in the rare operation shows up in that operation's own line. ## Two things not to do Do not reach for a metric per resolved field. One reconciliation document resolves tens of thousands of field positions; a counter or timer per field position is a cardinality problem and a cost problem, and the question it answers — where inside one execution the time went — is what per-resolver tracing exists for. Metrics answer "which operations, how often, how fast, how often failing"; that is a much smaller question and should stay small. Do not report a mean. The distribution of durations on a GraphQL endpoint is not merely skewed, it is a mixture of distributions that have nothing to do with each other, and a mean over a mixture is a number with no referent. Even after you split by operation, prefer a distribution over a mean — but the split matters far more than the statistic. ## Where the timer starts and stops One detail worth having an answer for: a GraphQL request is parsed, validated and then executed, and a failure at the first two stages costs almost nothing. If your timer covers all three, a flood of invalid documents will *lower* your latency while raising your error rate. Recording parse-plus-validate separately from execution keeps that honest. And when a response is delivered incrementally, "the request finished" has two candidate meanings — the initial payload and the final one — so pick one, name it in the metric, and be consistent.
- The endpoint p99 doubled overnight with no deploy. What do you check first?The mix, before the code. Break the same window down by operation and compare request rates rather than durations: if every operation's own latency line is flat and only the share of an expensive operation grew, the endpoint number moved for arithmetic reasons and nothing regressed. Month-end, a batch job, a new screen shipping a heavier document, or a retry storm on one operation all produce this. If you cannot do that breakdown, that is the finding.
- Why not simply record a duration per resolved field instead?Because the cardinality and the cost are wrong for metrics. A single deep document resolves tens of thousands of field positions, so a per-field timer means an enormous number of series and measurable overhead on every execution. The question "where inside this one execution did the time go" is answered by sampled per-resolver tracing, not by an always-on metric. Keep metrics to the operation grain and let tracing go deeper on a fraction of requests.
- Should parse and validation time sit inside the same timer as execution?Better separately. Parsing and validation are cheap and happen before any resolver runs, so a burst of invalid documents would otherwise pull your latency distribution downward while your failure count climbs — two signals moving in opposite directions for one cause. Recording parse-plus-validate and execution as two durations under the same dimensions keeps a validation flood legible and shows when document size, not backend work, is the cost.
One route is a till receipt with a single line reading "shopping". A pack of chewing gum and a chest freezer both went through it, so the average tells you about the trip, never about an item.
saying these in an interview costs you the question
- Calls the endpoint p99 the API's p99
- Blames a deploy when only the traffic mix moved
- Adds a metric per resolved field position
- Assumes one route means one workload
- Reports a mean over a mixture of workloads
- Thinks a flat endpoint metric means nothing regressed