skip to content

Why is a client-supplied operation name a risky GraphQL metric dimension?

level: seniorimportance: should knowfreq 46%

answer

  1. Ask who authors that label
  2. Two problems here, not one
  3. Unbounded values, and unverified values
  4. The server can parse out the truth
  5. Hash the document, bucket the rest

basics

~20 s

Because the caller authors it. Nothing bounds the value set, so a client that generates a fresh name per build multiplies your time series; and nothing verifies it, so an expensive document can arrive labelled as a cheap operation.

solid answer

~50 s

The operation name is a token the client writes into its own document and, optionally, repeats in the request body. It is therefore **unbounded** and **unverified**. Unbounded: nothing stops a build tool appending a content hash, or a screen interpolating a record id, so a payroll graph with 137 real operations can present tens of thousands of distinct label values within days, and a per-label time series is created for each. Unverified: a name is a free-text label on the document, never checked against what the document selects, so a caller can post an 8,400-row reconciliation under the name `EmployeeBadge` and corrupt both series. Derive the dimension server-side instead: a hash over the normalised document, or the persisted document identifier where one exists. Keep the client-declared name only as a hint, matched against an allowlist and bucketed to a single `unknown` value otherwise.

code

json · 5 lines
json
{
  "operationName": "EmployeeBadge",
  "query": "query EmployeeBadge($runId: ID!) { payRun(id: $runId) { payslips(first: 8400) { edges { node { deductionLines { amountCents benefitPlan { code } } } } } } }",
  "variables": { "runId": "run-2026-08-B" }
}

go deeper

for a junior

Know that the operation name is written by the client in its own document, not assigned by the server. That single fact is what makes it questionable as a label, and interviewers are checking you know where the value comes from.

for a middle

Explain both failure modes concretely: unbounded values create a series per distinct string, and unverified values mean the name need not describe what the document selects. Name the server-derived alternatives.

for a senior

Show you would ship the mitigation. Talk about canonical-printed document hashes, an allowlist with a single unknown bucket, treating a rise in unknown as a signal, and every downstream use — baselines, attribution, limits — that inherits a forged label.

for a principal

Own the policy: who registers operation names, whether documents are persisted, and what the organisation's cardinality budget for this signal is. Decide it before a client release quietly makes it a bill.

## Where the operation name comes from It helps to be exact about the artefact. An executable document may name its operations — `query EmployeeBadge($id: ID!) { ... }` — and a request may also carry an `operationName` selecting which of them to run. That selection is required when the document defines more than one operation. Both halves are written by the caller. The server checks that the named operation exists in the document; it does not, and cannot, check that the name is *apt*. So when someone proposes labelling request metrics with the operation name, the honest description is: we are letting every caller in the world write directly into our metric label set. Two distinct failures follow, and a strong answer names both. ## Failure one: the label set is unbounded Metric dimensions are only useful when the number of distinct values is small and stable, because each combination of values creates and retains a series. A hand-written GraphQL client is well behaved here — a payroll and benefits app might have 137 named operations across its screens, and that set changes at release cadence. But nothing enforces that discipline, and the ways it breaks are mundane rather than malicious: * A build step appends a content hash to every operation name for cache-busting, so `PayslipList` becomes `PayslipList_a91f3c` and every release invents a full generation of new series. * A screen builds its document text by string interpolation and lets an identifier leak into the name. * A third-party integration, an internal script or an automated scanner sends arbitrary documents with arbitrary names. One real shape of this incident: a graph with 137 registered operations observed 19,412 distinct values on that label inside 72 hours. Nothing was under attack; a client had shipped per-render names. The cost lands somewhere unpleasant — the metric pipeline, the storage bill, query latency on every dashboard built over that metric — and, worse, the metric stops being useful long before it stops working, because no single series has enough traffic to have a baseline. ## Failure two: the label is unverified This one is subtler and is the part candidates usually miss. Even bounded, the name is a *claim*, not an observation. A document named `EmployeeBadge` may select an 8,400-row page of payslips and expand each one's deduction lines. The server will happily execute it and your metric will happily attribute several seconds and a large amount of backend work to the badge operation. The consequences are not only cosmetic. Every downstream use of that dimension inherits the lie: per-operation latency baselines, cost attribution to a team, an allowlist keyed by name, rate limits per operation, and any objective computed per operation. Someone probing the endpoint does not even need bad intent to cause it — a copied-and-edited document that kept the old name does the same damage. Treat the name as caller-controlled input, because it is. ## What to dimension by instead Rank the candidates by who produces them. **Operation type** — `query`, `mutation`, `subscription`. The server derives it by parsing the document, so it is authoritative, and it has exactly three values. Always safe, always worth having. **A server-computed document hash.** Hash the document *after* parsing and printing it back in a canonical form, so that whitespace and formatting differences collapse to one value. This is a fact about what actually ran and cannot be forged: change the selection and you change the hash. It is not automatically bounded — a client that interpolates literals into query text produces a new hash per request — but it converts an unbounded label into an observable property of your traffic, and a spike in distinct hashes is itself a useful signal. **A persisted document identifier.** Where documents are registered ahead of time, or exchanged for a hash through the Automatic Persisted Queries handshake, the identifier the server already keys on is the natural dimension, and the value set is exactly the set of documents you have accepted. This is the strongest option and it is one of the underrated reasons to adopt persisted documents at all. **The client-declared name, sanitised.** Keep it — it is the only human-readable handle and it is what an engineer will search for — but pass it through a bounded map: if the name is in the known set for that document, use it; otherwise emit a single literal `unknown`. Then `unknown` becomes a first-class signal. A sudden rise in it means a new client, a broken build, or someone poking at the endpoint, and it costs you exactly one series to detect. A good final answer is a pairing: hash for identity, sanitised name for legibility, and the honest sentence that the name never travels alone.

  • Is the operation type a safe dimension by the same test?
    Yes, and it is the cleanest example of the test passing. The server determines whether the operation is a query, a mutation or a subscription by parsing the document it is about to execute, so the value is observed rather than claimed, and the value set is exactly three. It is also genuinely useful: subscriptions are long-lived and do not belong in the same request-duration series as request-response operations.
  • Your allowlist holds 137 names. What happens to the 138th?
    It becomes the literal value `unknown` — one extra series, not one per stranger. That keeps the label set bounded while turning the overflow into a signal in its own right: a rise in `unknown` means a client shipped something you have not registered, a build changed its naming, or someone is probing the endpoint. Passing the raw value through instead is exactly how a metric pipeline gets flooded by a single misbehaving caller.
  • Does hashing the document fully solve the cardinality problem?
    No, it bounds it by client behaviour rather than by policy. If documents are authored at build time, the distinct-hash count equals the number of shipped documents and is small. If a client interpolates literal values into query text instead of using variables, every request is a new document and a new hash. The difference is that the hash is now measurable — distinct-document count becomes a monitorable property — and persisted documents make the bound explicit.

saying these in an interview costs you the question

  • Treats the operation name as server-verified
  • Lets any caller string become a metric label
  • Assumes an operation name identifies what ran
  • Embeds a build id in every operation name
  • Blames the metrics backend for cardinality growth
  • Thinks a document hash survives any reformatting

context