skip to content

What does GraphQL field usage analytics record that an HTTP access log cannot show?

level: juniorimportance: nice to knowfreq 26%

answer

  1. One route, one status, no detail
  2. Instrument execution, not the transport
  3. Type dot field, per resolution
  4. Deduplicate to a set per request
  5. Tag it with client name and version

basics

~20 s

Field usage analytics records the schema coordinates - Type.field pairs such as Listing.priceHistory - that an execution actually resolved, and which client resolved them. An HTTP access log sees one URL, one method and one status for every operation alike.

solid answer

~40 s

A GraphQL API is usually one route, so transport logs describe every request identically: one path, one method, a `200`, some latency. Field usage analytics closes that gap by instrumenting execution instead of transport. As the executor resolves fields, it records the schema coordinate of each one - the `Type.field` pair, for example `Listing.priceHistory` - normally deduplicated to the distinct set a single execution touched, and tagged with the operation's identity and with a client name and version read from request headers by convention. Rolled up, that gives a per-coordinate answer to the question nobody can otherwise answer: has anything resolved this field lately, and which client build did it. None of this is in the GraphQL specification; it is an observability practice servers implement through execution hooks.

code

json · 15 lines
json
{
  "documentHash": "a91c7f2e4b",
  "operationType": "query",
  "clientName": "web",
  "clientVersion": "6.14.2",
  "resolvedCoordinates": [
    "Query.listing",
    "Listing.address",
    "Listing.askingPrice",
    "Listing.priceHistory",
    "PriceChange.amount",
    "PriceChange.changedAt"
  ],
  "at": "2026-09-03T09:41:17Z"
}

go deeper

for a junior

Recall the shape of the answer: one HTTP route means transport logs cannot tell operations apart, so usage is collected during execution and keyed by Type.field. Be able to give one concrete coordinate as an example.

for a middle

Explain the mechanics: an execution hook reads the parent type and field name, a per-request set deduplicates them so an unbounded list contributes once, and a client name and version from request headers turn a count into an owner.

for a senior

Show that you know the limits of the signal you are collecting - unattributed traffic, cached responses and sampling - and that the hook sits on the hot path, so its per-field cost is a real budget item you have measured.

for a principal

Own the policy question: what fraction of traffic can identify itself, whether unidentified callers are allowed at all, and what the organisation is entitled to conclude from a coordinate that has not been resolved.

## One route hides everything An HTTP-shaped API leaves a usable trail in transport logs: distinct paths, distinct methods, distinct status codes. A GraphQL API generally does not. Nearly every operation arrives at one route, carries its real content in a request body, and comes back `200` whether it succeeded, half-succeeded or failed outright. In a real-estate listings graph, a log line for a request that read one field - `Listing.address` - is byte-for-byte indistinguishable from a log line for a 19-level-deep document that walked a listing into its neighbourhood, into the agency, into that agency's other listings and each of their price histories. The transport layer knows a request happened. It does not know what was asked for, and it certainly does not know what was answered. That is a problem the moment anyone asks the maintenance question: *is anything still using this field?* Without an answer, a schema only ever grows. ## The unit of measurement: a schema coordinate Field usage analytics counts **schema coordinates**, written `Type.field`. `Listing.priceHistory` names the `priceHistory` field declared on the object type `Listing`. Three things it deliberately is not: - It is **not the response key**. If a client writes `history: priceHistory`, the alias changes only the key in the JSON response; the coordinate resolved is still `Listing.priceHistory`. - It is **not the operation name**. One operation selects dozens of coordinates, and the name is chosen by the client. - It is **not a path**. `listing.priceHistory.amount` is a response path through one particular execution; `PriceChange.amount` is the coordinate, and it is the same coordinate wherever in the graph that type is reached. The coordinate space is finite and small: it is bounded by the size of the schema, not by traffic. A mid-sized listings graph might have on the order of 1,247 output-field coordinates, and that number changes only when someone edits the schema. ## What a collector actually records The instrumentation point is execution, not the router and not the parser. A server exposes some hook that fires around field resolution; the collector takes the parent type name and the field name from the resolution's info, and adds the coordinate to a per-request accumulator. The accumulator should be a **set**, not a counter. A single document over a list that grew without a bound - every price change ever recorded on a listing - can resolve `PriceChange.amount` once per element; one request in this graph resolved 11,743 fields in total. If the goal is to know whether a coordinate is alive, that whole request should contribute exactly one entry per distinct coordinate. Counters are a separate, more expensive question. At the end of the request the collector emits one record: the set of coordinates, a stable identity for the document (typically a hash), the operation type, a timestamp, and the client identity. ## Client attribution Attribution is what turns "something used it" into "we know who to talk to". The widespread convention is a pair of request headers carrying a **client name** and a **client version**, set by every first-party caller. This is a convention only - no specification defines those headers, and the exact header names differ from stack to stack. Both values are self-reported, so they are excellent for directing engineering effort ("version 6.14.2 of the web client is the only thing still reading it") and worthless as a security control, since any caller can send any name. Requests that arrive without those headers land in an `unknown` bucket. The size of that bucket is the honest measure of how far your usage data can be trusted: if 40% of executions are unattributed, no coordinate can be declared client-free. ## What it buys, and what it does not What it buys is a factual answer to a question that is otherwise answered by mailing list and hope: *which clients resolved `Listing.priceHistory` in the last 87 days?* A `@deprecated` directive with a reason tells clients to stop using a field; usage data is the only thing that tells you whether they did. What it does not buy is any of the following. It is not part of the GraphQL specification, which defines the type system, validation and the execution algorithm and says nothing about reporting. It is not free - a hook on every field resolution sits on the hot path, which is exactly why per-request deduplication matters. And zero recorded usage is evidence, not proof: cached responses, sampling, unattributed traffic and rarely-taken code paths can all produce a zero for a field that is very much alive.

  • Where in the server would you hook to collect this?
    In the execution instrumentation, around field resolution - not inside each resolver by hand, which would miss every default-resolved field and every field someone adds later. The hook reads the parent type name and field name, adds the coordinate to a per-request accumulator, and a second hook at request end emits one record. Collecting in the router or a proxy is not an option: at that layer the operation is an opaque body.
  • Why record the distinct set of coordinates per execution rather than a count of resolutions?
    Because resolution counts scale with data, not with schema. A single document that pages a list which grew without a bound can resolve one coordinate thousands of times, so counts are dominated by list sizes and the telemetry volume is unbounded. The maintenance question - is this coordinate alive - is boolean, so a set bounded by schema size answers it at a fraction of the cost. Keep counts as a separate, samplable signal if you want them.
  • Does the GraphQL specification define any of this?
    No. The specification covers the type system, document validation, the execution algorithm and the response shape. Usage reporting is entirely an implementation concern. The only spec-sanctioned place a server may return its own extra data alongside a response is the `extensions` entry, and even there the specification deliberately leaves the contents undefined.

An access log is the turnstile count at the museum door. Field usage analytics is a sensor in every gallery, telling you which rooms nobody has walked into all year.

saying these in an interview costs you the question

  • Claims the access log's URL reveals which fields ran
  • Thinks usage reporting is part of the GraphQL specification
  • Counts operations and calls that field usage
  • Assumes the client-supplied operation name identifies the client
  • Records every field resolution, ignoring unbounded lists
  • Treats an aliased field as a different schema coordinate

context