skip to content

How do you bound the cost of GraphQL field usage collection without losing deletion evidence?

level: principalimportance: should knowfreq 38%

answer

  1. Volume scales with data, not traffic
  2. Deduplicate per request into a set
  3. Two streams, only one may be sampled
  4. Last-seen rows, bounded and kept forever
  5. Register client names, cap the labels

basics

~20 s

Split the signal in two. Capture the distinct coordinate set per execution unsampled and roll it into last-seen rows bounded by schema size; sample the volume and latency stream freely. Then bound the client dimension, because self-reported versions are unbounded.

solid answer

~50 s

The naive design emits one event per resolved field, which scales with data: a 19-level-deep document over a real-estate graph resolved 11,743 fields in a single request, most of them from a list that grew without a bound. Deduplicate per request into a **set** of coordinates and the record becomes bounded by schema size instead. Then separate the two questions the platform actually asks. *Is this coordinate alive* is boolean and rare-event sensitive, so it must be captured on 100% of executions and stored as a last-seen timestamp per coordinate, client and version - a table on the order of thousands of rows, kept forever. *How hot is this coordinate* is a distribution, so sample it and keep short retention. Finally, bound cardinality on the client dimension: self-reported versions grow without limit, so register client names at the edge and cap the label set with an explicit overflow bucket.

code

pseudocode · 15 lines
pseudocode
on field_resolved(info, request):
    coord = info.parentType.interned_coordinate(info.fieldName)  # no allocation
    request.coords.add(coord)                 # liveness: always, deduplicated
    if request.sampled:                       # volume: sampled
        counters[coord] += 1

on request_finished(request):
    client = (register.resolve(request.header("client-name")) or "unknown",
              request.header("client-version") or "unknown")
    for coord in request.coords:
        rollup.touch(coord, client, now())    # in-process, flushed per interval

on flush_interval():
    export_last_seen(rollup.drain())          # size = coordinates x clients
    export_counters(sampled_counters.drain()) # short retention

go deeper

for a junior

Take the headline away: recording every single field resolution produces far more data than recording each request, so collectors deduplicate to the distinct set of fields a request touched.

for a middle

Explain why the volume scales with data rather than traffic, and why a per-request set of coordinates is bounded by the size of the schema no matter how large the lists in the response are.

for a senior

Argue the split between an unsampled liveness signal and a sampled volume signal, size the last-seen table, and name the cardinality risk in the self-reported client-version dimension.

for a principal

Own the tradeoff end to end: what the graph pays continuously versus what it may conclude, who is allowed to enable sampling on which stream, whether unidentified callers are admitted at all, and how the whole argument changes for a graph with external consumers.

## The volume problem, stated honestly Field-level collection is the only telemetry in a GraphQL server whose volume is set by the *data*, not by the traffic. One request against a real-estate listings graph - a 19-level-deep document that walked a neighbourhood into its listings, each listing into a price history that had grown without a bound, and each price change into the agent who made it - resolved 11,743 fields. Emitting an event per resolution turns one request into four orders of magnitude more telemetry than a per-request log line, and the multiplier is chosen by whoever writes the query, not by you. So the first design decision is not about sampling. It is about the unit. ## Decision 1: the unit is a set, not a counter Deduplicate coordinates within a request. The unbounded price-history list then contributes exactly one entry for `PriceChange.amount`, no matter how many elements it had. The record's size is bounded by the number of distinct coordinates the schema declares - on the order of a thousand for a mid-sized graph - and is completely independent of traffic shape. This is not a compromise; for the question the data exists to answer it is lossless. "Was this coordinate resolved" is boolean per execution. ## Decision 2: split the two signals, because sampling is asymmetric There are two questions, and they have opposite statistical needs. - **Liveness** - is anything still resolving this coordinate, and who? Rare-event sensitive. Sampling is fatal: the coordinate you are trying to retire is by definition in the tail. - **Volume and latency** - how hot is this coordinate, what does it cost? A distribution. Sampling is fine and traditional. So run them as two streams. The liveness stream is captured on **100%** of executions and is cheap precisely because it is deduplicated and aggregated in-process. The volume stream can be sampled at whatever rate the telemetry budget allows. The failure mode this prevents is the one that actually happens: telemetry bills rise, someone turns on sampling globally, and every deletion decision made afterwards rests on evidence that quietly stopped being able to see rare usage. If the two streams are separate, that lever cannot be pulled by accident. ## Decision 3: store last-seen, not history For liveness you do not need a time series. You need, per `(coordinate, client name, client version)`, the timestamp it was last resolved. That is an **upsert into a bounded table**: schema coordinates times client builds. Even at 63 live client versions it is on the order of tens of thousands of rows, small enough to keep indefinitely. The payoff is that the observation window stops being a retention limit and becomes a query parameter. "Quiet for 87 days" and "quiet for 400 days" cost the same, and nobody has to notice that the raw pipeline only kept 30 days. ## Decision 4: bound the client dimension The coordinate dimension is bounded by the schema. The client dimension is not: client name and version arrive as self-reported headers, so cardinality grows with every build anyone ships, and a misconfigured caller can invent labels at will. The options trade precision against cardinality: - **Register client names** at the edge and map anything unrecognised to a single `unknown` label. This is the important one, because it caps the dimension without losing the identities you care about. - **Truncate the version** to major.minor. Cheap, but it costs exactly the precision you need to know which build to chase - and old builds are the whole reason the version is there. Prefer it only for the volume stream. - **Cap the label set with an explicit overflow bucket**, so cardinality has a ceiling and the overflow is visible rather than silently dropping labels. And measure attribution coverage as a first-class metric, because it bounds every claim the data can support. ## Decision 5: aggregate in-process, and pay attention to the hot path Export a rollup per instance per interval, not per request: the export volume becomes coordinates times clients per interval, independent of request rate. On the hot path itself, the hook runs once per resolved field, so avoid per-field allocation and string concatenation - intern the coordinate string once at schema build and add a pre-computed reference to the set. Budget it explicitly, measure it under a deep document, and treat the overhead as a number you can quote rather than an article of faith. ## The organisational tradeoff Underneath the mechanics is a decision only a lead makes: **you cannot delete on evidence you did not collect**, and the cost of collecting it is paid continuously by every service in the graph, while the benefit arrives sporadically as a schema someone is finally allowed to shrink. Framing that honestly - a small fixed observability tax against an unbounded accumulation of fields nobody dares remove - is what gets the unsampled liveness stream funded. The answer also changes with the graph's audience. For an internal graph you can require registered client identity and refuse traffic without it, and the data becomes near-perfect. For a partner or public graph you cannot; callers are anonymous by nature, the unknown bucket is permanently large, and no amount of collection will let usage data alone justify a removal. Knowing which world you are in - and saying so before anyone quotes a zero - is the judgement being tested.

  • What would you refuse to give up if the telemetry budget were cut in half?
    The unsampled liveness stream. It is the cheap half - deduplicated, aggregated in-process, bounded by schema size - and it is the only part that cannot be reconstructed later, because a rare resolution you did not record is gone. Cut sampling rate and retention on the volume and latency stream instead, and shorten raw execution retention to hours.
  • How do you keep the client-version dimension from exploding?
    Treat it as a registered label, not free text. Map client names through an allowlist at the edge so unrecognised callers collapse into one bucket, keep full version precision only on the liveness stream where the row count is bounded by builds, and truncate to major.minor on the sampled volume stream. Cap the label set with a visible overflow bucket so cardinality has a ceiling you can see.
  • How does the answer change for a graph with partner or public callers?
    Attribution stops being reliable, because you cannot compel an anonymous caller to identify itself. The unknown bucket stays permanently large, so usage data can support 'no first-party client reads this' and never 'nobody reads this'. Retirement then needs a different mechanism - a registered document set, contractual notice, or a deliberate, reversible removal - rather than a stronger telemetry pipeline.

saying these in an interview costs you the question

  • Emits one telemetry event per resolved field
  • Applies one global sampling rate to every stream
  • Lets self-reported client versions become unbounded labels
  • Stores raw executions instead of last-seen rows
  • Ignores the hook's per-field cost on the hot path
  • Assumes anonymous callers can ever be fully attributed

context