skip to content

Per-field GraphQL tracing emits thousands of spans per document; how do you afford that in production?

level: seniorimportance: should knowfreq 48%

answer

  1. Count the spans one list produces
  2. Most fields do almost no work
  3. Decide once per request, not per field
  4. Group by schema coordinate, not by path

basics

~20 s

Sample whole requests rather than individual fields, so a kept trace is complete. Skip spans for fields with no real resolver, collapse sibling list items into one aggregated span, and take steady-state numbers from per-coordinate aggregates instead of spans.

solid answer

~40 s

Start by counting: a span per resolved field means a seat-map read over 1,847 seats emits five figures of spans, most of them fields whose resolution is a property read off the parent. Four levers. **Sample the request, not the field** — the decision is made once per operation and inherited by every span, because half a trace answers nothing. **Do not instrument trivial fields**: span only those with a non-default resolver. **Collapse list siblings** into one span carrying a count, a total and a maximum. **Aggregate by schema coordinate for the steady state** — `Row.seats` is bounded, a response path with list indices is not. Then measure the instrumented-versus-not delta: a clock read and an allocation per field is work inside the hot loop even when the trace is dropped.

code

pseudocode · 22 lines
pseudocode
on_request_start(request):
    request.traced = sampler.keep(document_hash(request))
    request.buffer = []

on_field_start(request, field, path):
    if not request.traced:
        return NO_SPAN
    if field.resolver is DEFAULT_RESOLVER:
        return NO_SPAN                  # a property read; not worth a clock
    return open_span(path, field.coordinate)

on_field_end(request, span, duration):
    if span is NO_SPAN:
        return
    if is_list_item(span.path):
        request.buffer.aggregate(span.parent_path, duration)
    else:
        request.buffer.add(span)

on_request_end(request, elapsed, had_field_errors):
    if request.traced or elapsed > SLOW_MS or had_field_errors:
        export(request.buffer)

go deeper

for a junior

Be ready to say why a span for every resolved field is expensive at all: the count follows the size of the result, not the size of the document, so one ordinary page can produce thousands of them.

for a middle

Explain the mechanics of the cheap levers — why the sampling decision belongs to the whole request, and why a field resolved by the default behaviour costs more to instrument than to execute.

for a senior

Show that you would measure first. Count the spans a real document emits, benchmark the instrumentation overhead separately from the export bill, and keep slow and errored requests regardless of the sample.

for a principal

Own the budget and the split. Decide what is always-on and bounded versus sampled and detailed, set one sampling contract that every service in the graph honours, and be explicit about the visibility you are trading away.

## Where the volume comes from The span count of a per-field trace is not proportional to the size of the document. It is proportional to the size of the **result**, because the executor resolves each selected field once per parent object it is resolved against. A ticketing read that walks a seat map for one performance selects perhaps a dozen fields, and resolves 14 sections, 644 rows and 1,847 seats; at four selected fields per seat that is over 8,000 field resolutions and therefore over 8,000 spans, for what a user sees as one page. Add a list of upcoming performances to the same document and it multiplies again. That is the shape of the problem: not a few expensive spans, but an enormous number of nearly free ones. And the export cost, the storage cost and the query cost of a tracing backend all scale with span count, not with how interesting the spans are. ## Lever one: the sampling unit is the request The granularity of the keep-or-drop decision is the whole operation. A trace missing an arbitrary 90% of its fields is not a cheap trace, it is a useless one — the parent-child structure is broken, self times cannot be derived, and the field you are hunting is probably one of the missing ones. The decision is taken once, at the start of the request, and every span inside inherits it. The GraphQL-specific refinement is what you key the rate on. A flat rate across one endpoint is dominated by whatever operation runs most often, and the rare expensive report query — which is exactly the one you want traced — shows up in the sample almost never. Keying the rate on the operation identity (the document hash is the safe key; a client-supplied operation name is unbounded and forgeable) lets you sample the hot lookup at a very low rate and the rare heavy document at a very high one. A complementary rule is to buffer spans for the request and always emit when the request was slow or produced field errors, which costs memory per in-flight request in exchange for keeping the traces that matter. ## Lever two: most fields do not deserve a span Most fields in a large result are resolved by the default behaviour: look up a key on the parent object and return it. That resolution is nanoseconds of real work; wrapping it in a span costs a clock read, an allocation and an export payload, so the instrumentation is orders of magnitude more expensive than the thing it is measuring, and it contributes nothing but noise to the trace. Two filters work well. Instrument only fields whose resolver is not the default one — the server knows which those are when it builds the executable schema, so this is a build-time decision, not a per-request check. Or record intervals for everything but only *emit* spans above a threshold, keeping the sub-millisecond ones folded into their parent's self time. The first is cheaper and coarser; the second catches the field that is usually trivial and occasionally is not. ## Lever three: collapse the siblings The 1,847 spans on `performance.seatMap.sections.*.rows.*.seats.*.status` differ only in their list indices. Their individual identities almost never matter; their distribution always does. Emitting one span per sibling set, carrying a count, a total, a maximum and perhaps the path of the slowest member, keeps every question you actually ask of the data and removes three orders of magnitude of volume. This is also the fix for the read-the-numbers trap where a shared batched fetch parks its whole wait on one arbitrary sibling: the aggregate shows the true total for the set. ## Lever four: aggregates for the steady state, spans for the investigation Spans are for answering "why was *this* request slow". They are a poor substrate for "is `Row.seats` getting slower over the quarter". The distinction that matters is the key: a response path with list indices in it is an unbounded key, and no store will thank you for grouping on it, whereas the **schema coordinate** — the parent type and field name, `Row.seats` — is bounded by the size of the schema. Per-coordinate duration aggregates can run at 100% on every request at a fraction of the cost of tracing, and the sampled field traces exist to explain what the aggregates surface. ## Do not forget the on-CPU cost Even a span you throw away costs you. The tracer sits inside the executor's per-field path, so it pays a clock read and usually an allocation for every one of those 8,000 fields. Benchmark the same document with instrumentation disabled versus enabled-and-dropped; if the delta is significant, the fix is the trivial-field filter, not a lower sample rate, because sampling reduces export and storage but not the per-field work of deciding. ## And it compounds across services In a supergraph composed from 62 subgraphs, every subgraph is resolving fields and every subgraph is capable of tracing them, so one client operation can fan out into per-field traces in a dozen services at once. The rule that saves you is that the sampling decision is made once for the client operation and carried, so the whole distributed trace is kept or dropped together — independent per-service sampling produces a large volume of partial traces, which is the worst of both outcomes. ## What a strong answer sounds like Count the spans before proposing a fix; sample by request and key the rate on operation identity; stop instrumenting fields that do no work; aggregate sibling list items; keep bounded per-coordinate aggregates always on and reserve full field traces for a sample and for slow or errored requests; and measure the instrumentation's own overhead, not just the backend bill.

  • You drop the sample rate to a tenth and the endpoint's latency does not improve. Why not?
    Because sampling governs export and storage, not the work done inside the executor. If the instrumentation still opens and closes an interval for every resolved field in order to decide, the per-field clock reads and allocations are unchanged — only the payload leaving the process shrinks. The fix is to make the keep decision at the start of the request and skip span creation entirely for dropped requests, plus stop instrumenting fields with no real resolver.
  • Why key a sampling rate on a document hash rather than the operation name a client sends?
    The operation name is chosen by the caller, so it is unbounded and spoofable — two clients can send different documents under the same name, or a thousand names for one document, and either wrecks the rate you thought you configured. A hash of the normalised document text identifies the actual operation, is stable across clients, and cannot be steered by a caller wanting more or less of your tracing budget.
  • What do you lose by collapsing 1,847 sibling spans into one aggregate?
    The ability to point at a specific item after the fact — if one particular seat's fetch was pathological, the aggregate shows a maximum but may not tell you which one. Carrying the slowest member's response path on the aggregate recovers most of that at negligible cost. What you do not lose is the total, the count or the distribution shape, which is what nearly every latency question about a list actually needs.

saying these in an interview costs you the question

  • Samples individual fields, leaving broken partial traces
  • Spans every field including plain property reads
  • Groups long-term aggregates by response path with indices
  • Thinks a lower sample rate removes the in-process overhead
  • Lets each service sample the same operation independently
  • Uses a client-supplied operation name as the rate key

context