Your GraphQL endpoint passes every test but times out in production — how do you prove field fan-out is the cause?
answer
- One row hides a doubling
- Compare a field's count with its parent's
- The ratio equals the list length
- No single call is slow
- Budget backend calls per named operation
basics
~20 sCount resolver invocations per field path for one real request and compare each field's count with its parent's. A ratio matching the list length — not a slow single call — proves per-item fan-out, which one-row fixtures cannot reveal.
solid answer
~50 sStart from the observation that a one-row fixture makes 1+N equal two calls, so the tests were never capable of seeing this. Capture one real production request and tally resolver invocations by field path: if the enrolment's `student` field shows 564 invocations against 47 course parents and one term, the fan-out is arithmetic, not opinion. Corroborate from the data store's own statement log, where the same statement text repeats within one request id with only the bound key changing, and from the latency shape — total time tracks the requested page size rather than the document's length, and no single call is slow. Then reproduce it with a multi-row fixture and pin it with a per-operation call-count budget so the regression cannot return silently. The remedy itself is request-scoped batching, but the diagnosis is the count.
code
pseudocode · 11 lineson_resolver_start(requestId, fieldPath):
counters[requestId][fieldPath] += 1
on_request_end(requestId):
log(requestId, operationName, counters[requestId])
# term 1
# term.courses 1
# course.instructor 47 <- 47 / 1 = course page size
# course.enrolments 47
# enrolment.student 564 <- 564 / 47 = enrolment page sizego deeper
Recall that one row in a fixture makes the problem invisible, because 1+N is just two calls. Know that a passing test proves the response was right, not that the server was cheap to run.
Be ready to describe the evidence: per-field invocation counts for one request, and the ratio between a field's count and its parent's. Explain why the data store's statement log showing one repeated statement with varying keys says the same thing.
Demonstrate a full diagnosis under pressure — counts, latency distribution, per-operation attribution — and then the regression barrier: multi-row fixtures and a per-operation call-count budget. Interviewers look for the discipline of separating a slow query from a fan-out.
Own the systemic angle: cost is untested by default because suites assert responses, and one URL destroys route-level attribution. Be ready to argue for call-count budgets as a platform default and for how a single-endpoint API should be observed at all.
## Why the tests were never going to catch it With one course in the fixture, a 1+N field costs two calls. With two courses it costs three. Nothing about those numbers looks wrong, and an assertion on the response body — which is what integration tests almost always assert — is identical whether the server made 2 calls or 660. The test suite is not weak here; it is measuring the wrong dimension. Correctness of the response and cost of producing it are independent properties, and only one of them was ever asserted. Production differs in exactly the variable that matters: the page size real callers pass. A request for 47 courses with 12 enrolments each puts 564 enrolment objects in play, so a single backend-resolved field on the enrolment type runs 564 times. ## Proving it, in order **1. Tally resolver invocations by field path for one request.** This is the direct evidence and everything else is corroboration. Wrap resolution with a per-request counter keyed by the field's path, and print the tally at the end of the request: ``` term 1 term.courses 1 course.instructor 47 course.enrolments 47 enrolment.student 564 ``` The proof is not the size of any one number — it is the *ratio between adjacent rows*. 564/47 = 12 is the enrolment page size, and 47/1 = 47 is the course page size. A field whose count equals its parent's count times a list size is being resolved per item, by definition. **2. Read the data store's statement log for one request.** Correlate on the request id and look for the same statement text repeated with only the bound key varying. Hundreds of identical single-key lookups inside one request is the same finding from the other side of the wire, and it is the version that convinces a data-store owner. **3. Check the latency shape.** Per-item fan-out has a signature: no individual call is slow, the total is roughly the number of calls times the per-call latency, and the total moves with the requested page size rather than with anything about the document text. At a measured 62 ms per call, 660 calls is 41 seconds of backend work; against a pool of 8 connections that is around five seconds of wall clock, which is exactly the sort of number that trips a gateway timeout while every individual span looks healthy. If instead one span dominates and the rest are trivial, this is a slow query, not fan-out — a different problem with a different fix. **4. Attribute it to an operation, not to the endpoint.** All operations arrive at one URL, so endpoint-level metrics are useless for this. Break the counters down by operation name so you can say *which* client document is responsible; usually one or two are, and the rest are innocent. ## The regression that has no document change A case worth carrying into an interview, because it defeats the first instinct of blaming the caller. A deploy changed `Course.instructor` from `Instructor` to `Instructor!`. The old resolver returned null immediately for self-paced courses, which had no instructor — roughly three in four — and made no backend call at all for them. Once the field was non-null, returning null would have propagated an error over the whole course, so the resolver was changed to look up a placeholder record instead. Per-request instructor lookups went from about 12 to 47 on the identical document with the identical page size. The lesson is that fan-out is a property of *resolvers plus runtime shape*, not of the document. Nothing the caller sent changed. The schema change removed an early return, and a cheap field became a fully amplifying one. When you are hunting a regression, diff the resolvers and the schema alongside the client documents. ## Making it stay fixed Diagnosis without a regression barrier just books the same incident again later: - **Fixtures with more than one row.** Any list-returning field in a test fixture should hold at least three or four items. This single change makes 1+N visible as 5 calls versus 1 rather than 2 versus 1. - **A call-count budget per operation.** Assert in the test that a named operation causes at most *k* backend calls. It is the only assertion that fails when someone adds an innocent-looking nested field, and it fails loudly at review time rather than quietly in production. - **Ship the per-field counters.** Keep the tally in production behind a sampling flag so the next investigation starts with evidence rather than a rebuild. ## What not to conclude Two wrong turns are common. The first is capping the page size and declaring victory: that reduces the multiplier without touching the per-item behaviour, so the same document at a larger page brings the incident straight back. The second is blaming the client team for a greedy document. The document is a legitimate use of a schema that offered those fields; the server chose to resolve them one parent at a time. The fix belongs on the server, in request-scoped batching, and the counting above is what tells you which field to put it on first — almost always the deepest one, since that is where the product of the list sizes is largest.
- The trace shows one slow span and a few hundred fast ones. Is that fan-out?Not primarily. Fan-out's signature is many uniformly cheap calls whose sum is the latency, with no outlier. One dominant span points at a slow individual query — a missing index or a bad plan — and should be chased on its own terms. The two coexist often enough that you should measure both the count and the distribution before choosing a fix.
- Latency doubled overnight with no client release and no document change. Where do you look?At the server side of the shape. Fan-out is resolvers plus runtime data, so a schema or resolver change can remove an early return and turn a mostly-free field into a per-parent lookup, and a data change can grow the lists themselves. Diff the resolvers and the schema against the previous deploy, then compare per-field invocation counts across the two versions on the same document.
- Why not just cap the page size and close the incident?A cap lowers the multiplier without changing the per-item behaviour, so the endpoint is still 1+N and the incident returns as soon as a caller needs a larger page or someone nests one more level. It is a legitimate short-term mitigation to stop the bleeding, but reporting it as the fix hides an unbounded shape behind a number.
- How do you attribute the cost when every operation arrives at the same URL?Break your counters and timers down by operation name rather than by route, since the route is identical for everything. Requiring named operations from clients makes this reliable, and the per-operation view usually shows one or two documents responsible for most of the backend calls while the rest are unremarkable.
saying these in an interview costs you the question
- Trusts a single-row fixture to represent production cost
- Asserts only response shape and never backend call counts
- Blames the client's document instead of per-item resolution
- Calls a page-size cap the fix rather than a mitigation
- Uses endpoint-level metrics where every operation shares one URL
- Assumes fan-out means one slow query somewhere