skip to content

A subgraph gets one entity call carrying 4,812 representations. How do you bound that fan-out?

level: seniorimportance: should knowfreq 42%

answer

  1. Call count looked fine, size did not
  2. Two axes: how many, how big
  3. Cap the list that feeds it
  4. Split one huge call into bounded ones
  5. Alarm on p99 representation count

basics

~20 s

Bound the list that produces it: require a paginated argument with an enforced maximum, and score document cost at the router before planning. Then chunk oversized representation lists and let the subgraph reject calls above its own cap.

solid answer

~50 s

Batching converted an unbounded number of calls into one unbounded call, and the second problem hides better — the router's call count looks perfect. A 4,812-key fetch risks body-size and memory limits, a database `IN` list no planner enjoys, and an all-or-nothing failure that erases a whole branch with nothing cheap to retry. Fix it at the cause first: the list field feeding the fetch accepts `first: 5000` or no pagination at all, so require pagination and enforce a maximum page size, rejecting rather than truncating. Add a static cost or complexity limit at the router, which is the only place that sees the whole document before any subgraph is contacted — a widespread convention, not a specified rule. Then chunk the representation list into bounded calls, cap representation count inside the subgraph, and alarm on a p99 of entity-fetch size rather than on call rate.

code

pseudocode · 8 lines
pseudocode
MAX_PER_FETCH = 256
MAX_CONCURRENT = 4

function fetchEntities(subgraph, representations):
    chunks = split(representations, MAX_PER_FETCH)
    results = runWithLimit(MAX_CONCURRENT, chunks, chunk ->
        subgraph.call("_entities", chunk))
    return concatInInputOrder(results)

go deeper

for a junior

Recall that gathering many keys into one call is the fix for call count, and that the resulting call can itself grow large. Knowing page sizes should have a maximum is enough at this level.

for a middle

Explain the mechanics of the trade: one big call means one parse, one memory spike, one IN list and one all-or-nothing failure, whereas many small calls cost round trips but fail independently.

for a senior

Demonstrate that you have operated this. Name which limit sits at which layer, what the client sees when each trips, and the size-histogram metric that reveals amplification hidden inside a single successful request.

for a principal

Own the policy question: who is allowed to set page caps and cost budgets across independently owned subgraphs, how bulk-export use cases get a sanctioned path instead of abusing the interactive graph, and how limits are rolled out without breaking existing clients.

## How a 4,812-key fetch happens Batching cross-service calls turns an unbounded *number* of requests into an unbounded *size* of one request. Both are amplification; only the second one looks fine in a request-count dashboard. In a museum collection graph, a bulk-export client asks a **Catalog** subgraph for every artwork in a department and, for each, the condition report owned by a separate **Conservation** subgraph: ```graphql query DepartmentExport { department(code: "PAINT") { artworks(first: 5000) { accessionNumber conservation { status lastInspected } } } } ``` After deduplication the router builds one entity fetch carrying 4,812 representations. That is one HTTP call — the router's per-step call count is a healthy 1 — containing several hundred kilobytes of keys, asking one service to resolve nearly five thousand entities inside a single request whose failure is all-or-nothing. ## What actually breaks *Memory and body limits.* The representation list is parsed into memory on the subgraph side, and the response holds 4,812 entity objects at once. A body-size limit rejects the request outright; without one, a few concurrent exports are a heap problem. *Backend limits.* A batched lookup of 4,812 ids is frequently an `IN` list a database will not plan well, or refuses beyond a driver limit. The subgraph's internal batching quietly becomes the bottleneck it was supposed to remove. *Blast radius.* One hundred small calls that each fail independently degrade a page. One large call that fails erases the entire branch of the response — and if the field is non-null, the erasure climbs further. There is no partial progress and nothing useful to retry cheaply: a retry re-does all 4,812. *Tail latency owns the plan step.* The plan cannot proceed past that step until the whole call returns, so the step's latency is the slowest single request rather than the median of many. ## Where to put the bounds Work outwards from the cause. **Cap the list that feeds it.** Amplification size is a function of the list field above the entity fetch. A list field that accepts `first: 5000` — or accepts no pagination argument at all — is the actual defect. Require a paginated argument, enforce a maximum page size in the resolver, and reject rather than silently truncate, so clients discover the limit. **Score the document before planning.** A static cost or complexity limit evaluated at the router, before any subgraph is contacted, is the only control that sees the whole document. It multiplies list sizes down the nesting to a score and rejects the request outright. This is a widespread convention with several published formulations, not a specified part of GraphQL. **Chunk the representation list.** Where the router can split one oversized entity fetch into several bounded calls — say 256 keys each, issued with limited concurrency — you get back independent failure domains, retryable units and predictable memory, at the cost of more round trips. It is a tuning decision: too small and you have re-created the amplification you removed. **Defend inside the subgraph.** `_entities` is a real entry point, reachable by anything that can talk to the subgraph. A subgraph should enforce its own maximum representation count and its own internal chunking of database loads, independent of what the router promises. Treat the router as an untrusted caller for size purposes, because a misconfiguration upstream should not be able to OOM a service. **Measure the right thing.** The metric to graph is not calls per second, it is a histogram of representation-list size per subgraph per operation, with the 99th percentile alarmed. A p50 of 12 and a p99 of 4,812 is the signature of one client doing something the schema permits, and it is invisible in averages. ## The silent-failure shape The nastiest presentation is a federated subscription. Each event payload is resolved through the same plan as a query, so a subscription over a museum's conservation feed does an entity fetch *per event*. When one event's list grows past a limit, that event's payload fails to resolve — while the connection stays open and healthy. Curators reported updates "stopped"; the connection metric showed a live stream, and the errors were arriving inside payloads the client discarded. The diagnostic that resolved it was the entity-fetch size histogram, not the connection dashboard. ## The judgement being tested A senior answer does not just say "add a limit". It names *which* limit at *which* layer, says what the client experiences when it trips, and admits the trade: chunking buys isolation with round trips, and page-size caps buy safety by making some client's export slower and more paginated. It also separates the two amplification axes cleanly — call count and call size — and observes that fixing the first created the second.

  • Why might twenty smaller entity calls beat one call of 4,812 keys?
    Independent failure domains and retry units: one chunk failing costs one chunk, not the whole branch, and a retry re-does 256 keys rather than 4,812. Memory becomes predictable on both sides and backend lookups stay inside sane `IN`-list sizes. The cost is more round trips and more concurrency to manage, so chunk size is a tuning decision — too small and you have rebuilt the amplification you removed.
  • Should the limit live at the router or in the subgraph?
    Both, for different reasons. The router is the only component that sees the whole document, so cost scoring and page-size rejection belong there and give the client a clean, early error. But the entity field is a real entry point reachable by anything that can talk to the subgraph, so the subgraph must enforce its own maximum representation count — an upstream misconfiguration should not be able to exhaust a service's memory.
  • A federated subscription stopped delivering updates with no visible error. How could fan-out cause that?
    Each event payload is resolved through the same plan as a query, so a subscription does an entity fetch per event. If an event's list grows past a size or time limit, that event's payload fails to resolve while the connection stays open and healthy — the stream looks alive and the errors ride inside payloads a client may discard. The entity-fetch size histogram diagnoses it; the connection dashboard never will.
  • Which metric would have caught this before the incident?
    A histogram of representation-list size per subgraph per operation, alarmed at the 99th percentile. Call-rate and error-rate graphs stay flat because the amplification is inside one successful-looking request, and averages hide it — a p50 of 12 with a p99 of 4,812 is one client exercising something the schema permits.

Consolidating 4,812 parcels into one van solves the traffic problem right up to the moment the van will not start — and then nothing is delivered instead of most things.

saying these in an interview costs you the question

  • Says batching alone bounds cross-service fan-out
  • Raises the body-size limit and calls it fixed
  • Judges amplification by subgraph call count only
  • Silently truncates an oversized page instead of rejecting
  • Assumes a retry of a failed huge fetch is cheap
  • Puts the only size limit in the router

context