skip to content

You add a maximum-depth rule to a public GraphQL API and real traffic starts failing. How do you set the cap?

level: seniorimportance: should knowfreq 51%

answer

  1. Something legitimate is deeper than you think
  2. The tooling breaks before the product does
  3. Wrappers spend levels on structure
  4. Watch first, reject later
  5. Headroom above the legitimate maximum

basics

~20 s

Measure before enforcing. Log the depth of real documents through a full traffic and release cycle, then set the cap well above the legitimate maximum. Expect introspection and connection-wrapped traversals to be the deepest legitimate documents.

solid answer

~50 s

Two things break first, and both are predictable. The full introspection query nests `__schema { types { fields { type { ofType { ofType … } } } } }` past ten levels, so a cap of eight kills explorers, typed client generators and CI schema checks while product traffic looks fine — exempt the introspection meta-fields or set the cap above its measured depth. Second, a Relay-style connection spends three levels per hop on the connection field, `edges` and `node`, so a two-generation pedigree walk is already eight deep. The method: compute and log document depth without rejecting anything, run it through a peak of roughly 1,200 requests per minute and a full release cycle so pinned old clients appear, then set the cap with real headroom above the legitimate maximum and state both the limit and the observed depth in the error. And say it plainly: a depth cap is a floor, not a cost model.

code

graphql · 11 lines
graphql
query Descendants {
  animal(tag: "HFD-4471") {
    offspring(first: 25) {
      edges { node {
        offspring(first: 25) {
          edges { node { tag } }
        }
      } }
    }
  }
}

go deeper

for a junior

Know the headline surprise: the deepest documents hitting a public GraphQL API are usually legitimate — introspection and list-wrapping conventions — so a cap chosen by guesswork breaks tooling first. Measuring before enforcing is the safe instinct.

for a middle

Explain why the depth arithmetic is what it is: an unrolled introspection chain and three levels per connection hop for the connection field, edges and node. Be able to compute the depth of a document you are shown.

for a senior

Demonstrate the rollout discipline — instrument, observe a full traffic and release cycle, set headroom above the legitimate maximum, stage enforcement, alarm on rejections — and volunteer what the cap still fails to bound.

for a principal

Own the policy: which caller classes get which limits, what the documented process is when a legitimate client outgrows a cap, and how a blunt shape-based limit fits alongside weighted cost and page bounds in the overall abuse budget.

## The failure is predictable, so predict it before you ship Adding a maximum-depth validation rule to a public API is a change that looks free and is not. The cap rejects documents at validation, so anything it catches fails completely — no partial data, an `errors` entry and no `data` key — and the clients that break are usually not the ones you were worried about. Two categories break first, and both are avoidable if you know about them. ### Introspection is deeper than product traffic The standard full introspection query is one of the deepest documents any client will ever send. It walks `__schema { types { fields { type { ofType { ofType { … } } } } } }`, and the `ofType` chain is conventionally unrolled several levels to unwrap nested list and non-null wrappers such as `[Animal!]!`. Add the enclosing operation and you are comfortably past ten levels for a query whose execution cost is trivial. So a cap of eight, chosen because no product document goes past six, rejects introspection outright. The symptom is confusing: ordinary reads keep working, but in-browser explorers show an empty docs pane, typed client generators fail their schema fetch, and schema checks in CI go red — all simultaneously, all with a depth error nobody connects to the change. Handle it deliberately: either exempt documents whose root selections are the introspection meta-fields from the depth rule, or set the cap above the introspection query's measured depth. Do not discover this in production. ### Connection conventions spend depth on structure The Relay-style connection convention wraps every list in `edges` and `node`, so one hop costs three levels — the connection field, `edges`, `node` — before you reach a single real field: ```graphql query Descendants { animal(tag: "HFD-4471") { offspring(first: 25) { edges { node { offspring(first: 25) { edges { node { tag } } } } } } } } ``` Two generations of pedigree, one scalar field, already eight levels deep. A four-hop traversal is past fourteen. A cap tuned against a flat schema is simply wrong for a connection-shaped one, and the same is true of any schema whose domain is genuinely recursive — ancestry, org charts, bills of materials, threaded comments. Depth is not a proxy for cost in those schemas; a twelve-generation ancestry walk touches twelve records. ## Set the number from measurement, not from a default The method is the answer an interviewer wants. 1. **Instrument first, enforce later.** Compute the depth of every incoming document and record it as a metric and a log field, but reject nothing. Run that through a full traffic cycle — a weekday peak of around 1,200 requests per minute is a different population from a quiet Sunday — and a full release cycle, so pinned old mobile builds are represented. 2. **Read the distribution, not the mean.** Look at the maximum per client and at the high percentiles, and inspect the deepest documents by hand. Legitimate outliers and abusive outliers look different: one is a known operation from a known client, the other is a chain of the same field. 3. **Set the cap with real headroom above the legitimate maximum.** The purpose is to eliminate the absurd, not to trim the tail. If the deepest genuine document is 14, a cap of 15 is a false- rejection generator and a cap of 25 still refuses everything unbounded. 4. **Enforce in stages, and make the error legible.** Turn it on for unauthenticated traffic before first-party clients; state the limit and the observed value in the error message, because a bare "query too deep" turns every client-side debugging session into a guessing game. 5. **Alarm on rejections.** A sudden rise usually means a new client build, not an attack, and you want to see it before support does. Differentiate clients where you can. First-party callers ship a known, finite set of documents whose depth you can measure exactly at build time, so their cap can be tight. An open endpoint serving arbitrary partner traffic has no such luxury and needs a looser cap backed by other controls. ## Be honest about what the cap is worth A depth limit is a blunt instrument, and a strong candidate says so unprompted. It rejects legitimate deep documents and admits cheap-looking shallow ones; it measures document shape, not work. In this pedigree service the incident that actually hurt was not a deep document at all — it was an `offspring` list on a heavily used sire that grew without a bound over years of records, so a three-level document returned tens of thousands of nodes. No depth cap would have touched it; a page limit on the list argument would have. Treat the cap as a floor that removes the pathological cases cheaply, at validation, before any resolver runs — and pair it with a field-count limit, bounded list pages and a weighted cost model for the cases it cannot see. And say plainly that none of it is specified: the GraphQL specification defines document validity, not resource policy, so every one of these numbers is your house rule.

  • Why does a depth cap break in-browser explorers and codegen but leave normal reads working?
    Because the full introspection query is far deeper than typical product documents. It walks `__schema` into `types`, `fields`, `type` and a conventionally unrolled `ofType` chain that unwraps nested list and non-null wrappers, which pushes it past ten levels. Product documents rarely go past six. Either exempt the introspection meta-fields from the rule or set the cap above the introspection query's measured depth.
  • How would you roll this out without a window where you are either unprotected or breaking clients?
    Run it in report-only mode first: compute the depth of every document, emit it as a metric and a log field, reject nothing. Once you have a full traffic and release cycle of data, set the cap with headroom above the legitimate maximum and enable enforcement for unauthenticated traffic before first-party clients, with an alarm on rejection rate so a new client build shows up before support does.
  • Should first-party and third-party callers get the same depth cap?
    Not necessarily. A first-party client ships a known, finite set of documents whose depth you can measure exactly at build time, so its cap can sit just above that. An open endpoint serving arbitrary partner traffic cannot be bounded that way and needs a looser cap compensated by other controls. The cost is a second configuration and a second set of rejection metrics to watch.

saying these in an interview costs you the question

  • Picks a depth number from a blog post without measuring
  • Forgets that introspection is itself a deep document
  • Ignores the levels connection wrappers consume per hop
  • Enables enforcement without a report-only period first
  • Returns a bare too-deep error with no limit or observed value
  • Presents a depth cap as a complete cost control

context