How do you choose which GraphQL abuse controls to run at each phase for a public endpoint?
answer
- Legitimate traffic pays for every control
- Information decides the earliest phase
- Nothing outside the service parses documents
- Registered documents collapse the middle tiers
- Static scores bound shape, not work
basics
~20 sRank candidate controls by what each phase knows, what each costs on every legitimate request, and how certain each rejection is. Push size, rate and content-type outward, keep document-shaped limits in the service, and treat execution deadlines as the backstop nothing static can replace.
solid answer
~50 sStart from the traffic, not the checklist: at a 1,200-request-per-minute peak, every control is paid 1,200 times a minute by legitimate callers and only occasionally by an attacker, so the ordering question is really a cost-and-certainty question. Push outward what needs only bytes and headers - body size, connection rate, required `application/json` - accepting that no component in front of the parser can ever see a selection set. Keep document-shaped limits in the service, because they need the schema, and prefer the ones with crisp answers (depth, field count, operation count) over the ones that need calibration (static cost). Decide deliberately between an open endpoint with tuned limits and a registered-document endpoint that removes untrusted parsing entirely, because that choice is organisational, not technical: it buys enormous simplification and costs you release coupling and third-party clients. Then accept that every static score is an upper bound on shape, not a prediction of work, and keep a per-operation deadline as the only control that sees reality.
code
json · 7 lines{
"peakRequestsPerMinute": 1200,
"edge": { "maxBodyBytes": 65536, "requireContentType": "application/json" },
"registeredDocuments": { "firstParty": "required", "partner": "optional" },
"validation": { "maxDepth": 9, "maxFields": 400, "maxOperations": 1, "costBudget": 5000, "mode": "report-only" },
"execution": { "deadlineMillis": 900, "maxConcurrentBackendCalls": 24 }
}go deeper
Focus on knowing that these limits exist and are chosen by the team, and that some belong at the HTTP layer while others need the parsed document.
Be able to name a sensible starting set and justify each placement by the information available at that phase rather than by habit.
Demonstrate calibration: measure the real distribution, roll limits out in report-only mode, and alert on rejection rate as a client-regression signal.
Own the strategic fork between an open, cost-bounded endpoint and a registered-document one, the release coupling it implies, and the residual risk that only execution-phase controls can address.
### Frame it as a budget, not a checklist Controls are usually presented as a list to adopt. That framing produces stacks that are simultaneously over-defended in cheap places and undefended where it matters. The useful frame is that every control has three properties: the *information* it needs (which fixes its earliest possible phase), the *cost it imposes on legitimate traffic* (paid on every request), and the *certainty of its verdict* (how often it rejects something that was fine). At a 1,200-request-per-minute peak, a control adding two milliseconds costs about 2.4 seconds of aggregate latency a minute, forever, to stop something that may arrive twice a week. That is often still worth it - but it is a decision, and it should be made explicitly. ### Tier by tier, with what each tier cannot do **Outside the service** - a reverse proxy, a load balancer or a CDN - you can enforce anything computable from bytes and headers: request size, connection and request rate, TLS and origin policy, and a required `application/json` content type. What you cannot do there is anything about the document, because that component does not parse GraphQL and does not hold the schema. Teams periodically try to move the depth cap outward; the honest answer is that doing so means running a second GraphQL implementation with a second copy of the schema to keep in sync, and the sync is where it will fail. One trap belongs to this tier specifically. Response caching keyed on the request body is attractive here and quietly wrong for a graph where different fields have wildly different freshness needs. A cached seat-map response makes `seatsRemaining` a field that returns stale data - a caller sees 47 seats on a tier that sold out ninety seconds ago and starts a checkout that cannot complete. Freshness is a per-field property, and a cache placed before the parse cannot see fields at all. **At the service edge, before parsing**, you authenticate the caller and, if you have one, resolve a registered document identifier. This is the highest-leverage placement available, because it decides whether attacker-controlled text ever reaches your parser. **At validation**, you get the document and the schema and none of the runtime world. Prefer controls with crisp verdicts here - a depth cap, a total field-count cap, a cap on operations per document, a refusal of introspection - and treat a static cost score as the expensive, powerful, high-maintenance option it is: it needs per-field weights, it needs re-tuning whenever the schema changes, and it produces the false rejections that generate the support tickets. **At execution**, place the controls that need reality: a per-operation deadline, per-object authorization, caps on concurrent backend work, and batching that bounds the calls one field can fan out into. ### The one strategic fork Everything above assumes an open endpoint. The alternative reshapes the whole stack: register every document the first-party clients use, and reject anything not on the list. Depth, breadth, alias and cost limits become nearly redundant, because the set of possible documents is finite and reviewable before release. What you pay is coupling - a client cannot ship a new query without a registration step - and the loss of arbitrary third-party callers, which for many products is the entire point of exposing a graph. The honest answer is usually both: register the documents your own applications use and admit them by identifier, while keeping a limited, cost-bounded open path for partners, with different budgets on each. Say which you are choosing and why, because that is the judgement the question is asking about. ### Calibrating rather than guessing Caps chosen from a blog post reject real traffic. Measure first: log observed depth, field count and static cost for a week of real operations, set the initial caps above the observed 99.9th percentile, and run every new limit in report-only mode before it rejects anything. Alert on rejection rate, not just on attacks, because a spike after a client release is far more likely to be your own team than an adversary. ### What to concede A good principal answer names the residual risk. A static score bounds the *shape* of a document, not the work it causes; a shallow, cheap-looking operation over a badly indexed table will pass every pre-execution control you have. Deadlines and backend concurrency caps are what catch it, which is why they are not optional even on an endpoint with an excellent cost limiter. And no static control ever compensates for a missing authorization check - a perfectly bounded, promptly executed, correctly sized request that reads someone else's data is still a breach.
- You can only fund two controls this quarter. Which two, and why those?A per-operation deadline and a depth-plus-field-count cap. The deadline is the only control that bounds real work rather than document shape, so it catches the expensive operations every static rule misses, and it needs no calibration. The static cap is the cheapest possible rejection of the pathological documents that cyclic type references make legal, needs no per-field weighting, and produces almost no false rejections. Cost scoring is more powerful and much more expensive to own; it can wait.
- How do you pick the initial numbers rather than copying them?Measure. Instrument depth, field count and static cost for every real operation for a week, look at the distribution rather than the mean, and set the first caps above the observed 99.9th percentile. Ship every limit in report-only mode, watch what it would have rejected, and only then enforce. Keep alerting on the rejection rate afterwards, because a spike almost always means a client release rather than an attack.
- What residual risk survives a well-placed control set, and what do you do about it?Operations that look cheap and are not. Every pre-execution control scores document shape, so a shallow selection over an unindexed table or a fan-out hidden behind one field passes them all. The answers are execution-phase: a deadline, caps on concurrent backend calls, per-request batching to bound fan-out, and field-level latency metrics so the expensive fields are visible. That is also why a cost limiter never removes the need for a timeout.
- When is registering every document the wrong answer?When arbitrary third-party callers are the product. Registration assumes you control every client and can gate its release, which is true for your own applications and false for partners writing their own operations. It also bites on pinned mobile versions you cannot force-upgrade, since their registered documents must keep working for as long as those builds exist. The usual resolution is two paths with different budgets rather than one policy for everyone.
saying these in an interview costs you the question
- Adopting every control from a list without ranking them
- Assuming an edge component can enforce document depth
- Copying depth and cost numbers instead of measuring traffic
- Believing a cost limiter removes the need for a timeout
- Caching whole responses where field freshness differs
- Treating document registration as free of release coupling