Your GraphQL cost limiter began rejecting a shipped mobile app's query after a schema change — how do you diagnose and fix it?
answer
- Rejections started at a release, not a peak
- Log the score, budget and top path
- Replay the document against both configs
- Wrapper fields sit inside the multiplier
- Fix the weight, not the budget
basics
~20 sLog every operation's score, its budget and the path that contributed most, then replay the rejected document against the weight tables before and after the deploy. A rename that adds charged levels inside a list multiplier is the usual cause.
solid answer
~50 sStart from evidence, not theory: the limiter must emit the computed score, the budget and the top-contributing path per operation, otherwise a rejection is indistinguishable from any other 4xx. Correlate the rejections with client version and operation name — a cliff that starts at a deploy and only affects one app build tells you the document changed, the weights changed, or both. Then replay the exact document against the old and new configuration and diff the two scores. The classic cause is structural: renaming a plain list field to a connection puts `edges` and `node` **inside** the slice multiplier, so a 288-item slice suddenly pays for two more charged levels and the score roughly doubles. Fix it by charging structural wrapper fields nothing, not by raising the global budget — that weakens the control for every caller. Then gate future documents in CI against the deployed weights.
code
json · 14 lines{
"errors": [
{
"message": "Operation exceeds the cost budget",
"extensions": {
"code": "COST_LIMIT_EXCEEDED",
"score": 110811,
"budget": 75000,
"topContributor": "site.arrays.inverters.samples.edges.node"
}
}
],
"data": null
}go deeper
Know that a cost rejection is a server-side policy decision, not a bug in the client, and that the first thing to look for is what changed — the document, the weights or the budget — rather than retrying.
Be able to recompute a document's score by hand and show where the points moved, including why fields nested inside a list slice are charged once per item and fields beside it are charged once.
Demonstrate the operational half: the instrumentation that makes a rejection diagnosable, correlating rejections with release and operation name, and choosing the fix that does not weaken the control for every other caller.
Own the policy question. A weight table is a published interface with compatibility consequences, so it needs versioning, review with the schema, an observe-mode rollout and a pre-release check against the client's document corpus.
## The incident A solar-array telemetry graph enforces a cost budget of 75,000 points per request. Mobile release 4.11 shipped on a Tuesday. Within forty minutes of the phased rollout reaching about a fifth of devices, one operation's rejection counter went from zero to roughly 6,900 per minute, and the support queue filled with "the history screen is empty". The schema change that preceded it looked harmless. `Inverter.readings: [Reading!]!` was renamed and reshaped into `Inverter.samples`, a connection following the Relay server specification's convention, and the mobile client was updated in the same release to select the new field. Before: ```graphql readings(first: 288) { wattage capturedAt } ``` After: ```graphql samples(first: 288) { edges { node { wattage capturedAt } } } ``` ## Why the score doubled With every field weighted 1, and the enclosing document selecting 12 arrays per site and 8 inverters per array: | level | before | after | |---|---|---| | per-item selection | 2 (`wattage`, `capturedAt`) | 4 (`edges`, `node`, and the two scalars) | | the list field | 1 + 288 × 2 = 577 | 1 + 288 × 4 = 1,153 | | one inverter | 578 | 1,154 | | `inverters(first: 8)` | 4,625 | 9,233 | | one array | 4,626 | 9,234 | | `arrays(first: 12)` | 55,513 | 110,809 | | **document total** | **55,515** | **110,811** | The client asks for exactly the same 27,648 readings as before. Two purely structural fields — `edges` and `node` — landed *inside* a 288× multiplier, at 96 inverters, and added 55,296 points to a 75,000-point budget. Nothing about the work changed; the pricing of the shape did. ## Diagnosing it The first question in a review of this incident is not "what changed" but "how long did it take you to know". A cost limiter that returns a bare error and logs nothing is unowned. The instrumentation that makes this a ten-minute diagnosis: - **Emit the score with every operation**, accepted or rejected, together with the budget in force. A score is only useful as a distribution. - **Emit the top contributing path** — here `site.arrays.inverters.samples.edges.node` — because the interesting fact is *where* the points are, not that there were too many. - **Tag by operation name and client build**, so a rejection cliff can be read against a release. - **Keep a replay tool**: feed a stored document and its variables through the current weight table offline and print the same breakdown. Diffing that output across two config revisions is what turns a guess into a cause. The signature here is unambiguous once those exist: rejections are confined to one operation and one client build, they start at a release rather than at a traffic peak, and the score for that document changed while the schema's own semantics did not. ## Fixing it Short-term, the release cannot be recalled, so something on the server must move. The options, worst to best: 1. **Raise the global budget.** Fast, and it weakens the control for every caller including the abusive one. Avoid it as anything but a fire-break with an expiry. 2. **Raise the budget for that caller.** Acceptable as a stopgap when the traffic is first-party and the real cost is known to be unchanged. 3. **Fix the weights.** `edges`, `node` and `pageInfo` are structural wrappers, not work. Charging them zero — or folding a connection's whole wrapper into the connection field's own weight — restores the score to what the work actually is. This is the correct fix, and it generalises to every connection in the schema, not just this one. Longer-term the failure is a process failure. A weight table is an interface: change it and previously valid documents become invalid. So version weights with the schema and review them in the same change, roll weight changes out in **observe mode** first — score, log, do not reject — for a bake period, and score the client's committed document corpus in CI against the deployed weights, failing the build when any document exceeds some fraction of the budget. A document at 74% of budget in CI is a release you can still stop. ## The general shape of a false rejection The score is an upper bound on *authorised* work, so it over-charges whenever the real work is smaller than the document permits: a slice of 288 the backend silently caps at 50, a field weighted for its worst case that is usually served from a warm cache, an abstract-type selection scored as the sum of branches when only one can apply. The remedy is calibration — compare configured weights against measured cost per field on real traffic and adjust the weights — not budget inflation. Raising the budget makes every document cheaper, including the one you built the control to stop.
- Why is raising the budget the wrong first response to a false rejection?Because the budget is a single global dial: raising it until one legitimate document passes makes every other document cheaper too, including the ones the control exists to refuse. It also hides the real defect, which is that a weight no longer reflects the work. Prefer fixing the weight, or granting the affected caller its own budget, and treat a global raise as a fire-break with an expiry date and a ticket attached.
- How would you have caught this before the mobile release shipped?By scoring the client's committed document corpus in CI against the weight table that is actually deployed, and failing the build when a document crosses some fraction of the budget — 70% is a reasonable line. That converts an outage into a red build. Pair it with rolling weight changes out in observe mode, where the limiter computes and logs the score but does not reject, so you can see the new distribution against real traffic before it can refuse anything.
- How do you tell a false rejection from an abusive document in the logs?By the shape of the traffic, not the score. A legitimate document is one of a handful of named operations, comes from a known client build, and appears at a stable rate; its score is stable too, and jumps at a deploy. Abuse is usually novel document text, anonymous or one-off callers, scores far above the budget rather than just over it, and a rate that climbs. Logging the operation name and client build is what makes the two separable at all.
The freight did not get heavier; the carrier started charging per pallet as well as per crate, and the same shipment stopped fitting the budget.
saying these in an interview costs you the question
- Raises the global budget until the document passes
- Rejects with no score, budget or path logged
- Treats a weight change as a non-breaking change
- Assumes the schema change altered the real work
- Ships weight changes straight into enforcing mode
- Charges structural wrapper fields as if they were work