How do you set a GraphQL cost budget and allocate points across callers?
answer
- A point needs an exchange rate
- Observe mode before enforcing mode
- Two limits: per request and per window
- Tier the budget, publish the weights
- Owning team proposes, platform reviews
basics
~20 sAnchor a point in something measurable — one object fetched, or a millisecond of service time. Set the per-request cap above the p99 score of legitimate traffic, then budget each caller in points per window rather than requests per window.
solid answer
~50 sA point is a currency you invent, so the first job is an exchange rate: decide that one point means roughly one object fetched, or one millisecond of measured service time, and score real traffic in observe mode for a bake period before enforcing anything. The per-request cap then comes from the measured distribution — comfortably above the p99 of legitimate documents, not from a round number someone liked. One cap is never enough: a caller that cannot send one 300,000-point document will send forty 8,000-point ones, so the second limit is a **points-per-window** budget per caller, which makes cost the billed unit instead of requests. Tier those budgets — first-party app, contracted partners, anonymous traffic — publish the weights and return the score and remaining budget on rejection so integrators can self-correct, and review field weights alongside schema changes so the numbers do not drift.
go deeper
Know that the budget is a number someone chose, not a specified constant, and that a document is refused when its score exceeds it. Being able to read a rejection's score and budget is enough at this level.
Explain why a per-request cap alone is insufficient and what a per-caller points-per-window budget adds, and describe how you would gather the score distribution before switching a limiter into enforcing mode.
Show the rollout discipline: observe mode, a bake period covering a real business cycle, a cap set from the measured p99 of legitimate traffic, and rejections that carry enough detail for an integrator to fix their own document.
Own the tradeoffs end to end — the exchange rate between points and capacity, tiering by caller class, who proposes and who reviews a field's weight across many teams, and when a cost model is the wrong tool because every document is already known.
## A point has to mean something Complexity points are invented units, and an invented unit with no anchor produces theatre: a budget of "1000" that nobody can defend, tuned upward every time somebody complains. The first decision is the exchange rate, and there are only two honest choices. **One point ≈ one object fetched.** Easy to explain, easy for a client to predict from its own document, and blind to the fact that some objects are 400× more expensive than others — which you then correct with per-field weights. **One point ≈ one millisecond of measured service time.** Truer to capacity, much harder for a client to predict, and it needs continuous recalibration as the backend changes underneath it. Most teams pick the first and lean on weights. What matters is that you can finish the sentence "this endpoint can serve about N points per second" with a number derived from a load test, because that sentence is what turns a budget from a guess into capacity planning. ## Measure before you enforce Never launch a cost limiter enforcing. Run it in observe mode — compute the score, log it with the operation name and the caller, reject nothing — for long enough to cover a weekly cycle and a month-end. What you want out of that bake is the score distribution *per caller*, not a single average. The per-request cap belongs comfortably above the p99 of legitimate traffic, with headroom for the documents your own clients will write next quarter. Publish the resulting number and the weight table; a limit callers cannot predict is a limit they will trip. ## One limit is never enough A per-request cap stops the single catastrophic document. It does nothing about volume. A caller refused one 300,000-point document will send forty at 8,000, and your endpoint is just as loaded. So the second control is a **points-per-window budget per caller** — a refilling allowance, charged per operation, rather than a request counter. This is the change of unit that matters: on a REST API requests-per-minute is a rough proxy for load because routes have bounded shapes; on one GraphQL endpoint it is nearly meaningless, because two requests to the same URL can differ by five orders of magnitude. Billing in points also removes the perverse incentive to fragment: splitting one document into ten costs the same points, so clients stop gaming the shape. A refinement worth knowing: because the score is an upper bound, you can charge the estimate on admission and **refund the difference** once the actual work is known — the caller who asked for 288 readings and received 50 gets most of the points back. It makes budgets far more usable for honest clients at the price of a reconciliation path. It is a design, not a specified behaviour. ## Tiering, and treating the budget as a product surface Different callers deserve different allowances, and pretending otherwise either strangles your own app or hands an anonymous caller the same power as it: - **First-party clients** — documents you can inspect and measure ahead of shipping, so generous budgets with a pre-release check are reasonable. - **Contracted partners** — a budget that is part of the agreement, sized from their real workload and visible to them. - **Anonymous or trial traffic** — tight, and tight is defensible if the failure is informative. Whatever the tier, the rejection has to be actionable. Returning the score, the budget and the largest contributing path, and documenting the weights, is the difference between an integrator fixing their document in ten minutes and filing a ticket. That is a product decision as much as a security one. ## The organisational problem In a supergraph composed from 62 subgraphs, nobody has a global view of what a field costs. The team that resolves a field knows its cost; the platform team owns the budget and the blast radius. The workable split is that the owning team proposes a field's weight, the platform reviews it in the same change as the schema, and new fields land on a conservative default rather than a free one. Without that review, weights drift toward zero — every team's own field feels cheap — and the budget quietly stops binding. Two habits keep it honest: a periodic recalibration comparing configured weights against measured per-field cost on real traffic, and treating a weight change as a compatibility event, because raising a weight can invalidate documents that clients have already shipped. ## What you are trading away Static scoring over-charges by construction, so a budget tight enough to stop abuse will refuse some legitimate documents; that is the trade, and the mitigation is calibration and per-caller budgets, not a bigger global number. Over-restriction has its own cost: integrators fragment their traffic, poll more often and cache worse, so you push load around rather than removing it. And cost control is not authorization and not an execution deadline — it bounds what a caller may ask for, not what they may see or how long a single field may take. If your graph is first-party only, the honest alternative is to skip scoring altogether and admit only documents you have measured; a cost model earns its keep precisely when you cannot enumerate the documents in advance.
- Why budget points per window instead of requests per window?Because on a single GraphQL endpoint a request is not a unit of work — two calls to the same URL can differ by five orders of magnitude, so a request counter constrains almost nothing. Charging points makes cost the billed unit, which also removes the incentive to fragment: splitting one document into ten costs the same. A per-request cap stops the one catastrophic document; only a per-window point budget stops forty medium ones a second.
- How do you handle the gap between the charged estimate and the work actually done?Charge the estimate on admission, since that is all you know before execution, then reconcile: when the real work is known, refund the difference to the caller's window budget. The caller who asked for 288 readings and got 50 gets most of the points back, which makes budgets usable for honest clients without weakening admission control. It costs you a reconciliation path and some bookkeeping, and it is a design choice, not specified behaviour.
- Who should own a field's cost weight in a graph built by many teams?The team that resolves the field proposes it — they are the only ones who know what it costs — and the platform team reviews it in the same change as the schema, because they own the budget and the blast radius. New fields should land on a conservative default rather than a free one. Without review, weights drift toward zero, since every team's own field feels cheap to that team, and the budget silently stops binding.
It is a data plan, not a call limit: you sell gigabytes per month, because counting downloads tells you nothing about how much was moved.
saying these in an interview costs you the question
- Picks a round budget with no measurement behind it
- Caps per request only, leaving repetition unbounded
- Uses one budget for internal and anonymous callers alike
- Treats a point as a unit with no physical meaning
- Raises the budget whenever a caller complains
- Calls cost control an authorization mechanism