How do you set latency and availability SLOs for one shared GraphQL endpoint?
answer
- The unit is not the endpoint
- A weighted average nobody owns
- You cannot promise arbitrary work
- Partial responses need a written rule
- One owner per objective-bearing operation
basics
~10 sNot at the endpoint. Choose a small set of user-facing operations from a known document set, give each its own objective, and write down what counts as success when a response is partial.
solid answer
~50 sAn endpoint-level objective on a GraphQL API is a weighted average dominated by whichever operation is cheapest and most frequent, so it is green during outages of everything that matters. Set objectives **per operation**, and accept the three preconditions that implies. First, the document set must be knowable — you cannot promise a latency for a document a caller composes on the fly, because the caller chose the work. Second, the identity you objectivise on must be server-derived, not the client-declared name, or the objective can be misattributed. Third, success must be a written predicate per operation: a request error is a failure, but a field error that nulls an optional side panel may not be, and leaving that implicit means two teams compute different numbers from the same responses. Then pick a dozen operations a human waits on, not all 137, and give the rest coarse health monitoring instead.
code
pseudocode · 19 linesobjectives = {
"PayslipDetail": {
latency_ms_p99: 800,
success_ratio: 0.999,
# a nulled optional panel is still a usable screen
tolerated_error_paths: [ ["payslip", "benefitSummary"] ]
},
"SubmitBenefitEnrollment": {
latency_ms_p99: 2500,
success_ratio: 0.9995,
tolerated_error_paths: [] # a mutation is all-or-nothing here
}
}
is_good(response, spec):
if not response.has_key("data"): return false # request error
for e in response.errors:
if e.path not in spec.tolerated_error_paths: return false
return truego deeper
Know that a single GraphQL endpoint mixes cheap and expensive operations, so an overall success or latency target says little about any real user journey. That is the foundation the rest of this judgement rests on.
Be able to explain why the objective's unit is the operation, and what has to be true for that to work: server-derived operation identity, and a decision about whether a response carrying both data and errors counts as success.
Show you would implement it. Talk about picking the handful of operations a human waits on, encoding the success predicate as data beside each one, and why a known or persisted document set is a precondition for any latency promise.
Own the whole scheme: which promises the organisation makes, who owns each, how budgets are attributed when one operation spans several teams, and how you stop an endpoint-level liveness signal from being mistaken for a user-facing objective.
## Why the endpoint is the wrong unit An objective is a promise to a user about a thing they do. On a path-per-resource API the route is a passable proxy for that thing. On a GraphQL endpoint it is not: one route carries a badge lookup that renders on every screen and a payroll reconciliation an administrator runs at month-end. An endpoint-wide success ratio is the traffic-weighted mean of unrelated promises, and because the cheap operation dominates the weights, the aggregate can sit above target through a total failure of the expensive one. Nobody owns it, nobody can act on it, and it will be green on the worst day of the quarter. So the real question is not *what number* but *what unit*, and the unit is the operation. Getting there forces three decisions that are the substance of this answer. ## Decision one: which operations are objective-bearing, and can they be? Start with a filter that is structural rather than editorial. **You can only objectivise a document set you know.** GraphQL hands the caller control of the selection, so the amount of work behind one request is a function of what the caller wrote. There is no honest latency promise for an arbitrary document — the same root field, selected two levels deeper, is a different workload. That makes a known document set (registered documents, or a persisted-document identifier the server already keys on) a precondition for a latency objective, not a nice-to-have. Where callers compose freely, the most you can promise is availability of the *endpoint*, plus limits that bound the work. Then apply judgement. A payroll and benefits graph with 137 named operations does not need 137 objectives; that is an unmaintainable surface and, incidentally, its own cardinality problem. Choose the ones where a human is waiting and a failure is visible to them: the payslip view, the benefits enrolment mutation, the reconciliation report. A dozen is a realistic number. Everything else gets monitoring and alerting without a formal objective. Separate mutations from queries while you are at it. Availability of a benefits enrolment mutation and latency of a dashboard query are different promises with different consequences, and mutations often deserve a tighter availability target and a looser latency one. ## Decision two: what counts as a failure This is where most GraphQL objectives quietly fall apart, because the response envelope has a middle state that a status-based SLI does not. Three outcomes exist: execution never began (no `data` entry), execution produced a complete result, and execution produced `data` with a non-empty `errors` list. The third is a judgement call and it must be *written down*, per operation: * A request error is unambiguously a failure — the caller got nothing. * A field error on a position the screen cannot function without is a failure. * A field error that nulls an optional panel — a benefits-provider summary alongside a payslip that rendered fine — may legitimately be a success. Encode that predicate as data next to the operation, not as tribal knowledge in a dashboard query, because the moment two teams evaluate partial responses differently the numbers stop being comparable and the budget stops being a shared currency. A pleasant side effect: writing it forces schema authors to say which fields are load-bearing, which is a design conversation worth having anyway. A related trap: a schema that models expected failures as typed result payloads produces no `errors` entry, so those requests are structurally successful. Decide deliberately whether your objective measures execution health (they are successes) or user outcomes (they are not), and be consistent. ## Decision three: ownership when one operation spans several teams A single document routinely selects fields owned by three teams. When one team's resolver fails, the request fails, and the budget it burns belongs to an operation, not to a resolver. The workable arrangement is that the **operation** has one owner accountable for its objective, while per-error attribution — which field path failed, which service behind it — routes the page and the follow-up to the team that broke it. Splitting a budget evenly across contributing teams sounds fair and produces the opposite of accountability: everyone is partly responsible and nobody acts. This is also the argument for keeping the objective set small. Each objective needs an owner who can actually change the thing being measured; if you cannot name that person, the objective is decoration. ## What to keep at the endpoint level Keep a coarse endpoint signal, but demote it. "Is the process accepting and answering requests at all" is a real question, it is cheap, and it catches whole-service failures fast. Just do not let it be the number anybody promises, and do not let a green endpoint line be taken as evidence that anything a user does is working.
- A caller composes an ad-hoc document nobody has seen before. Can it sit inside a latency objective?Not honestly. The caller chose the selection, so it chose the work, and the same root field selected two levels deeper is a different workload with a different floor. For freely composed documents the promise you can make is endpoint availability plus bounds — depth and cost limits, timeouts, pagination caps — rather than a latency target. A registered or persisted document set is what converts "some GraphQL request" into a thing with a stable cost you can promise against.
- One operation selects fields owned by three teams. Whose error budget burns when one of them fails?The operation owner's. A budget is attached to a user-visible promise, and the promise is the operation; splitting it across contributing teams diffuses accountability until nobody acts. Per-error attribution still matters, but it does a different job: the failing field path and the service behind it route the page and the fix to the team that broke it, while the operation's owner stays accountable for the objective and for deciding whether that dependency should be tolerated or made optional.
- Do you keep any endpoint-wide indicator at all?Yes, demoted. "Is the endpoint accepting requests and answering them" is cheap, catches whole-service failures quickly, and is a reasonable thing to alert on. What it must not be is the number anyone promises, or evidence in an incident review that users were fine. State plainly that it is a liveness signal, keep it visually separate from the per-operation objectives, and make sure nobody's dashboard puts a green endpoint line where a user-facing objective should be.
saying these in an interview costs you the question
- Sets one endpoint-wide objective and calls it done
- Counts every partial response as a total failure
- Promises latency for arbitrary client-composed documents
- Gives all 137 operations equal formal objectives
- Leaves the partial-success predicate implicit
- Splits one operation's budget evenly across teams