skip to content

Why can't per-object authorization be enforced as a GraphQL validation rule?

level: seniorimportance: should knowfreq 42%

answer

  1. Two questions hide under one word
  2. One needs the document, one needs the object
  3. Nothing is loaded before execution
  4. Fetching inside a rule breaks four things
  5. Request error versus field error

basics

~20 s

Validation runs over the document and schema before any resolver executes, so the object does not exist yet. Whether this caller may read this particular order can only be decided once the order has been loaded - during execution.

solid answer

~50 s

Split authorization into two questions with different answers. "May this caller ever select this field?" depends on the document plus the caller's identity, both known before execution, so it can be enforced as a rule over the parsed document and rejects cheaply. "May this caller see *this* object?" depends on the object - its owner, its venue, its state - and nothing before execution has loaded it. A validation rule that tried would have to fetch data itself, which makes validation impure, runs outside the operation's deadline, and duplicates the fetch the resolver is about to do anyway. The placement also changes what the caller sees: a validation rejection is a request error, so the whole operation is refused with no `data` at all, while an execution-time denial is a field error that leaves the rest of the response intact. Coarse static checks early, instance checks on the resolved object.

code

graphql · 7 lines
graphql
query OrderDetail($id: ID!) {
  order(id: $id) {
    reference
    buyerEmail
    internalRiskScore
  }
}

go deeper

for a junior

Remember the order of events: nothing has been loaded when a document is validated, so any check about a specific record has to wait for execution.

for a middle

Be able to split the static field-permission question from the instance question and say precisely what information each one needs.

for a senior

Demonstrate the production consequences: an I/O-performing validation rule sits outside the deadline and duplicates loads, and validation denial versus field denial changes what a client can still render.

for a principal

Own the layering decision - which checks are cheap early exits, which seam carries the authoritative check, and how you keep new traversals from bypassing it as the schema grows.

### Two different questions wearing the same word "Authorization" on a GraphQL endpoint covers two questions that need different machinery, and conflating them is the mistake this topic exists to catch. The **static** question is about the document: may a caller with this identity mention this field at all? On a ticketing and venue schema, `Order.internalRiskScore` might be selectable only by fraud-team callers; nobody else should be able to name it, ever, for any order. Everything that question needs is present before execution - the parsed document lists the field, the schema says which type it belongs to, and the caller was authenticated at the HTTP layer. So it can be enforced as a rule over the parsed document, and it is worth doing there: the rejection costs a parse and a tree walk instead of a round trip to the order store. The **instance** question is about an object: may this caller read *order 8842*? That depends on facts stored in the order - who bought it, which venue it belongs to, whether it was transferred. None of those facts exist before a resolver loads them. A validation rule inspecting the document sees `order(id: $id) { buyerEmail }` and, because variable values are coerced as part of executing the request rather than validating it, may not even see which order is being asked for. ### Why "just fetch it during validation" is the wrong fix The tempting shortcut is to have the rule load the object and decide. It goes wrong in four ways at once. It makes validation impure and slow. Validation is meant to be a fast, deterministic function of document and schema - a few hundred microseconds of tree walking. A rule that performs I/O turns your cheapest rejection stage into a network-bound one, and does it for every request including the ones that were going to be rejected by the depth cap two rules later. It runs outside the operation's bounds. The per-operation deadline covers execution. Work done in a validation rule is happening before that clock starts, so a slow authorization fetch is unbounded by the very control you added to bound slow work. It duplicates the fetch. The resolver is about to load the same object. Now the object is loaded twice, once outside any per-request batching or caching the execution phase provides, and the two loads can disagree - the second one seeing a newer state than the one you authorized against. And it does not generalise. The document names entry points, not the objects the traversal will reach. A rule reading the document cannot know which orders the seat-hold traversal will end up touching, so it can only authorize the fields it can see named - which is precisely the static question again. ### The placement changes what the client sees This is the detail that separates a candidate who has run this in production. A rejection during validation is a *request error*: the operation is refused as a whole, the response carries an `errors` list and no `data` key, and the client receives nothing at all. A denial during execution is a *field error*: the field's value is unavailable, the surrounding response still carries every other field the caller was entitled to. For the static check, whole-operation refusal is right - the document was never acceptable. For the instance check, refusing the whole operation would be a poor experience and often wrong: a caller listing thirty orders on their account page, one of which was transferred away, should get the other twenty-nine. ### Where the check actually goes at execution Once you accept that instance authorization is an execution-phase concern, the design question becomes *which* execution-phase seam it attaches to: the resolver that produced the object, a wrapper applied uniformly to the fields that need it, or the data layer that loads the object in the first place. The lower it sits, the harder it is to bypass by adding a new traversal - but that design choice is downstream of the placement fact, and the placement fact is what an interviewer is testing here. ### The shape of a strong answer Name the split - static field permission versus instance permission. Locate each: document-plus-identity is available pre-execution, object state is not. Explain concretely why a fetching validation rule is worse than useless: impure, untimed, duplicated, and still unable to see the objects a traversal will reach. Then close on the observable difference, request error versus field error, because that is the part a team feels the day they get it wrong.

  • What does the caller see differently when the denial happens at validation rather than during execution?
    A validation denial is a request error: the operation is rejected as a whole, and the response contains an `errors` list with no `data` key, so the caller gets none of the fields they were entitled to. An execution denial is a field error: the offending field is unavailable and reported in `errors` with a path, while the rest of the response is delivered normally. For a list where one element is forbidden, that difference is everything.
  • Is there anything you gain by ALSO doing the static check, given the execution check is authoritative?
    Yes - cost and clarity. A document that names a field this caller may never select is rejected after a parse and a tree walk, before any resolver runs or any backend is touched, which matters most under load. It also gives an unambiguous, whole-operation answer for a document that was never acceptable, rather than a response full of null holes. The execution check remains authoritative; the static one is a cheap early exit.
  • Why is a validation rule that loads data to make its decision worse than doing nothing there?
    It converts the cheapest phase into a network-bound one for every request, including those about to be rejected by a later static rule. It runs before the operation's deadline starts, so the fetch is unbounded by the timeout you rely on. It duplicates the load the resolver is about to perform, outside per-request batching, and the two reads can see different states. And it still cannot cover objects reached later in the traversal.

saying these in an interview costs you the question

  • Claiming a validation rule can check object ownership
  • Fetching data inside a validation rule to authorize it
  • Believing validation sees variable values such as an id
  • Thinking a denied field always aborts the whole operation
  • Treating authentication at the edge as authorization
  • Assuming one check on the entry field covers the traversal

context