Where should authorization live in a federated graph — the router, each subgraph, or the services behind them?
answer
- Ask what each layer can know
- Coarse at the edge, precise at the data
- A denial before planning saves fan-out
- Autonomy buys drift
- Standardise the shape of a denial
basics
~20 sSplit it by what each layer can know. A router has the credential, the document and the schema but no data, so it can only enforce coarse rules. Per-record rules must live in the subgraph or below.
solid answer
~50 sDecide by information, not ideology. The router holds the caller's credential, the operation and the composed schema, so it can authenticate and refuse a field whose declared requirement the caller does not meet — cheaply, before planning, so a denied field never becomes a fetch. It cannot decide whether a record belongs to the caller, because it holds no data. Those rules belong in the subgraph, or better in a caller-scoped data layer that both root fields and entity resolution pass through. Centralising gives one auditable place and bounded blast radius, but queues nine teams' policy changes behind a 4-person platform team and tempts subgraphs to stop checking. Distributing gives precision and autonomy at the price of drift, and in a supergraph the weakest subgraph sets the real floor. Most graphs run both layers and standardise the shape of a denial across them.
go deeper
Recall the basic division: the edge can check who you are and coarse permissions, while whether a particular record is yours can only be decided where the data is loaded.
Explain why a router cannot evaluate per-record rules — it has the credential, the document and the schema but no loaded data — and why every service still verifies identity for itself.
Argue the layering concretely: authenticate and refuse coarse requirements before planning, re-verify and enforce record-level rules at the data, and keep the shape of a denial consistent across services.
Own the organisational tradeoff. Be ready to say what a small platform team standardises versus what product teams own, how you detect drift across many subgraphs, and why the weakest service sets the graph's effective floor.
## Start from what each layer can know Every sensible answer here falls out of one question: what information does this layer actually have at the moment it must decide? The **router** has the caller's credential, the operation document, and the composed schema. It knows the caller has scope `donations:read` and that the document selects `Donation.donorNote`. It has no data. It cannot know whether donation `dn_84213` belongs to this supporter, because nothing has been loaded yet — and if you make it load something to find out, you have quietly built a second data service inside the router. A **subgraph** has the caller's identity, the field being resolved, and the object it is resolving it on. It can evaluate per-record and relationship rules, because it is holding the record. It knows nothing about the other eight subgraphs' rules. The **service or store behind the subgraph** can enforce scoping in a way no resolver can escape, and is the only layer where "unfiltered read" can be made unexpressible. So the split is not ideological. Coarse, schema-shaped rules — is the caller authenticated, does this field require a scope they lack — are decidable at the router. Instance-level rules are not decidable anywhere above the data. ## What centralising buys, and what it costs Enforcing at the router gives one auditable place, one implementation, and denials that happen **before** planning: a rejected field never becomes a fetch, so a hostile or merely careless request costs no downstream work at all. Blast radius is bounded by a component the platform team already owns. The costs are real. The router cannot express the rules that actually matter most of the time. It becomes a queue: every policy change from nine product teams lands on a 4-person platform team's backlog. And a single central enforcement point invites the belief that subgraphs no longer need their own checks — which is false the moment anything can reach a subgraph directly, or the moment the router enters through `_entities` rather than a root field. ## What distributing buys, and what it costs Enforcing in each subgraph puts the rule next to the data and next to the team that understands it, which is the whole reason the graph was federated. Rules can be as precise as the domain needs, and they hold regardless of which entry point was used, if they were written that way. The cost is drift. Nine teams produce nine interpretations of "an administrator", nine deny shapes, and — statistically — at least one subgraph whose ownership check lives only in a root-field resolver and is therefore absent from the entity path. In a supergraph the weakest subgraph sets the real floor, because the router will happily reach the data through whichever service answers. ## The layered answer, and the part people forget Most graphs land on: authenticate at the router and reject unauthenticated traffic there; express coarse, declarative requirements in the schema so they are visible and enforceable before planning; re-verify identity in every subgraph rather than trusting a header; own instance-level rules in the subgraph or below, at a layer both root fields and entity resolution pass through. The part people forget is that the **shape of a denial** is a graph-wide contract. We changed one subgraph's nullability on a Tuesday deploy — `Donation.amountMinor` went from nullable to non-null, which read as a tightening — and a redaction strategy that had been returning `null` for amounts a caller could not see began erasing whole campaign branches instead of individual fields. Roughly 3,100 supporters saw an empty page rather than a partially redacted one. Nothing about authorization changed; a nullability decision in one subgraph collided with a denial convention in another. Standardise the convention — null for redaction on deliberately nullable fields, or an error with a path — and make it reviewable. ## What a small platform team should own Own the mechanism, not the policy. That means: the shared verification middleware every subgraph mounts; the router's header allowlist and strip rules; the rule that entity resolution runs the same checks as root fields, with a conformance test each subgraph runs in its own pipeline; the deny shape; and a review checklist that asks which entry points reach a field. Do not own the content of nine domains' rules — that is the bottleneck the federation was supposed to remove. ## Knowing when it has drifted Drift is not visible from the supergraph, so probe below it: periodically send `_entities` documents directly at each subgraph as an unrelated caller and assert refusal, and treat any subgraph that answers as an incident rather than a ticket. Composition validates types, not rules; there is no composition check that will tell you a subgraph forgot to authorize.
- Are there federation directives for declaring authorization requirements in a subgraph schema?Later editions of the composition specification define declarative directives — `@authenticated` and `@requiresScopes` among them — that let a subgraph state a coarse requirement on a field or type in its own SDL, which a router can enforce before planning. Support depends on the federation version and on the router, they express scope-shaped requirements only rather than per-record rules, and they are a composition convention rather than anything the GraphQL specification defines.
- What should a small platform team own here, and what should it deliberately not own?Own the mechanism: shared credential verification, the router's header allowlist and strip rules, the requirement that entity resolution runs the same checks as root fields, a conformance test each subgraph runs in its own pipeline, and the graph-wide deny shape. Do not own nine domains' policy content — that recreates the central bottleneck federation existed to remove, and the team will be the least qualified reviewer of each rule.
- How would you detect that one subgraph's authorization has drifted from the rest?Probe below the supergraph, because from the client side there is only one graph. Periodically post `_entities` documents directly at each subgraph as a caller with no relationship to the data and assert refusal, and treat an answer as an incident. Composition validates types, not rules, so nothing in the build will tell you a subgraph forgot to check; only an active probe or a code review will.
saying these in an interview costs you the question
- Wants all authorization in the router because it is one place
- Leaves every rule to subgraph teams and calls it autonomy
- Assumes a router can check whether a record belongs to the caller
- Treats declarative auth directives as part of the GraphQL specification
- Lets each subgraph pick null-versus-error redaction independently