skip to content

What goes wrong when a federated entity's @key is not actually unique?

level: seniorimportance: should knowfreq 44%

answer

  1. Composition never looks at data
  2. Two objects, one identity
  3. Ask what the key is unique within
  4. Cache entries collapse together
  5. A repeatable directive makes migration possible

basics

~20 s

Nothing detects it. Composition takes uniqueness on faith, so colliding key values make two objects look like one: lookups return whichever row matched first, and any layer keyed on the identity serves one viewer the other's data.

solid answer

~50 s

Uniqueness is an **unchecked promise**. Composition validates types, fields and reachability; it never inspects data, so a key that identifies two objects composes cleanly and ships. At runtime the damage is quiet. When the router carries an identity to another subgraph, that service looks up by the key and returns whichever row it found — so a document joining across services can assemble one response out of two different accounts. Worse, anything that treats the identity as a cache key — a response cache in front of a subgraph, an entity-level cache in the router — collapses two objects into one entry, and a cache that served another viewer's row is a data-leak incident, not a correctness bug. The symptoms are intermittent, correlate with cache warmth rather than with any particular query, and look like a caching defect. The cause is a schema that declared a branch-unique account number as a graph-wide identity.

code

graphql · 14 lines
graphql
# Before: unique per branch, declared as a graph-wide identity
type Account @key(fields: "accountNumber") {
  accountNumber: String!
  displayName: String!
}

# Migration step: both keys valid at once, references move one by one
type Account
  @key(fields: "accountNumber")
  @key(fields: "sortCode accountNumber") {
  sortCode: String!
  accountNumber: String!
  displayName: String!
}

go deeper

for a junior

Remember that a key is a claim about data, not something the tooling verifies. If two objects can share the key value, the graph will treat them as one object.

for a middle

Explain the mechanics of the failure: a lookup by a colliding key returns an arbitrary match, no error is raised anywhere, and layers keyed on the identity merge the two objects into one entry.

for a senior

Be ready to diagnose it from production symptoms — intermittent, cache-correlated, no errors logged — and to reach the uniqueness-scope question rather than chasing the caching layer that merely exposed it.

for a principal

Own the policy: which keys the graph permits at all, whether a natural key is ever acceptable without a storage constraint behind it, and how a key migration is staged across teams that deploy independently.

## The promise nobody verifies When a subgraph writes `type Account @key(fields: "accountNumber")`, it is making a factual claim about data: *this value identifies exactly one account*. Composition checks a great deal — that the field exists, that it is resolvable here, that it takes no arguments, that the merged graph is satisfiable — but it cannot check that claim, because it never sees a single row. Neither can the router. Uniqueness is enforced, if at all, by the service that owns the data, and only if someone put a constraint there on purpose. So the failure mode of a weak key is not a build error. It is a graph that composes, deploys, and is wrong for a subset of requests. ## What a collision actually does Consider an 11-service bank statements graph in which six subgraphs reference `Account`. Account numbers are unique **within a sort code**, not globally: two branches can both hold `40286611`. The accounts subgraph keys on `accountNumber` alone. Now trace a request. A document asks for an account and its statements. The router fetches from the accounts subgraph, carries the identity forward, and asks the statements subgraph to resolve an account from `accountNumber: "40286611"`. That service looks up the value and finds two rows. Whatever its lookup does — take the first, take the most recent, take whichever the index returned — it returns one of them. The response is a single, well-formed `Account`: the display name from one branch's customer and the statement lines from another's. There is no error. The envelope has no `errors` entry, nothing is null, and every field is the right type. The graph confidently returns a coherent lie. ## Why caching turns it from a bug into an incident It gets worse where the identity is used as a cache key. Caching entity lookups by their identity is a widespread implementation technique — an entity-level cache in the router, a response cache in front of a subgraph — and it is not something the composition specification defines, so its behaviour is entirely a property of whatever you deployed. Every one of those layers takes the same thing on faith: distinct identities mean distinct objects. When two accounts share an identity, they share a cache entry. The first request to arrive populates it; the second gets the first one's data. Now the wrong-row problem is not merely intermittent, it is *sticky*, and it crosses a customer boundary: one viewer sees another viewer's statement row. That is a confidentiality incident with an audit trail, not a data-quality ticket. ## Why it is hard to diagnose Every signal points away from the schema. * The rate is tiny and non-obvious — the share of accounts whose number collides, so perhaps one request in a few thousand. * It correlates with cache warmth and traffic mix, not with a query shape, so replaying the same document usually returns the right data. * It disappears when caching is turned off, which reads as a caching defect and sends the investigation to the wrong team. * Nothing logs an error, because from every service's point of view every request succeeded. The question that cracks it is not "what is wrong with the cache", it is **"what is the uniqueness scope of every key in this graph?"** For each entity, ask what the declared key is unique *within* — globally, per tenant, per region, per branch — and compare that to what the graph assumed. A key that is unique per tenant in a multi-tenant graph is the same bug wearing different clothes, and it is common, because a natural key almost always has a scope the schema author forgot to mention. ## Fixing it without a coordinated deploy Widening or replacing a key looks like a breaking change across every referencing subgraph, and if done as one edit it is. Alternative keys make it survivable, because `@key` is repeatable: 1. Add the correct key alongside the wrong one on the owning subgraph: `@key(fields: "accountNumber") @key(fields: "sortCode accountNumber")`. The graph now accepts either identity, and every existing reference keeps working. 2. Move referencing subgraphs one at a time onto the new key, deploying independently. 3. When nothing references the old key any more, remove it and the incident class goes with it. Step 1 is the important one: it converts an all-at-once schema break into a migration each of the six referencing teams can schedule for itself. ## The prevention that actually holds The durable fix lives in the owning service, not in the schema: a real uniqueness constraint on the columns behind the key, so a colliding row fails to be written rather than failing to be distinguished. Two habits go with it — prefer a surrogate identifier over a natural one whenever the natural value's uniqueness scope is anything narrower than the whole graph, and treat every `@key` in review as a claim someone must be able to point at a constraint for.

  • How would you change an entity's key across six referencing subgraphs without a coordinated deploy?
    Add the new key alongside the old one — `@key` is repeatable, so the entity temporarily accepts both identities. Referencing subgraphs migrate to the new key independently, on their own release schedules. When nothing uses the old key, delete it. Replacing the key in a single edit would break every reference at once and force eleven services to ship together.
  • Where should uniqueness actually be enforced, if not in composition?
    In the service that owns the entity: a unique constraint in its storage on exactly the columns the key names, so a colliding record cannot be written in the first place. Schema review is the second line — every `@key` is a factual claim, and the reviewer's question is which constraint backs it. Nothing in federation will ever check it for you.
  • Why does this class of bug so often get filed against caching?
    Because caching is what makes it visible and sticky. Without a cache the wrong row appears only when a lookup happens to match the other record; with one, a single collision is stored and served repeatedly. Disabling the cache makes the symptom mostly vanish, which looks like proof the cache was at fault — while the weak identity that permitted it stays in the schema.

Declaring a key is like writing a name on a shared filing cabinet: the cabinet does not check that the name is unique, so the day two customers share one, every department starts filing into the same drawer and reading each other's papers.

saying these in an interview costs you the question

  • Assumes composition validates that a key is unique
  • Thinks the router compares full objects before joining
  • Blames the cache when the schema declared a weak key
  • Uses a natural key unique only per tenant or branch
  • Fixes it by disabling caching and moving on
  • Changes the key in every subgraph in one deploy

context