When a federation router builds one _entities call, why does it deduplicate representations?
answer
- The same object appears many times
- Positions versus distinct keys
- Coalescing before the call goes out
- One result spliced into many slots
- An optimization, not a specified rule
basics
~20 sBecause a list usually names the same entity many times: 143 artworks by 37 artists yield 37 distinct keys. Sending each key once cuts payload and subgraph work. It is a router optimization, not a rule the federation specification imposes.
solid answer
~40 sBefore issuing an entity fetch the router has one representation per *response position*, and those repeat heavily — 143 artwork positions may reference only 37 distinct artists. Deduplication collapses structurally identical representations (same `__typename`, same key values) into one, keeps a map back to every position that referenced them, and splices the single returned object into all of them. The saving tracks the skew: smaller request body, and subgraph work proportional to distinct entities rather than positions. Crucially this is an implementation optimization. Apollo Federation's subgraph specification defines `_entities(representations:)` and the requirement that the returned list align with the input list; it never requires the router to deduplicate. So a subgraph must tolerate duplicate representations, must return one element per input position, and must never treat entity resolutions as a per-view usage count.
code
json · 7 lines{
"representations": [
{ "__typename": "Artist", "id": "PER-8811" },
{ "__typename": "Artist", "id": "PER-8811" },
{ "__typename": "Artist", "id": "PER-2094" }
]
}go deeper
Recall that many list positions often point at the same entity, and that asking about a key once is cheaper than asking 143 times. You are not expected to know how the collapsing is implemented.
Explain the mechanism: representations are grouped by type plus key values, the duplicate keys drop out, and the single result is fanned back into every position. Be ready to say that this is optimization, not specification.
Show what depends on not depending on it: subgraphs must tolerate duplicates, entity resolvers must be side-effect free, and any usage metric built on entity resolutions is measuring router behaviour rather than client behaviour.
Own the general rule this illustrates — a subgraph contract must be written against what the composition specification guarantees, never against what today's router happens to optimize, or an upgrade quietly changes the semantics of your services.
## The same entity, many positions In a federated response, one entity is frequently referenced from many places in the same list. In a museum collection graph, an exhibition of 143 artworks may be the work of only 37 artists — a retrospective is often a dozen artists and one prolific studio. When the client asks for each artwork's artist name, held in a **People** subgraph while **Catalog** owns the artwork, the router has 143 *positions* to fill but only 37 *distinct keys* to ask about. Deduplication is the step that collapses the second number into the first. Before issuing the entity fetch, the router groups the representations it built — each an object of `__typename` plus the key fields — and drops the repeats: ```json { "representations": [ { "__typename": "Artist", "id": "PER-8811" }, { "__typename": "Artist", "id": "PER-8811" }, { "__typename": "Artist", "id": "PER-2094" } ] } ``` becomes two representations, not three. The subgraph resolves 37 artists, the router keeps a map from each key back to every response position that referenced it, and the single returned object is spliced into all of them. ## Why it is worth doing The saving is proportional to the reference skew, and real graphs are skewed: authors, artists, tenants, categories, currencies, conservators. Three costs shrink together. The request body shrinks — 4,812 bytes of repeated ids is not nothing when representation lists run to thousands. The subgraph's work shrinks, because its work is proportional to distinct keys rather than to positions, whether it loads them one by one or in a batch. And the response shrinks, because the entity list comes back with 37 elements rather than 143 copies of the same object. Note what does *not* change: the client's response. Deduplication happens entirely below the response-shaping layer. The document asked for an artist under each of 143 artworks and gets 143 artist objects; whether they were fetched once or 143 times is invisible. ## Specified, or convention? This is the part that separates a candidate who has read the specification from one who has read a blog post. The Apollo Federation subgraph specification defines the `_entities` field, the `_Any` scalar for representations, and the requirement that the returned list align positionally with the input list. It does **not** require a router to deduplicate. Deduplication is a router-side optimization: some implementations do it, some do it only for simple keys, some make it configurable, and none of that changes whether a subgraph is correct. The practical consequences all follow from that: *A subgraph must tolerate duplicates.* If the same representation arrives twice, return the entity twice, once per position. Failing, collapsing the list, or returning a shorter list breaks the positional contract. *A subgraph must not count entity resolutions as usage.* If a conservation-audit table writes one row per resolved entity, the count answers "how many distinct keys were asked for", not "how many times a curator looked at this artist" — and it changes silently when the router's deduplication behaviour changes. Usage belongs in a metric derived from the client document, not from entity resolution. *A subgraph must still batch internally.* Thirty-seven keys in one call is thirty-seven lookups unless the subgraph loads them together. ## When identical entities fail to coalesce Two representations coalesce only when they are actually identical. Several situations break that: **Composite and nested keys.** A key of `@key(fields: "provenanceLot { id } accessionNumber")` produces a nested object; whether two such objects compare equal depends on the router normalizing them, not on the entity being the same. **Extra fields pulled in by a `@requires`.** When a field the subgraph resolves declares that it requires sibling fields, those values ride along inside the representation. Two references to the same artist can then carry different accompanying values and are, correctly, not the same representation. **Multiple keys on one type.** A type may declare more than one key; two references reaching it through different keys produce structurally different representations for the same underlying row. ## The interview register This is not a question anyone's offer turns on — it is a curiosity that reveals whether you picture the entity fetch concretely. The strong answer is two sentences of mechanism and one of caveat: the router collapses repeated keys because the subgraph's cost tracks distinct entities, and because it is an optimization rather than a specified rule, nothing in a subgraph may depend on it happening.
- What must a subgraph do if the router does not deduplicate?Behave exactly as it would otherwise: return one element per input position, in input order, repeats included. A duplicate key is not an error and must not collapse the list, because the router matches results to positions by index. Internally the repeats should cost one lookup, since a per-request batched load naturally coalesces them — that is where a subgraph recovers the saving the router chose not to make.
- When can two references to the same entity fail to coalesce?When the representations are not structurally identical. Composite or nested keys serialize to objects whose equality depends on normalization; a field declaring `@requires` pulls sibling values into the representation, so two references to one artist can carry different accompanying data; and a type with several keys can be reached through different ones. All three produce distinct representations for the same underlying row.
- Does deduplication change what the client sees in the response?No. It happens entirely below response shaping. The document asked for an artist under each of 143 artworks and gets 143 artist objects, identical in content to the unbatched result. Deduplication is invisible in the response and should be invisible in behaviour — which is why side effects in an entity resolver are a mistake.
A waiter taking 143 orders that are only 37 different dishes: the kitchen ticket lists 37 dishes, but 143 plates still leave the pass.
saying these in an interview costs you the question
- Says the federation specification requires deduplication
- Assumes each response position gets its own entity fetch
- Thinks deduplication changes the client's response shape
- Counts entity resolutions as per-view usage
- Assumes representations carrying required fields still coalesce
- Lets a subgraph reject duplicate representations as invalid