skip to content

Why must a subgraph's _entities resolver return its results in the input order?

level: seniorimportance: should knowfreq 47%

answer

  1. Position is the only correspondence
  2. Results carry no key to match against
  3. One slot per input, same length
  4. Null means unresolvable in this subgraph
  5. A missing row must not shorten the list

basics

~20 s

The router matches results to representations by position, not by identity. A reordered or shortened list silently attaches one entity's fields to a different entity. Return exactly one element per input, in order, with null where the entity could not be resolved.

solid answer

~50 s

`_entities` returns `[_Entity]!` — a list with no keys in it, and the router never re-reads identity from the results, because the client's selection set need not even include the key fields. Position is the only correspondence there is. So the resolver's contract is one slot per representation, in the same order, `null` where this subgraph cannot resolve that reference. The bug arrives through the store: `WHERE nct_id IN (...)` returns rows in the database's order and simply omits rows that do not exist, so handing that result straight back both reorders and shortens the list. If representation 14 of 23 is missing, every entity after it shifts up one and the response carries the wrong trial's data under the right trial's id — no error anywhere. Build the result by mapping over the representations and indexing into what the store returned.

code

pseudocode · 10 lines
pseudocode
resolve Query._entities(representations):
    # the store returns rows in ITS order and omits misses,
    # so never hand its result back directly
    ids  = representations.map(r -> r["nctId"])
    byId = trialStore.findAllByNctId(ids).indexBy(t -> t.nctId)

    return representations.map(r ->
        r["__typename"] == "Trial"
            ? byId[r["nctId"]]      # null when the store had no row
            : null)

go deeper

for a junior

Remember the rule as a shape: the list you return has one element for every representation you were given, in the same order, with null where you found nothing.

for a middle

Explain why position is the only link — the returned objects carry no key the router could match on — and show the map-over-representations implementation that is correct by construction.

for a senior

Diagnose the silent version: a 200 response with one entity's fields under another's id, traced to a store that reorders and omits. Cover the null-versus-throw choice and the test that would have caught it.

for a principal

Frame it as a contract that no type checker enforces, and argue for where it should be enforced once — a shared resolver adapter plus a fixture batch containing a miss — rather than re-learned in every subgraph a new team writes.

## The contract, and why it has to be positional The reserved field is declared roughly like this: ```graphql type Query { _entities(representations: [_Any!]!): [_Entity]! } ``` Note the two nullability positions. The **list** is non-null: a subgraph must always return a list. The **elements** are nullable: any individual slot may be null. Now ask how the router pairs a returned object with the representation that asked for it. It cannot read a key out of the result, because the outbound query only selected the fields the plan needs: ```graphql query($representations: [_Any!]!) { _entities(representations: $representations) { ... on Trial { enrollmentTarget primaryOutcome } } } ``` There is no `nctId` in that selection, and there is no requirement that there ever be one. The response is a bare list of objects. The only correspondence available is *index n of the result answers index n of the input*, and that is exactly what the federation specification requires: the returned list must have the same length as `representations` and be in the same order. ## How the alignment gets broken Almost always through the data store, because the natural implementation reads like this: ```pseudocode resolve Query._entities(representations): ids = representations.map(r -> r["nctId"]) return trialStore.findAllByNctId(ids) # WRONG ``` Two independent things go wrong with that line. **Order.** `SELECT ... WHERE nct_id IN (...)` returns rows in whatever order the planner chose — index order, insertion order, whatever the join produced — not the order of the `IN` list. A cache or a search index has even less obligation. **Length.** Rows that do not exist do not come back. A batch of 23 representations where the trial at index 14 was withdrawn from the registry returns 22 rows. Every entity from index 14 onward now answers the wrong representation, and index 22 has no answer at all — some servers then pad, others return a short list the router rejects. The symptom is the ugly kind: no exception, no error entry, HTTP 200, a well-formed response, and `NCT-05128` showing another trial's enrolment target. It survives code review, it survives a unit test written against a two-element batch where both rows exist, and it appears in production only when the batch contains a miss — which is to say, at the worst possible moment. ## The shape that is correct by construction Drive the output from the *input list*, never from the store's result: ```pseudocode resolve Query._entities(representations): ids = representations.filter(r -> r["__typename"] == "Trial") .map(r -> r["nctId"]) byId = trialStore.findAllByNctId(ids).indexBy(t -> t.nctId) return representations.map(r -> r["__typename"] == "Trial" ? byId[r["nctId"]] : null) ``` This is the same discipline a batching loader demands of its batch function — the list you return is a rewrite of the list you were given, not a repackaging of what came back from storage. Two further habits keep it honest: never deduplicate the output (the router may legitimately send the same key twice, and it expects an answer at both indexes), and never sort for tidiness. ## What a null slot means, and what an exception means A null at index n says: *this subgraph cannot resolve that reference*. It is the specified way to report an unresolvable representation — the trial was withdrawn, the id is from a stale cache, the reference was never valid. The router places null where that object sat in the response, and what the client sees from there is the ordinary null-propagation story of the composed schema. Throwing is different and much blunter. If the resolver raises for the whole field, the field error hits `[_Entity]!`, which is non-null, so the whole `_entities` result becomes null and the entire batch of 23 is lost rather than one slot. That is why 'one bad representation kills the fetch' is a real production failure mode: a subgraph that lets a malformed key throw out of the loop turns a single dud reference into a wholesale hole in the response. ## Testing it The test that catches this costs three lines: send a batch of at least three representations where the middle one does not exist, and assert both the length and the identity of each slot. Asserting length alone passes a padded-but-reordered list; asserting identity requires selecting a key field in the test document so you can see which entity landed where. A property test over shuffled batches is cheap here and catches the store-order variant that only shows up under a different query plan.

  • The same representation appears twice in one batch. May the resolver collapse it to a single result?
    No. The output list is indexed against the input list, so a duplicate input needs an answer at both indexes. Deduplicate internally if you like — load the entity once and place the same object at both positions — but never shorten the list. Collapsing duplicates is the same defect as dropping a miss, and it shifts everything downstream by one.
  • How does returning null for a slot differ from raising an error for it?
    A null says the reference is unresolvable here and costs exactly that slot. An error raised for the whole field hits the non-null list type, so the entire batch collapses to null and every entity in it is lost. If a single representation is genuinely bad, prefer a null slot, or a field error scoped to that index, over an exception that escapes the loop.
  • What test reliably catches a misaligned _entities implementation?
    Send at least three representations with a non-existent one in the middle, and assert both the length of the returned list and the identity of each slot — selecting a key field in the test document so the assertion can tell which entity landed where. Shuffling the batch across runs catches the variant where the store's ordering happens to match on small inputs.

It is a coat-check rack, not a mailroom: the router keeps only the ticket numbers, so an attendant who hands back the coats in a different order, or skips a hook, gives everyone after it the wrong coat.

saying these in an interview costs you the question

  • Returns the store's rows directly from the resolver
  • Filters unresolvable references out of the list
  • Assumes the router matches results by key
  • Deduplicates repeated representations in the output
  • Throws for one bad representation in a batch
  • Tests only batches where every entity exists

context