skip to content

A federated subgraph deployed hours ago, but its new field is still missing from the graph — why?

level: seniorimportance: should knowfreq 40%

answer

  1. A whole supergraph or nothing
  2. The graph freezes, it does not fall over
  3. Code shipped, contract did not
  4. Someone else's conflict blocks yours
  5. Alert on the age of the composition

basics

~20 s

Almost always the schema failed to compose. Composition is all-or-nothing, so no supergraph was produced and the router kept serving the last one that composed. The service is running new code while the graph still describes the old schema.

solid answer

~50 s

Composition emits either a complete supergraph or a list of errors — never a supergraph with the offending part removed. So a conflict does not take the graph down; it freezes it. The router keeps serving the last supergraph it successfully received, every existing query keeps working, and the only visible symptom is that something new is missing. Meanwhile the service itself deployed, which is the dangerous half: the code, its database and its behaviour have moved on while the graph still advertises the previous schema. Diagnose it by comparing the schema the running subgraph reports with the definition of that type in the composed supergraph; if they differ, look at the composition result for the most recent publish. The failing conflict is often in a different subgraph entirely, since one bad publish blocks every publish after it.

code

graphql · 9 lines
graphql
# Irrigation subgraph, running in production
type Mutation {
  openValve(paddockId: ID!, seconds: Int!, requestId: String!): ValveCommand!
}

# The composed supergraph the router still serves
type Mutation {
  openValve(paddockId: ID!, seconds: Int!): ValveCommand!
}

go deeper

for a junior

Understand that deploying your service and getting your schema into the graph are two different events. If a new field is missing, check whether the graph was rebuilt, before suspecting the router or a cache.

for a middle

Explain the all-or-nothing outcome and what the router does with it: no supergraph is produced, the previous one keeps serving, and the symptom is a missing field rather than an error anyone sees.

for a senior

Show the diagnosis and the risk. Compare the running subgraph's schema with the composed one, find whose conflict blocked the publish, and reason about code that is live while its contract is not — especially writes that depend on an argument nobody can send.

for a principal

Own the release design: compose before rolling code, page on a stale composition, roll back a failing publish rather than leaving the whole graph frozen, and set the rule that new code must behave correctly under the previous contract.

## Failing closed means freezing, not breaking Apollo Federation composition has exactly two outcomes: a complete supergraph schema, or errors. There is no third mode where it emits a supergraph with the conflicting field, type or subgraph left out. That is a deliberate choice — a graph that quietly shrinks would break clients with no signal and no owner — but it produces a failure shape that catches teams out, because nothing appears to be wrong. When composition fails, the router does not stop and does not start rejecting traffic. It keeps serving the last supergraph it successfully obtained. Every existing operation still validates and still executes. Dashboards are green. The only symptom is negative: a field, an argument or a whole type that a team shipped this morning is not in the graph, and a client's query naming it fails validation with an "unknown field" style request error rather than anything that points at composition. ## The dangerous half is skew, not absence The missing field is an inconvenience. The real hazard is that the two halves of the deploy have separated: the subgraph is running new code, new persistence and new behaviour, while the supergraph the router plans against still describes the previous schema. Everything the router sends is planned from a stale contract. A farm-sensor graph shows how expensive that gets. The Irrigation team ships an idempotency argument on a valve command so that a retried request is recognised rather than replayed: ```graphql # Irrigation subgraph, newly deployed type Mutation { openValve(paddockId: ID!, seconds: Int!, requestId: String!): ValveCommand! } ``` The same publish also carried an unrelated change to a shared measurement type that conflicts with another subgraph's copy, so composition fails and no new supergraph is produced. The router's schema still exposes `openValve(paddockId:, seconds:)` with no `requestId`, so no client can send one. That evening a scheduling client times out mid-command and retries, exactly as designed — and because the deduplication the new code depends on is keyed on an argument the graph will not let anyone send, the valve command is written twice and the paddock is irrigated twice. The subgraph's own tests all passed. The defect lives entirely in the gap between what the service now expects and what the graph still advertises. ## Diagnosing it The sequence is short and worth being able to recite: 1. **Compare the two schemas.** Ask the running subgraph for the schema it reports, and pull the same type out of the composed supergraph. If the field is in the first and not the second, this is a composition problem and not a routing or caching one. 2. **Look at the composition result for the latest publish**, not at the subgraph. The errors name the conflicting declarations. 3. **Check whose conflict it is.** On a 62-subgraph supergraph this is the step people skip. Composition runs over the whole set, so one team's unfixed conflict blocks every publish after it — your field is missing because of a type mismatch in a service you have never deployed to. 4. **Check how long the graph has been frozen.** The gap between the newest published subgraph schema and the composed one is the real blast radius: every schema change from every team since the failure is also missing. ## Operating a graph so this cannot bite - **Compose before you roll the code.** Running composition against the candidate schema as part of the pipeline turns a silent freeze into a failed build owned by the team that caused it. - **Treat a composition failure as a page, not a warning.** It is the one production condition whose symptom is that nothing happens. Alert on the newest published subgraph schema not being present in the composed supergraph, and on the age of the last successful composition. - **Deploy so the two halves can separate safely.** If code can ship before its schema lands, it must behave correctly with the old contract — an argument the client cannot yet send must have a defined behaviour when absent, and a write that depends on a new argument for correctness has to be safe without it. - **Keep failed publishes from queueing.** Because one conflict blocks everyone, an unfixed composition failure is a shared outage of the release process. Roll the offending subgraph schema back rather than leaving it failing while its owner investigates. ## Say what is specified and what is convention The all-or-nothing outcome of composition is federation's composition specification. What happens *around* it — where composition runs, whether the router is fed by a registry or by a file, whether a failure pages anybody — is operational convention that differs from one setup to the next, and the honest answer describes the mechanism and then says which parts are choices your platform made.

  • Why does composition refuse to publish a supergraph with just the conflicting field removed?
    Because that turns one team's mistake into a silent break for every client selecting the field, with no error anyone owns. A graph that shrinks quietly also erodes: each unfixed conflict would remove a little more. Refusing to emit anything keeps the failure attached to the change that caused it, at the cost of blocking publishes until it is fixed.
  • What alert would have caught this before a client did?
    Two. One comparing the newest published subgraph schema against the composed supergraph, firing when a publish is not represented in the served graph. One on the age of the last successful composition, because the symptom of the failure is that nothing changes. Both are cheap, and neither depends on a client noticing a missing field.
  • The subgraph is deployed and the schema will not compose for a while. What do you do first?
    Get the graph unblocked before you debug. Roll the offending schema publish back so composition succeeds again, since until it does every other team's changes are stuck behind it. Then confirm the deployed code is safe against the old contract — particularly any write whose correctness depends on an argument clients cannot yet send — and fix the conflict at leisure.

The presses keep running yesterday's edition because today's copy failed proofing — nobody sees an error, they just never see the new page.

saying these in an interview costs you the question

  • Expects the router to serve a partially composed graph
  • Thinks a composition failure takes the graph down
  • Looks only at their own subgraph for the conflict
  • Assumes deploying the service publishes the schema
  • Ignores that the code now expects a contract clients cannot use
  • Treats a failed composition as a warning to fix later

context