In a federated graph with many services, which subgraph failures should be allowed to fail a whole request?
answer
- Thinner, or misleading?
- Not importance — consequence of silence
- Tier the field, not the service
- Nullability is how the policy is enforced
- Completeness, not uptime, is the number
basics
~20 sOnly the ones whose absence would make the rest of the response wrong or unsafe to act on. Everything else degrades to a hole. Tier the subgraphs by that test, then encode each tier in nullability, budgets and the client contract.
solid answer
~50 sTreat it as a correctness question, not an availability one. For each contributing service ask: if this field is missing, is the remaining response merely thinner, or is it misleading? Enrichment — billing totals, related-case suggestions, activity feeds — is thinner, so it should be nullable and degrade to a hole. A service whose data qualifies whether the rest may be used at all, such as a conflicts or entitlement check in a legal graph, must fail the branch loudly rather than vanish, because a screen that silently omits it looks identical to one that passed. Then make the tiering enforceable: nullability at ownership seams is the mechanism, a schema review gate stops a team quietly upgrading its own field, latency budgets follow the tiers, and the graph's SLO is measured as response completeness per branch rather than uptime.
code
graphql · 11 linestype Case @key(fields: "id") {
id: ID!
caseNumber: String!
# tier two: absence is thinner, degrade to a hole
billing: Billing
relatedCases: [Case!]
# tier one: absence is misleading, fail the branch
conflictCheck: ConflictCheck!
}go deeper
Know that some missing fields are harmless and some are dangerous, and that a screen showing nothing where a warning belongs is worse than a screen that refuses to load. Recall that the schema is where that choice gets recorded.
Explain how a tier becomes real: nullability at the ownership seam decides whether a failure leaves a hole or erases the branch, and the client has to render the hole as explicitly unavailable rather than as an empty value.
Show how you would run it — fault injection per service, per-branch completeness metrics, error redaction at the router, and per-tier latency budgets — and diagnose which of eleven services is degrading a screen from the error paths alone.
Own the whole policy: the test that sorts fields into tiers, the CI gate that stops a team quietly coupling a core type to its own service, the SLO expressed as completeness rather than uptime, and the willingness to say that some branches must never degrade silently.
## Why this is a lead's decision Every individual choice here belongs to a team: the Deadlines team picks its own nullability, the Billing team picks its own timeout. But the property that matters — what a caseworker sees when the graph is degraded — is the product of eleven such choices, and no team can see it. Someone has to own the policy across the graph, and that is the question an interviewer is really asking. ## The test that does the tiering Not importance. Importance is a losing argument: every team believes its field is important, and the debate never converges. Use a sharper test: **if this field is absent, is the response merely thinner, or is it misleading?** In an 11-service legal case-file graph the answer sorts services quickly. *Thinner.* `Case.billing` — unbilled hours. Missing, the screen has a gap where a number goes. A caseworker reads a case file perfectly well without it. Same for related-case suggestions, activity feeds, document thumbnail counts. *Misleading.* `Case.conflictCheck`. A case screen that shows no conflicts because the Conflicts service was down is byte-for-byte indistinguishable from a case screen that shows no conflicts because there are none. The absence changes what a human will do next, and it does so silently. The same holds for anything expressing an entitlement, a legal hold, or a restriction: the missing state is the permissive-looking state. So the tier is decided by the consequence of a silent absence, and it lands on a *field*, not a service: one subgraph can own fields in both tiers, and the tiering must follow the field. ## Encoding the tiers so they hold A policy that lives in a wiki decays in a quarter. Encode it. **Nullability is the mechanism.** Tier-two fields are nullable at the ownership seam, so a failure leaves a hole. Tier-one fields keep their Non-Null marker deliberately — accepting that a failure erases the branch, which is exactly the intent, because a blank screen with an error is safer than a plausible-looking one. This inverts the usual advice, and saying so explicitly is what separates a policy from a rule of thumb. **A schema gate enforces it.** A composition-time check in CI that flags any Non-Null field added to a type the subgraph does not own, requiring an explicit sign-off with the tier recorded. Composition itself will never raise this: a valid type is a valid type, and availability coupling is invisible to it. **Budgets follow the tiers.** Tier-two fields get tighter deadlines, because giving up early on them is cheap. Tier-one fields deserve the slack, since abandoning them fails the branch anyway. **The client contract is written down.** Partial data is a normal response, and the client must be built to render per branch with an explicit unavailable marker where a tier-two hole appears — never a blank space that reads as zero. For tier-one, the contract is the opposite: no silent fallback, no cached last-known value. ## What leaves the router Errors that reach clients are part of the same policy. A subgraph's raw message can carry its service name, an internal hostname, a driver error naming a table. Forwarded verbatim, an unauthenticated caller can map an eleven-service topology by sending badly shaped operations and reading the wreckage. The rule: a stable, client-meaningful code in the error, a correlation id, and everything else kept server-side. Neither the GraphQL specification nor the composition specification says anything about this — it is entirely your policy, which is why it is so often nobody's. ## Measuring the right number Status-code availability is meaningless here: a graph that has lost four of eleven services still answers 200. The number to run against is **completeness** — the share of operations returning no error entries at all, split by the `path` prefix that identifies the branch, and separately the share of *tier-one* operations that returned complete. Those two curves are what tell you whether the degradation strategy is working, and they are what an executive should see instead of a green uptime tile. ## Proving it before the incident The policy is a hypothesis until it is exercised. Fault injection in a pre-production graph — kill one subgraph, screenshot the screen — is worth more than any document, because the failure mode people actually get wrong is aesthetic: a hole that renders as a plausible zero. Run it per service, at least when a service joins the graph and when a core type gains a contributed field. ## The honest tradeoff to name Degradation buys availability with ambiguity. Every hole you allow is a place where a user cannot tell missing from empty, and eleven services' worth of holes is a lot of ambiguity to hand to someone making a decision from the screen. The judgement is choosing where ambiguity is cheaper than unavailability, and being willing to say that in some branches — the ones that gate a decision — it never is.
- How much of a failing subgraph should appear in the error a client receives?A stable code the client can branch on, a human-readable but non-specific message, and a correlation id. Not the service name, not its host, not the driver text a data store produced. Forwarded verbatim, those turn every malformed operation into a free map of an eleven-service topology. Keep the detail on the server, keyed by the correlation id, where your own on-call can reach it and a caller cannot.
- A team wants to add a Non-Null field to a core type from a new, unproven subgraph. What do you say?Nullable at the seam until the service has an operating history, with the inner type kept strict so clients still get a well-shaped object when it resolves. Tightening a field from nullable to Non-Null later is a safe, non-breaking direction for callers; loosening it after clients have been generated against the strict type is not. The asymmetry means the cautious choice costs nothing and the confident one is hard to undo.
- How would you show an executive that the degradation policy is working?Two curves, neither of them uptime. The share of operations returning with no error entries — response completeness — broken down by the branch that failed, and the share of tier-one operations that returned complete. The first shows how much the graph is degrading; the second shows whether it is degrading in places that are safe. A status-code availability tile stays green through a four-service outage and should not be on the dashboard at all.
- Should the router ever serve a stale cached value in place of a failed branch?For tier-two fields, sometimes, and only if the response says so — a timestamp on the branch, so the client can label it rather than present stale data as current. For tier-one fields, never: a cached entitlement or conflict result is exactly the plausible-looking answer the tiering exists to prevent. The deciding question is the same one used for the tiering: is stale merely thinner, or is it misleading?
A case bundle may arrive without the billing tab and still be usable, but if the conflicts memo is missing the clerk must stop rather than hand over a bundle that looks cleared.
saying these in an interview costs you the question
- Tiers services by importance rather than by consequence of absence
- Assumes every failure should degrade to partial data
- Leaves the policy in a document with no schema gate
- Forwards raw subgraph error text to clients
- Measures the graph's health with status-code availability
- Renders a missing branch as a zero with no marker