How would you set one count contract for connections across dozens of independently owned subgraphs?
answer
- Composition checks shapes, never meanings
- The same field name, three different promises
- Encode the promise in the name
- Enforce at publish, not in review
- Adding one is cheap; removing one is not
basics
~20 sFix meaning by name, not by hope: one name reserved for exact counts, distinct names for capped and estimated ones, mandatory descriptions, nullability everywhere. Enforce it with schema linting at publish time, because composition never checks semantics.
solid answer
~50 sComposing a supergraph from many subgraphs validates names, types and nullability — never what a number means. So in a 62-subgraph supergraph, `totalCount: Int!` can be an exact tally in one subgraph and a capped estimate in another, and no client can tell. The lever is a naming convention promoted to a rule: reserve the conventional name for a count that is exact at read time, require any capped or estimated count to carry a different name that says so, require a description stating the flavour and the cost class, and require the field to be nullable so a subgraph can decline under load without failing the connection. Enforce it as a lint at schema publish rather than in code review. Then instrument field usage, because the case for retiring an expensive count is almost always evidence that nobody reads it — and retiring one is a breaking change you must stage.
code
graphql · 18 linestype TranslationUnitConnection {
edges: [TranslationUnitEdge!]!
pageInfo: PageInfo!
"""EXACT at read time, same filter and viewer scope as the page.
Cost class: live-unbounded. Null if the time budget is exceeded."""
totalCount: Int
}
type SegmentMatchConnection {
edges: [SegmentMatchEdge!]!
pageInfo: PageInfo!
"""CAPPED at 10000; countIsCapped tells the client the ceiling was hit.
Cost class: bounded."""
cappedTotalCount: Int
countIsCapped: Boolean!
}go deeper
Take away the core idea: two schemas can use the same count field name for different promises, and only the field's description tells you which you are reading.
Be able to say what composition and introspection do and do not check, and why nullability on a count field is a deliberate escape hatch rather than sloppiness.
Show how you would enforce this in practice: lint at schema publish, mandatory descriptions, per-field usage and timing, and a deprecation window before any count field is removed.
Own the tradeoff between one enforced vocabulary and many teams' autonomy, and be explicit that a count is cheap to add and expensive to retract — which is why the contract must exist before the fields do.
## Why this is a governance problem, not an implementation one One team can decide what its count means in an afternoon. Sixty-two teams publishing into one supergraph cannot, because nothing in the machinery adjudicates meaning. Composition under a federation specification checks that field names, types, nullability and directive usage are consistent enough to plan queries against; it has no opinion about whether an `Int!` is an exact tally, a tally capped at 10,000, or a statistical estimate that drifts by four percent. Introspection cannot expose the difference either — a client sees `totalCount: Int!` and nothing more. The result, left alone, is predictable. Every team reaches for the same conventional field name because it is the one their consumers recognise, and each attaches whatever semantics their storage made affordable. A client team then writes one shared list component against `totalCount` and ships a UI that renders an exact figure for some connections, a capped ceiling for others, and a wrong-looking estimate for the rest. Nobody lied; the vocabulary simply had no owner. ## The contract worth writing The contract has to be enforceable by a machine reading SDL, because that is the only checkpoint that scales past a handful of teams. Four rules do most of the work: **One name, one meaning.** The conventional name is reserved for a count that is exact at read time, over the same filter and the same authorization scope as the page. Any other flavour gets a different, self-describing name. This is the whole contract in one line, and it is the rule to state first in an interview. **Cost is declared, not discovered.** Every count field carries a description naming its flavour and its cost class — precomputed, capped at a stated ceiling, live and unbounded. That description is the only channel through which semantics can reach a consumer, so make it mandatory. **Every count is nullable.** Non-null is a promise the field can always be produced. Under a federation router, a field that fails inside a subgraph can propagate and cost the caller a whole branch of the response; a nullable count lets a subgraph decline under a time budget while the page still returns. Nullability is also nearly impossible to loosen later without breaking clients, so make it the default from the start. **No unbounded count on a nested connection by default.** A connection under a list of parents multiplies the count, so the rule is that a per-parent count must be precomputed, capped, or absent. ## Making it stick Write the rules as schema lint that runs when a subgraph publishes to the registry, not as guidance in a wiki. A lint that rejects a count field with no description, or a field named for an exact count whose description says "approximate", is worth more than any number of review comments. Grant exemptions explicitly and record them, so "we had a reason" survives the engineer who had it. Then instrument. Per-field usage from the router or the subgraphs tells you which clients select which count, and field-level timing tells you what each one costs. Those two numbers together are what turn an argument into a decision: a count that costs 1,447 ms and is selected by one internal dashboard that renders it as "about 200k" is a count you can downgrade to an estimate. ## The part people underestimate: you cannot take it back cheaply Once a count is published, retracting it is a breaking change with an unusually long tail. A document that still selects a removed field fails validation outright — the whole request is rejected before execution, so the client does not degrade, it goes dark. Worse, a client pinned to a removed field may be one you cannot redeploy: a build-time-registered document, a mobile release in the wild, an automation someone wrote against the endpoint two years ago. Deprecating the field, watching usage fall to zero, and only then removing it is not ceremony, it is the only safe order. This is exactly why the contract is worth setting before the fields exist. A count field is the cheapest thing in the world to add — one integer, one resolver — and among the more expensive to remove. ## The judgement to voice The unserious answer is "ban exact counts". A total is a real product requirement and some of them are cheap. The lead's job is to make the price visible at the point of design and the choice reversible afterwards: named grades so nobody is misled, mandatory descriptions so the price is legible, nullability so load has an escape hatch, lint so the rules hold across teams who will never meet, usage data so retirement is evidence-driven, and deprecation windows so retirement does not break clients you cannot reach.
- What exactly happens to a client whose shipped document still selects a count field you removed?The request fails validation before execution: the field is not defined on that type, so the server returns a request error and no data at all. It is not a degraded response with one missing key — the entire operation is rejected. That is why removal must follow deprecation plus usage evidence, and why a client you cannot redeploy, such as a released mobile build or a registered build-time document, effectively pins the field until that client is gone.
- Would you standardise counts by putting them in a shared schema module every subgraph imports?A shared definition helps with shape but not with semantics: the type is common, yet each subgraph still writes the resolver behind it, so one can be exact and another capped. Shared modules are still worth it for consistent naming and descriptions, but they must be paired with the publish-time lint and with per-field usage data. Standardising the vocabulary is easy; the enforcement is the actual work.
- How do you decide whether a specific expensive count should exist at all?With two measurements and one product question. Field usage tells you which clients select it and how often; field timing tells you what it costs when they do. The product question is what the number is used for — a rendered approximation, a navigation control, or a figure of record. A count that is costly, rarely selected, and displayed as a rounded approximation should become an estimate; one that drives numbered navigation has to be paid for or the navigation changed.
It is like sixty-two suppliers stamping "net weight" on their crates, each measuring it differently. The fix is not to inspect every crate but to define the term and refuse shipments that do not state which definition they used.
saying these in an interview costs you the question
- Assuming composition validates what a number means
- Trusting a shared field name to imply shared semantics
- Making count fields non-null across every subgraph
- Removing a published count field without usage evidence
- Enforcing conventions only through code review
- Believing a removed field merely returns null to old clients