How do you set a breaking-change policy for a GraphQL schema with un-updatable clients?
answer
- Start from the callers, not the schema
- Upgrade latency of the slowest client
- Different surfaces can carry different rules
- Publish the client's half too
- Never breaking has its own bill
basics
~10 sDerive the rules from consumer upgrade latency, not from the schema. Segment the graph, freeze only what slow clients reach, publish the client obligations that make additive change safe, and name the escape valve.
solid answer
~50 sAn edit is breaking relative to somebody, so start with a census of consumers ranked by **upgrade latency**: an internal web client redeploys in minutes, a native build for site coordinators lives on devices for a year, a sponsor's integration moves on a contract, and a monthly batch job can hide a break past your rollback window. The policy is a function of the longest tail *per surface*, not per schema — treating one endpoint as one compatibility regime drags the whole graph down to the slowest caller. Then publish the client's half of the contract: a fallback branch for unrecognised output enum values, no assumption that a nullable field is populated, named operations so usage is attributable. Finally price the alternative honestly. 'Never break anything' buys parallel near-duplicate fields nobody can choose between, so name the escape valve — who authorises a break, on what evidence, over what window, and which legal or security reasons override it.
go deeper
You are unlikely to be asked this, but understand the premise: whether an edit is breaking depends on who is calling and how quickly they can ship a fix, not on the schema alone.
Be able to say why an internal web client and a shipped mobile build deserve different treatment, and why documents you cannot see make a removal a judgement call rather than a lookup.
Bring the operational half: which callers you can actually attribute traffic to, what evidence would clear a removal, and what client-side obligations you would insist on before calling additive change safe.
Own both costs. Argue the segmentation of the graph, the escape valve and its authoriser, and the long-run price of a schema that can only ever grow — a candidate who only defends compatibility has answered half.
## Policy starts with the callers, not the schema The question "is this edit breaking?" has no answer that lives inside the schema. An edit is breaking relative to somebody, and the only thing that turns a class of edit into a *policy* is knowing who that somebody is and how fast they can move. So the first artefact is not a rulebook, it is a census of consumers with one column that matters: **upgrade latency**, the time between you shipping a change and the last caller that must adapt having adapted. Typical bands, using a clinical-trial registry graph as the example: * **Your own web client.** Redeploys in minutes, from a repository you control. Its documents are knowable and changeable in the same change as the schema. * **A native application for site coordinators.** Old versions persist for months; a hospital-managed device fleet may sit on a build for a year or more. You cannot force an upgrade and you often cannot even measure adoption precisely. * **A sponsor's partner integration.** Changes when the sponsor schedules engineering work, which is a contractual conversation, not a technical one. * **Batch and reporting jobs.** Frequently the forgotten class: they run monthly, so a break can lie dormant past your rollback window. The policy is a function of the longest tail that touches a given part of the graph. Which is the first real insight to state in an interview: **write the policy per surface, not per schema.** A region of the graph that only your own web client selects can evolve aggressively, because breakage is a same-day fix by the same team. The types a two-year-old device build reaches are effectively frozen except for additions. Treating one endpoint as one compatibility regime forces the whole graph down to the slowest caller's rules, and that is how schemas ossify. ## Publish the client's half of the contract Compatibility is a two-sided property, and most organisations only ever write down the server's side. State the obligations that make additive change genuinely safe, and treat them as a condition of being a supported consumer: * Handle an unrecognised value in any output enum with a defined fallback rather than an exception. * Never assume a nullable field will be populated, and never assume a list is non-empty. * Name every operation, so usage can be attributed to a caller rather than to an anonymous document. * Select only the fields you use, because a selection is a dependency you are asking to be preserved. A client that meets those four is one you can ship additively to for years. A client that meets none of them makes even a value addition an incident, and no amount of server-side discipline fixes that. ## Decide the price of never breaking anything The reflexive answer — "we simply never make breaking changes" — is a position, not a policy, and it has a bill. Its currency is parallel surface: `investigator` beside `leadInvestigator` beside `investigatorV2`, three shapes with drifting semantics, all resolved, all tested, all needing an answer to "which one do I use?" for every new engineer. Every field you cannot remove is a field you maintain forever, and the aggregate cost lands not on the team that added it but on everyone who reads the schema afterwards. So a serious policy names the escape valve as explicitly as the default. Something like: additive by default; a break to a frozen surface requires a named owner, evidence about who is calling, a window tied to the version-adoption curve rather than to the calendar, and a decision recorded about what happens to callers who do not move. And it names the override — a field that leaks participant identifiers under a privacy obligation, or a value that is actively wrong — where correctness or legal exposure beats compatibility and the break happens now, with communication rather than a window. ## What a principal is actually being asked The interviewer is not checking whether you can list breaking-change classes; that is a mid-level question. They are listening for four things: that you derive the rules from consumer upgrade latency rather than from taste; that you are willing to segment the graph rather than applying one regime; that you can articulate the cost of *not* breaking as clearly as the cost of breaking; and that you know a policy without usage evidence is a document nobody can enforce. The candidate who says "never break anything" and stops has answered half the question — and the half they skipped is the one that produces a schema people can still understand in three years.
- How long a window would you give a removal when the slowest caller is a native application?Tie it to the version-adoption curve rather than the calendar: the window ends when the share of traffic from builds that still select the field falls below a threshold you agreed in advance, with a floor for callers you cannot measure. A fixed ninety days is arbitrary; it either strands a long tail or freezes the schema for months longer than the data warrants.
- Would you ever break the schema deliberately rather than add a parallel field?Yes, in two cases. When the existing field is actively wrong or unsafe — leaking an identifier you are obliged to stop returning — correctness beats compatibility and the break happens with communication rather than a window. And when the parallel-field bill is already too high: a surface where three near-duplicates exist is costing every future reader more than one coordinated migration would.
- How does the policy change for an internal schema with three known consumers?It gets far more aggressive, because upgrade latency collapses. With every document in repositories you control, a rename is a same-change edit across callers and a removal is a scheduling problem rather than a contract. The discipline that remains is knowing the callers are still exactly those three, which is an ownership question, not a technical one.
- What makes a breaking-change policy unenforceable in practice?Not knowing who calls what. Without attributable usage a policy reduces to intuition: nobody can clear a removal, so nothing is ever removed, and the schema grows monotonically. The second failure is having no named owner for exceptions, which turns every hard case into a debate rather than a decision.
saying these in an interview costs you the question
- Answers 'never break anything' with no cost acknowledged
- Writes one compatibility rule for every consumer
- Assumes all clients redeploy when the server does
- Offers a schema version number as the whole answer
- Ignores legal or security breaks that cannot wait
- Forgets batch and reporting callers entirely