skip to content

You own a GraphQL graph twelve teams publish to — what fails the schema-check gate, what only warns, and who may override?

level: principalimportance: should knowfreq 31%

answer

  1. decide what the gate protects
  2. noise is how a gate dies
  3. the hatch costs visibility, not friction
  4. watch the override rate both ways
  5. one window cannot fit every consumer

basics

~20 s

Hard-fail only breaking edits with observed usage; require a named reviewer for the dangerous class and for breaking edits with credible zero-usage evidence; report nothing for safe additions. Make overrides possible but recorded, and treat the override rate as the gate's health metric.

solid answer

~50 s

Decide first what the gate is *for*: preventing an unannounced break to a live consumer, not enforcing taste. That yields a narrow policy. **Hard-fail** a breaking edit with observed usage inside a valid window — no override, because the evidence says someone is calling it. **Require a named approver** for the dangerous class and for breaking edits backed by credible zero-usage evidence. **Say nothing** about safe additions and description-only edits; noise is how a gate loses its readers. Run the check at publish time, not only on the pull request, because the baseline moves under a branch. Overrides must exist and must be recorded with evidence and an approver — a gate with no escape hatch gets routed around, and a routed-around gate is worse than none because the organisation still believes it is protected. Watch the **override rate**: if a large fraction of checks are being overridden, the policy is miscalibrated, not the teams.

code

pseudocode · 13 lines
pseudocode
policy for graph "parcel-locker":

  SAFE                                 -> pass, report nothing
  description-only / formatting        -> filtered before reporting
  DANGEROUS                            -> pass, requires a reviewer from the owning team
  BREAKING + usage observed in window  -> hard fail, no override offered
  BREAKING + credible zero usage       -> pass, named approver + evidence on the publish

run at: pull request (feedback) and publish (decision, baseline resolved now)

health:
  override_rate       = overrides / checks        # reviewed monthly, both directions
  stale_deprecations  = fields deprecated > 180d  # a graph that only grows has failed too

go deeper

for a junior

You will not set this policy, but know what it implies for you: additive changes sail through, a removal needs evidence and someone's approval, and a green check does not mean a change is safe for callers.

for a middle

Be able to argue why the middle class needs a human rather than a rule, and why a gate that reports cosmetic edits ends up ignored. Know that the check runs at publish time, not only on your pull request.

for a senior

Show the operational reading: which findings you would suppress, what evidence a removal must carry, and how you would detect that people are overriding by reflex rather than by judgement.

for a principal

Own the tradeoff. State what the gate protects and what it deliberately does not, keep approval out of a central queue, treat the override rate as a policy signal in both directions, and name the semantic changes no gate can catch.

## Start from what the gate is for With twelve teams publishing into one parcel-locker graph, the gate is an organisational instrument, and the first move is to say plainly what it protects: **a live consumer must never be broken without a decision someone made on purpose.** That is a narrower goal than "the schema is good", and the narrowness is the point — every additional job you give the gate (naming conventions, description coverage, house style) costs it credibility on the one job that matters. ## The policy, class by class **Safe.** Report nothing. Additive edits are the normal traffic of a healthy graph, and a gate that comments on every one trains people to skim. **Description-only and formatting edits.** Filter them out entirely. This is not pedantry: the fastest way to get a real breaking finding waved through is to bury it under twenty cosmetic ones. **Dangerous.** Pass, but require a reviewer from the owning team who acknowledges it. These are type-legal and cannot be adjudicated by a machine — whether a new enum value hurts depends entirely on how consumers read the field, which no diff can know. Attaching a human is the whole value of the middle class. **Breaking with observed usage.** Hard fail, and do not offer an override in the tool. Evidence says a live consumer selects it; the way forward is deprecation, coordination and time, not a button. **Breaking with credible zero-usage evidence.** Pass with a named approver, and require the evidence — window, coverage, which consumer classes it covers — to be attached to the publish record. This is the case the whole system exists to handle well: without it, the graph can only grow. ## Where the gate runs Run it twice. On the pull request, so the author gets the finding while they still have context. And again **at publish time**, resolving the baseline at that moment, because a branch cleared on Monday can be published on Thursday against a schema two other teams have moved since. Only the publish-time run is a statement about what is actually happening. ## The escape hatch is not optional A gate with no override is a gate that will be bypassed — by publishing out of band, by a team standing up its own endpoint, or by someone quietly disabling the step. Every one of those is worse than an override, because the organisation goes on believing it is protected. So build the hatch, and make its cost **visibility, not friction**: an override requires a named approver and a written reason attached to the publish, and it shows up in a place the graph's owners read. Also decide *who* approves. In a twelve-team graph, routing every override through the platform team makes a queue that becomes the bottleneck for everyone's release. The workable shape is that the owning team approves its own overrides, the platform team owns the *policy* and reviews the record after the fact, and only a specific class — breaking edits on the graph's most widely consumed types — needs a second pair of eyes from outside. ## The number to watch The health metric is the **override rate**: overrides divided by total checks, reviewed on a regular cadence. It reads in both directions. Persistently high, and the policy is miscalibrated — usually the gate is failing on findings that are not real risks, and people are clicking through by reflex, which means the next genuine one goes through too. Persistently zero across all twelve teams, and either the graph is purely additive (check whether anything has ever been deprecated for more than half a year) or teams are avoiding the removals they should be making. A second useful number is the deprecation age distribution — how many fields have carried a deprecation for longer than two quarters. A graph that only ever grows is not a graph with a strict gate; it is a graph whose gate has stopped anyone from cleaning up. ## What the gate structurally cannot do Be explicit about this with the teams, or the gate manufactures false confidence. A structural diff sees types, never meaning. A field that changes units, a list whose default order flips, a cursor whose encoding changes, an argument whose runtime validation tightens — all leave the schema byte-identical and all break consumers. No policy setting catches them; only announcement discipline and a changelog do. The gate is one control among several, and saying so is more principal than tightening it further. ## Consumer tiers Finally, twelve teams do not have one consumer profile. An internal console you redeploy today and a kiosk fleet on a quarterly firmware train cannot share one removal window. Either the policy carries per-consumer-class windows, or the graph is tiered — parts consumed only by fast-moving internal clients evolve on a short clock, parts on the embedded contract move on the fleet's. Pretending one window fits both either paralyses the fast half or breaks the slow half, and choosing which of those you are doing is exactly the call the role owns.

  • Why not simply hard-fail every breaking verdict and remove the override entirely?
    Because the graph then can only grow. Fields nobody calls keep their resolvers, their backing queries and their place in everyone's mental model forever, and the teams who need a legitimate removal will find a route around the gate — out-of-band publishing, a separate endpoint, a disabled step. You lose the removal *and* the visibility. A recorded override with named accountability is strictly better than a bypass you cannot see.
  • The override rate has climbed steadily for a quarter. How do you read that?
    As a defect in the policy first, not in the teams. Sample the overrides and classify them: if most are findings that were never real risks — cosmetic edits, dangerous-class noise on types with one internal consumer — the gate is crying wolf and people are clicking through by reflex, which means the next genuine breaking finding goes through too. Tighten what is reported before tightening what is enforced.
  • How do you keep the platform team from becoming the bottleneck for twelve teams' releases?
    Separate policy from approval. The platform team owns the classification rules, the evidence standard and the review of the record; the owning team approves its own overrides in-line. Reserve cross-team sign-off for a small, named set — breaking edits on the most widely consumed types. Anything that puts one team in the synchronous path of every other team's deploy will be optimised around within a quarter.
  • What would you tell teams the gate does not protect them from?
    Semantic change. Identical types with different meaning — new units, a different default ordering, a re-encoded opaque cursor, tighter runtime validation on an argument — pass every structural check and break consumers exactly as hard as a removal. Those need announcement, a changelog and consumer contact. A team that reads green as 'compatible' rather than 'no structural change' has been given false confidence by the gate itself.

A smoke alarm that goes off when you make toast gets its battery taken out. The design goal is not maximum sensitivity, it is a signal people still act on at three in the morning.

saying these in an interview costs you the question

  • Hard-fails every breaking verdict with no escape hatch
  • Reports formatting and description edits alongside real findings
  • Routes every override through one central approver
  • Treats a green gate as proof of compatibility
  • Applies one removal window to every consumer class
  • Never measures how often the gate is overridden

context