skip to content

A GraphQL schema check flags a field removal as breaking, yet recorded usage shows zero calls — what would convince you to ship it?

level: seniorimportance: must knowfreq 52%

answer

  1. the verdict is about possibility
  2. zero, over how long, at what coverage
  3. the slowest client sets the window
  4. sampling cannot prove absence
  5. record the override and the evidence

basics

~20 s

The verdict says a document could select the field, not that one does. Usage evidence overrides it only if the observation window outlives your slowest-updating client, the traffic was not sampled, and the recording covers every path that reaches the schema.

solid answer

~50 s

A breaking verdict is a statement about **possibility**: some valid document selects this field, so removing it could invalidate one. Usage evidence is what turns that into a fact about your traffic — but zero observations only means something if the evidence is trustworthy on three axes. **Window**: it must be at least as long as the release cycle of your slowest-updating consumer; a fourteen-day window says nothing about kiosk firmware that ships quarterly. **Coverage**: sampled traffic cannot prove absence — at a 1,200-request-per-minute peak, a 1% sample will routinely miss an operation that runs a few times a day. **Reach**: only what arrived at this server was counted, so responses served from a cache layer, queued offline clients and ad-hoc scripts may be invisible. If all three hold, ship it: deprecate first, let the window run, remove, and record the override with the evidence and a named approver so the decision is auditable.

code

pseudocode · 11 lines
pseudocode
field    = "Compartment.sizeCode"        # deprecated 41 days ago
observed = usage_count(field, window)    # 0

window   = 41 days
coverage = 1.0                           # unsampled; a 1% sample at 1,200 rpm peak proves nothing
slowest  = kiosk_firmware.release_interval   # ~92 days  <-- window does not cover it

if observed == 0 and coverage == 1.0 and window >= slowest:
    remove(field)                        # record evidence, window, coverage, approver
else:
    keep_deprecated(field)               # here: coordinate with the firmware release train

go deeper

for a junior

Know that a breaking verdict says a client could be affected, not that one is, and that the safe order is deprecate, wait, then remove — never remove first and watch what happens.

for a middle

Be able to state what makes a zero-usage reading meaningful: how long it was observed for, whether all traffic or a sample was recorded, and whether every consumer reaches this server at all.

for a senior

Demonstrate the judgement: tie the window to the slowest consumer's release cycle, refuse to reason about absence from sampled data, name the invisible callers, and describe the rollback and error-rate watch after the publish.

for a principal

Own the policy behind the individual call. Decide what evidence is sufficient to override a gate, who is allowed to sign it, how the decision is recorded for the next incident, and what you do for consumer fleets you cannot force to upgrade at all.

## What the verdict actually asserts Read the finding literally. A structural check has no traffic data; it knows only the type system. "Breaking: `Compartment.sizeCode` removed" means *there exists a valid document that would stop being valid* — it does not mean anyone sends one. Treating that as a hard stop means a graph that can only grow, and a schema that only ever grows is its own production problem: dead fields keep resolvers, indexes and on-call knowledge alive for nobody. So usage evidence is not a way of cheating the gate. It is the missing half of the input. ## Axis one: the window must outlive the slowest client Usage is observed over a window, and the window has to be long enough that every *live generation of every consumer* has had a chance to appear in it. That is a property of your clients' release cadence, not a round number someone liked. On a parcel-locker graph the consumers are wildly different. A web console redeploys several times a week — a few days covers it. A courier mobile app updates on the stores' cadence, with a long tail of users who never update. The locker kiosks run firmware pushed roughly once a quarter, and a bank that has been offline for maintenance may not have called in weeks. A 30-day window covers the console comfortably, covers the courier app poorly, and covers the kiosks not at all. The practical version of the rule: **window ≥ the longest interval over which a consumer generation you cannot force-upgrade might reappear.** Where that number is unacceptably long — a quarter, a year — the honest answer is that the window is not your instrument, and the field's removal has to be coordinated with the fleet's release train rather than inferred from silence. ## Axis two: absence of evidence needs full coverage Sampling and proof of absence do not mix. If operations are recorded at 1% — a reasonable rate at a 1,200-request-per-minute peak — then an operation issued a handful of times a day has a comfortable chance of never being sampled across an entire month. Zero observations under sampling is not zero usage; it is a number consistent with both zero and "once a day". Before acting on a zero, establish what fraction of operations was recorded. If the answer is "we sample", either raise the rate for the specific decision, or reason about the bound the sample supports rather than about zero. A useful sanity check is a positive control: pick a field you *know* is rarely used and confirm the record shows it. A record that shows zero for everything rare is measuring nothing. ## Axis three: reach — what never arrived is never counted Usage is recorded where the operation is executed. Anything that satisfies a caller without reaching that point is invisible: a response served from a caching layer in front of the service, a client that queues operations offline and has not drained, a batch job that runs on the first of the quarter, an internal script someone wrote against the endpoint by hand. None of them show up, and the last one is the most common real-world surprise, because it belongs to nobody and is in no repository. Do not try to enumerate consumers from memory. Ask the inverse question — *what would have to be true for this field to be used and not appear in the record?* — and check each answer. ## Sequencing, not a single decision Even with all three axes satisfied, removal is a sequence: 1. **Deprecate first** so the field carries a machine-readable warning and a stated replacement, and consumers have something to act on. 2. **Let the window run** from the deprecation, not from the day someone decided to delete it. 3. **Remove**, with the check's breaking verdict explicitly overridden and the override recorded — the evidence, the window, the coverage, the approver. This matters far more than it sounds: the next person to ask "why did this disappear" is often an incident responder, and "a green pipeline" is not an answer. 4. **Watch the request-error rate after the publish**, per client where you can. A removal shows up as documents failing validation, which is loud and immediate rather than a slow degradation. Have the schema rollback path ready, and know that rolling the schema back does not un-fail the requests already rejected. ## The judgement to voice in an interview The interesting answer is not "check usage". It is **naming what would make the zero untrustworthy** — short window against a slow fleet, sampled traffic, a caching layer absorbing traffic, an unknown consumer — and then saying which of those you can actually rule out in the system in front of you. A candidate who says "usage was zero for thirty days, so we removed it" without qualifying the window against the client fleet is describing how the outage happened, not how the decision was made.

  • How would you decide the observation window rather than defaulting to thirty days?
    Enumerate the consumer classes and take the worst release interval you cannot force. A web console you deploy yourself needs days. A mobile app on store cadence with a slow-updating tail needs months, or per-version evidence. Embedded clients need the window to span at least one full firmware cycle plus the time an offline unit might stay dark. If the resulting number is unacceptable, the removal is a coordination problem with the release train, not a data problem.
  • You removed the field and error rates jumped. What is the failure mode and what does rolling back actually fix?
    Documents selecting the removed field now fail validation, so the entire request is rejected — not a null field, the whole operation. Republishing the previous schema and redeploying stops new failures immediately, which is why the rollback path matters more than the forward one. It does not repair anything already rejected, and any client that hard-failed and cached an error state may need its own recovery.
  • Is there a safer intermediate step between deprecating a field and deleting it?
    You can stop populating it — have it always return null — but be honest that this is a behaviour break, not a safe step: it is invisible to a structural diff and can be worse than removal because callers get a plausible-looking wrong answer instead of a loud error. It only helps when the field is nullable and consumers demonstrably tolerate null. Otherwise the honest options are keep it or remove it.
  • The check has no usage data at all for this graph. Does that change the answer?
    Yes — it means you have no evidence, not that the field is unused, and the two must not be conflated. Without a record the only sound routes are coordination (announce, deprecate, confirm with each known consumer) or instrumenting first and waiting out a real window. Removing on the grounds that nobody complained is deciding on absence of complaint, which is the weakest evidence available.

Watching an empty road for a week proves the road is unused only if you were watching continuously, and only if the truck that uses it does not come once a quarter.

saying these in an interview costs you the question

  • Treats zero observations as proof the field is unused
  • Uses a fixed thirty-day window regardless of client cadence
  • Draws conclusions about absence from sampled traffic
  • Forgets clients whose traffic never reaches this server
  • Removes without deprecating or announcing first
  • Overrides the gate without recording evidence or approver

context