Any team in your organisation can relax its own schema name's compatibility mode to unblock a release. What does that autonomy cost?
answer
- the mode is just configuration
- cheapest action defeats strongest guarantee
- cost lands on the consuming teams
- strict with no exit breeds invisible bypasses
- attributed and expiring, not forbidden
basics
~20 sThe guarantee stops being a property of the platform and becomes a per-name, per-day unknown. Consumers can no longer reason about what is enforced anywhere, the cost of an override lands on teams who never made it, and the evidence usually disappears with the setting.
solid answer
~40 sA compatibility mode is configuration, so the gate is exactly as strong as the control over who may change it. When the producing team holds that lever, three things follow. The guarantee becomes local and unknowable: a consumer team cannot tell whether the stream it depends on was enforced last Tuesday. The cost is externalised — the team that relaxes the mode ships on time, while the teams that decode pay in failed consumers and dead-letter backlogs. And the evidence is thin: a setting flipped and flipped back leaves nothing behind unless the registry records it. The answer is not to remove the lever, because a hard central mandate breeds bypasses like new names and side-channel writers. It is to make the override loud, attributed, time-boxed and visible to the dependants it affects.
go deeper
Recall that the compatibility rule is a setting, so someone can turn it off, and that doing so affects everyone reading the stream.
Explain why a relaxed setting restored the next day leaves no evidence, and why the incident it causes is then blamed on the registry.
Argue the asymmetry: the producing team captures the schedule benefit while the consuming teams absorb the failures, which is a control-placement problem.
Commit to a design — default strict, keep the escape hatch, make it attributed and expiring — and defend it against the strict-mandate alternative on the grounds of invisible bypasses.
This is a governance question wearing a technical costume. Nothing about the comparison changes; what changes is who is allowed to switch it off, and what happens around them when they do. ## The lever A mode is a setting on a name. Relaxing it is a small, legitimate-looking action, usually taken under deadline pressure by a competent engineer who has convinced themselves the change is fine. It requires no code review, leaves no diff in a repository, and — unless the registry is instrumented — no durable trace. That is the asymmetry a lead has to reason about: the **strongest guarantee in the pipeline is defeated by the cheapest action in it**. ## What the organisation actually loses - **A global property becomes a local one.** 'Our streams are compatibility-checked' stops being true of the platform and becomes true of whichever names happen to be configured today. Nobody can state the guarantee without an audit. - **The cost is paid by people who did not choose it.** The producing team gets its release; the consuming teams get decode failures, dead-letter backlogs and a reprocessing job. Externalised costs are systematically over-incurred, and that is a predictable outcome of the control placement, not a moral failing. - **The evidence evaporates.** Relaxed to ship, restored on Monday, and the incident three weeks later has no explanation. Post-incident review finds the mode correct and concludes the registry failed. - **The precedent spreads faster than the practice.** Once an override is known to be available and unremarked, it becomes the normal response to schedule pressure across teams that never discussed it. ## Why 'take the lever away' is the wrong reflex A hard central mandate with no escape hatch is not the safe option; it is a different risk: 1. **It creates bypasses.** A rejected registration is answered with a new name, a second stream, or a writer that never consults the registry — all of which are worse than a recorded override, because they are invisible. 2. **It blocks legitimate work.** Genuinely breaking changes are sometimes correct: a contract that was wrong from the start, a stream with a known and tiny consumer set, a coordinated migration where both sides ship together. 3. **It concentrates judgement in a team without context.** A central approver cannot know the consumer set of every stream, so approval degrades into rubber-stamping, which is an override with extra latency. ## Making the override expensive rather than impossible The design goal is to keep the action available and make it **loud, attributed and temporary**: 1. **Attribute it.** Who relaxed which name, when, and why, recorded by the registry rather than by convention. 2. **Time-box it.** The relaxed setting expires and reverts by itself. Overrides that must be renewed are noticed; overrides that persist silently are the ones that cause the incident a year later. 3. **Notify the dependants.** The teams reading that stream learn before the incompatible version lands, not during it. This alone converts most overrides back into conversations. 4. **Separate the roles.** The team that owns the name is not always the team shipping the change; requiring the name's owner to agree is a light check that keeps context local. 5. **Measure the rate, not the instances.** A count of active overrides, and their age, is the health metric. A rising count means the default is wrong for some class of streams, which is information, not misbehaviour. ## Three placements of the control | Placement | What you gain | What you pay | |---|---|---| | Central mandate, no override | A statable guarantee | Bypasses you cannot see; blocked legitimate work | | Free per-team override | Speed, local judgement | No global guarantee; externalised cost; no evidence | | Attributed, expiring override | Both, approximately | Instrumentation and the discipline to watch the metric | ## The judgement to voice in an interview Say plainly that this is a control-placement decision, not a tooling one, and that the failure mode of maximum strictness is invisible bypasses rather than compliance. Then commit: default strict, keep the escape hatch, and spend the effort on attribution and expiry rather than on approval workflow. A guarantee nobody can circumvent does not exist in an organisation that can create a new name in a second; a guarantee whose exceptions are visible and short-lived does.
- Why is a forbidden override often worse than a recorded one?Because the alternative routes remain open and are silent: a fresh name with no history, a parallel stream, or a writer that never registers. A recorded override leaves a dependant list and an expiry; a bypass leaves nothing, and you learn about it from the consumer that broke.
- What single piece of instrumentation most improves this situation?An automatic expiry on a relaxed setting, with attribution. It converts an invisible permanent state into a visible temporary one, forces a renewal that someone must justify, and makes the count of active overrides a metric you can act on.
- When is deliberately breaking a contract the right call?When the consumer set is small, known and coordinated, or the contract was wrong in a way that costs more to preserve than to break. The test is whether you can enumerate the readers and their rollout order; if you cannot, the change is not coordinated, it is unmeasured.
saying these in an interview costs you the question
- Assumes a configured mode is as binding as a technical constraint
- Believes forbidding overrides eliminates breaking changes
- Ignores that the cost lands on consuming teams
- Treats a relaxed setting as temporary without an expiry
- Never tells the dependants of a stream about an override
- Reads a rising override count as misbehaviour rather than signal