A platform team is debating whether to mandate FULL compatibility mode for every schema subject across the organization. What are the concrete costs of that policy, and in what situations would a looser mode, or no registry at all, be the better call?
answer
- FULL = additive-only, forever, both directions
- cost scales with number of independent consumers and unpredictable deploy order
- single producer/consumer in lockstep pipeline needs less strictness
- transient non-replayed topics don't need backward guarantee
- registry itself has a fixed operational cost, not free below some team size
basics
~20 sFULL mode is the safest but the most restrictive: it forces every schema change to be additive with defaults, which slows down teams that need to rename fields, tighten types, or make big changes. For internal, tightly-coordinated systems or truly transient data, a looser policy or skipping the registry entirely can be the more pragmatic choice.
solid answer
~1 minMandating FULL compatibility everywhere trades organizational safety for evolution speed: because FULL requires both new-reads-old and old-reads-new to hold, essentially every field must be optional with a sensible default forever, which rules out renames, tightening a field from optional to required, or narrowing a type, without a new subject (a de facto major version). That's the right call for shared, multi-consumer platform topics with unpredictable deploy order, but it's overkill, and a genuine productivity tax, for topics with a single producer and single consumer deployed in lockstep (e.g., two services in the same deploy pipeline, or an internal topic entirely owned by one team) where BACKWARD or even NONE with strong integration tests is perfectly safe and much faster to iterate on. It's also often unnecessary for topics carrying short-lived, non-replayed data where old messages are never read by new code anyway, since one of FULL's two guarantees is simply moot. A registry (in any mode) can also be entirely the wrong tool for very small systems or prototypes where the coordination overhead of registering and versioning schemas outweighs the risk it protects against, at which point plain JSON with disciplined manual review is a reasonable, if less enforced, alternative.
go deeper
Should sense that stricter modes make some changes (like renames) harder, without needing to design a full policy.
Should explain concretely why FULL rules out renames/required-field-tightening and name at least one situation where a looser mode is fine.
Should reason about deploy-order predictability, consumer count, and data replay as the variables that decide the right policy per subject, plus multi-step migration patterns.
Should design and justify an org-wide schema governance policy differentiating shared platform topics from internal/team-owned ones, anticipate the friction/escape-hatch failure mode, and know when skipping a registry entirely is the pragmatic call.
## The policy dial Compatibility mode is fundamentally a **policy dial** that trades safety against how much schema change is still expressible. FULL compatibility, requiring simultaneous BACKWARD (new reader can parse old data) and FORWARD (old reader can parse new data), is the strongest guarantee a registry can enforce short of forbidding change altogether, but that strength has a real, recurring cost: nearly every evolution has to be purely additive, with every new field carrying a default and every removed field having carried one, forever. - **Renaming a field** is impossible under FULL without going through an awkward add-new-field-then-deprecate-old-field dance stretched over multiple releases, because a rename is structurally equivalent to removing one field and adding an unrelated one, which almost always violates one direction or the other. - **Tightening an optional field to required**, or narrowing a numeric type from long to int, similarly tends to violate FULL because it changes what a reader on one side can assume that a writer on the other side actually provides. ## The case for mandating it everywhere The organizational case for mandating FULL everywhere is understandable: it removes the need for anyone to reason about deploy order, which is genuinely hard to guarantee across dozens of independently-released services, especially ones owned by different teams with different release cadences, or in blue-green/canary deployments where old and new producer code can be running simultaneously for minutes or hours. For a shared platform topic like a company-wide `user-events` stream consumed by a dozen teams with no coordination mechanism between them, FULL is close to mandatory; anything looser bets on discipline that a large, decoupled organization structurally cannot enforce. ## Where the mandate stops paying But applying that same policy uniformly to every subject in the registry, including internal, single-producer/single-consumer topics, imposes the cost without buying the corresponding benefit. Consider two microservices owned by the same team, deployed via the same CI/CD pipeline in a fixed order (producer migration runs, then producer deploys, then consumer deploys, always in that sequence, enforced by the pipeline itself). Here the deploy order is not unpredictable, it's guaranteed by tooling, so BACKWARD compatibility (or in some cases even relaxed checking backed by strong contract or integration tests) delivers equivalent real-world safety with far less friction: the team can rename a field or make a previously-optional field required in a single release, because they control both ends and know exactly when each side ships. Forcing FULL here means the team pays the **additive-only tax** for a risk (arbitrary deploy order) that doesn't actually exist in their pipeline. ## Transient, never-replayed data Another case where a looser policy is correct is topics carrying transient, non-replayed data, such as a short-retention topic used purely for real-time fan-out where messages are consumed within seconds and never replayed from history (retention measured in minutes, not the days or weeks typical of an event-sourced or audit-style topic). Because old messages are never read by a newly-deployed consumer in that world, the BACKWARD half of FULL's guarantee, new reader parses old data, is protecting against a scenario that structurally cannot happen. Enforcing it anyway is pure overhead. ## Whether to run a registry at all There's also a smaller-scale, more foundational question: whether to run a schema registry at all. For a small system, an early-stage product, or a handful of services in active flux where the schema itself is still being discovered, the fixed cost of standing up and operating a registry can outweigh the risk it mitigates, especially when there are only one or two consumers who can simply coordinate via code review and a shared repo. That fixed cost is: - an additional highly-available service, - subject naming conventions, - CI wiring for schema registration, - onboarding every new engineer to the registry workflow. In that phase, plain JSON messages with a lightweight JSON Schema validated in tests, or even just disciplined manual review of message shape changes in pull requests, is a legitimate, lower-ceremony alternative; teams often introduce a full registry later, once the number of independent consumers and the pain of uncoordinated schema drift both cross a threshold that justifies the operational investment. ## Failure modes on each side - The failure mode on the **'too strict everywhere'** side is organizational friction that erodes trust in the registry itself: engineers who repeatedly hit FULL-mode rejections for changes they know are actually safe in their specific deploy topology start reaching for escape hatches, like registering a brand-new subject for every minor change to dodge the compatibility check entirely, which defeats the purpose of having compatibility guarantees at all and litters the registry with orphaned, unmaintained subject histories. - The failure mode on the **'too loose or absent'** side is the classic uncoordinated-breaking-change incident: a producer ships a change with no registry (or NONE mode) enforcing anything, a consumer several deploys behind starts throwing deserialization errors in production, and the team discovers the dependency only from an incident, not from a build failure. ## Where it shows up A concrete real-world shape: a mid-size company runs FULL_TRANSITIVE on all shared platform-team topics (user events, billing events) consumed by many downstream squads, but explicitly carves out BACKWARD-only (or even registry-optional, contract-test-covered) policies for topics internal to a single squad's own service mesh, documented explicitly in their platform's schema governance guide so engineers aren't surprised by which rule applies where, and don't waste time lobbying to loosen a shared-topic policy that exists precisely because of the coordination problem across teams they don't control.
- Why does renaming a field almost always violate FULL compatibility even though it feels like a minor, low-risk change?A rename is structurally indistinguishable from removing one field and adding a differently-named one, and adding a field under FULL requires it to have a default so an old reader parsing new data still works, while the removed field breaks a reader that still expects it under the old name. The registry has no semantic understanding that the old and new names refer to 'the same' logical field, so it evaluates the change as two separate, individually risky structural edits.
- What's a safe multi-step pattern for renaming a field on a subject that's locked to FULL compatibility?Add the new field name alongside the old one (dual-write, both populated by the producer, both optional with defaults), wait until all consumers have migrated to reading the new field, then remove the old field name in a later, separate schema version once it's confirmed no live consumer still depends on it. Each individual step stays additive-with-defaults and passes FULL checks on its own.
- How would you decide, as a platform team, which topics get FULL_TRANSITIVE versus a looser policy?The key variables are number of independent consumer teams, predictability of deploy order between producer and consumers, and whether historical messages are ever replayed by new consumer code; topics with many independently-deployed consumers and replayed history warrant FULL_TRANSITIVE, while topics with a single coordinated producer/consumer pair or purely transient non-replayed data can safely use a looser policy, ideally documented as an explicit, discoverable governance rule rather than left to each team's guess.
Mandating FULL compatibility everywhere is like requiring every door in a building, including a private supply closet only one person ever uses, to meet public-fire-exit code: the right call for the main lobby doors everyone passes through unpredictably, but needless cost for a closet whose one user already knows exactly how it works.
saying these in an interview costs you the question
- Claims FULL compatibility should always be used everywhere with no situational judgment
- Cannot explain why a field rename is hard under FULL mode
- Doesn't recognize that a registry itself has an operational cost that can outweigh its benefit for very small systems
- Assumes compatibility mode choice is purely technical with no organizational/coordination dimension
- Suggests routing around FULL-mode rejections by registering a brand-new subject for every small change as a normal practice