What's the difference between a 'weak schema' approach and a 'strong schema registry' approach to versioning events within an event-sourced store, and what does each cost you?
answer
- weak = untyped JSON, no gate, discipline-based
- strong = registered schema per version + enforced compatibility mode
- cost paid early (strong) vs late (weak)
- scales with number of independent teams/consumers
- registry can become a bottleneck or gamed with a blob field
basics
~20 sA weak schema just serializes events as loosely-typed JSON with a version number and trusts application code to handle whatever shape shows up. A strong schema registry enforces a formal, checked schema for every event version up front, catching incompatible changes before they can even be written.
solid answer
~40 sWeak-schema event stores treat each event as effectively untyped data, JSON with a `type` and `version` field, with no enforcement at write time — compatibility is entirely a discipline and testing problem for the teams writing upcasters. This is flexible with near-zero tooling overhead, but nothing stops a developer from writing an incompatible event that breaks readers until discovered in production. A strong schema registry, Avro or Protobuf style, often paired with something like Confluent Schema Registry, requires every event schema to be registered and enforces compatibility rules at write time — an incompatible schema change is rejected before it can ever be published. That safety costs process overhead and some rigidity. The choice is essentially: pay the cost upfront and continuously versus pay it unpredictably later when a bad change reaches production.
go deeper
Should know there are two broad approaches — trust-based JSON versioning versus a formal enforced registry — without needing to name specific tools or compatibility modes.
Should be able to describe what a schema registry actually checks, the compatibility mode, and give one concrete cost of each approach.
Should reason about when each approach fits an organization's scale, single team versus many independent consumers, and describe realistic failure modes for both, including registry-as-bottleneck and the blob-field escape hatch.
Should be able to design or choose a schema governance strategy across an organization, including compatibility-mode policy, registry operability and HA, and a plan for legitimate breaking changes that doesn't just disable enforcement.
## The structural question Every event-sourced system has to answer a structural question before it writes a single event: how strictly is the shape of an event enforced, and by what mechanism? 'Weak schema' and 'strong schema registry' are the two ends of that spectrum, and the choice shapes how schema evolution mistakes get caught — **early and cheaply, or late and expensively.** ## How each one actually works - **In a weak-schema system**, events are typically serialized as JSON carrying a `type` name and a `version` number, but no external authority validates that a given payload actually conforms to any declared shape. The application code that writes the event decides what fields to include; the application code that reads it decides what to expect, and if those disagree — a field is misspelled, a type doesn't match, a required field is missing — nothing catches it until deserialization fails or, worse, silently produces wrong values at read time, potentially in production, potentially months after the bad event was written. Compatibility is entirely a matter of team discipline and testing; there is no gate. - **In a strong-schema-registry system**, the pattern popularized by Avro and Protobuf usage with Kafka via tools like Confluent Schema Registry, every event type has a formally registered schema per version, and the registry enforces a declared compatibility mode — backward, forward, or full — at the moment a writer tries to publish a new schema version. If a proposed schema violates the compatibility rule, say removing a field that lacked a default, the registry rejects the registration outright, before any event using that broken schema can ever be produced. ## Why both patterns exist The reason both patterns exist, rather than one clearly dominating, is that they optimize for different failure timing and different organizational scales. 1. **Weak schema exists because most systems start small**, with one team owning both writers and readers, where the overhead of a registry and a registration workflow is pure friction relative to the risk — you can just be careful, write good upcasters, and test them, because a mistake's blast radius is contained to your own team. 2. **Strong schema registries exist because at larger scale**, with many independent teams reading and writing the same event types on different deploy schedules, informal discipline doesn't scale: nobody can review every other team's event changes for compatibility by hand, and a single incompatible change from one team can silently break a consumer owned by a completely different team. ## The trade-off The trade-off is symmetric and real on both sides. - **Weak schema costs nothing in tooling or process** and stays maximally flexible, but the cost is deferred and often more expensive: incompatible changes are caught by production failures, not a pre-write gate, and by the time you notice a consumer has been silently defaulting a missing field for two weeks, you may need a data-repair job on top of the schema fix. - **Strong schema pays the cost upfront and continuously:** someone has to run and operate the registry, every schema change goes through a registration step, and the enforced compatibility mode can feel rigid when a genuinely necessary breaking change now needs a conscious version bump and coordinated consumer upgrades rather than a quiet rename. ## Failure modes In production, weak-schema failure modes look like: - a typo'd field name shipped without anyone noticing until a downstream projection silently stops updating; - two teams independently 'fixing' the same event type in incompatible ways because there's no single source of truth for the current schema; - and upcasters written reactively, after a break is discovered, rather than proactively. Strong-schema-registry failure modes look different: - a team blocked mid-incident because a legitimately necessary breaking change is rejected by the compatibility mode with no fast path; - the registry itself becoming a bottleneck or single point of failure if not made highly available; - and teams gaming the system by registering overly permissive schemas, like a top-level catch-all blob field, specifically to bypass enforcement, which reintroduces weak-schema risk inside a nominally strong-schema system. A well-known real-world reference is Confluent Schema Registry's compatibility modes, `BACKWARD`, `FORWARD`, `FULL` and their transitive variants, applied to Avro schemas on Kafka topics — the most widely deployed version of exactly this pattern, though the same registry-and-enforce idea shows up in any platform wanting publish-time safety instead of runtime discovery.
- What does a schema registry's 'backward compatibility' enforcement mode specifically check before allowing a new schema version to be registered?It checks that a reader using the new schema can still correctly read data written under the previous schema version, for example that any field being removed had a default, or that a newly required field also has one. If the new schema would fail to parse old data, registration is rejected.
- How does a schema registry avoid becoming a single point of failure for every event write in the system?In practice it's deployed as a highly-available clustered service, and clients cache schema lookups aggressively, often permanently since a given schema ID never changes once registered, so steady-state writes don't need a live round-trip to the registry — only genuinely new schema versions require a registration call.
- Can a team 'cheat' a strong schema registry's enforcement, and what's the risk of doing so?Yes — a common way is registering a schema with a permissive catch-all field, like an untyped map or JSON blob, that technically satisfies the registry's compatibility rules while letting the team put anything inside it. The risk is that this reintroduces all the weak-schema failure modes inside a system that appears to have strong guarantees, which is often worse because nobody expects to need the same discipline there.
Weak schema is like a shared team notebook where anyone can write whatever they want and you just trust everyone to be careful; a strong schema registry is like a form-processing office that rejects any submitted form on the spot if it doesn't match an approved template — safer, but now every new form design has to go through approval.
saying these in an interview costs you the question
- Says a schema registry validates business logic, not just structural compatibility
- Thinks weak schema means 'no versioning at all' rather than versioning without enforcement
- Doesn't mention that compatibility checks happen at publish/registration time, before the event is written
- Assumes a strong schema registry has no operational cost or single-point-of-failure risk
- Can't explain a scenario where weak schema is the reasonable choice