How do you evolve the schema of a domain event published to a Kafka topic without breaking existing consumers? Discuss compatibility modes and what changes are safe.
answer
- Schema Registry + Avro/Protobuf/JSON; id embedded in record
- BACKWARD = upgrade consumers first (default)
- FORWARD = upgrade producers first
- FULL = both; TRANSITIVE = check all prior versions
- Safe: add field w/ default, remove defaulted field
- Breaking change → new version/topic + dual-publish
basics
~20 sUse a schema registry (Avro/Protobuf/JSON Schema) and a compatibility mode. The safest common choice is BACKWARD: new consumers can read old data. Safe changes are adding optional fields with defaults and removing optional fields; renaming or removing required fields breaks compatibility.
solid answer
~60 sPut event schemas in a **Schema Registry** (Confluent or compatible) with **Avro**, **Protobuf**, or **JSON Schema**, and let the registry enforce a **compatibility mode** at registration time. Modes: **BACKWARD** (new schema can read data written with the previous schema — lets you upgrade consumers first; the most common default), **FORWARD** (old schema can read data written with the new one — upgrade producers first), **FULL** (both), and the `*_TRANSITIVE` variants that check against *all* prior versions, not just the latest. With Avro, **adding a field with a default** and **removing a field that has a default** are backward-compatible; adding a *required* field without a default, or removing a required field, is not. Renames are dangerous unless you use Avro **aliases**. Practically: never reuse a field for a new meaning, version your event type, prefer additive changes, and for breaking changes publish a new topic or a new event version (e.g. `OrderPlaced` v2) and dual-publish during migration. The registry's schema id is embedded in each record so consumers fetch the exact writer schema.
go deeper
Know that a schema registry plus a compatibility mode protects consumers and that you add optional fields, not required ones.
Explain BACKWARD vs FORWARD and which deploy order each enables.
Apply transitive modes, Avro defaults/aliases, and a dual-publish/versioning plan for breaking changes.
Set org-wide registry policy (default mode, subject strategy, CI enforcement) and treat event schemas as governed public APIs.
## Why event schemas must evolve carefully Events on a Kafka topic are a **contract** between a producer and *many, unknown, independently-deployed* consumers — and because Kafka **retains** records (sometimes forever, e.g. compacted topics), a consumer may read events written **months ago** under an **old** schema. So you cannot just 'change the class'. You need rules that guarantee old and new readers/writers interoperate. ## Schema Registry mechanics A **Schema Registry** stores versioned schemas and assigns each a global **schema id**. Serializers (Avro/Protobuf/JSON Schema) embed that id (a magic byte + 4-byte id) in **every record's** bytes. A consumer reads the id, fetches the **writer's schema**, and deserializes — Avro then resolves the writer schema against the consumer's **reader schema** using its resolution rules. The registry **rejects** a new schema at registration time if it violates the configured **compatibility mode**, turning a runtime break into a deploy-time failure. ## Compatibility modes (the core interview content) - **BACKWARD**: a consumer using the **new** schema can read data written with the **previous** schema. ⇒ You can **upgrade consumers first**, then producers. Allowed: **delete fields**, **add optional fields (with defaults)**. - **FORWARD**: a consumer using the **old** schema can read data written with the **new** schema. ⇒ **Upgrade producers first**. Allowed: **add fields**, **delete optional fields**. - **FULL**: both backward and forward — only additions/removals of optional (defaulted) fields. - **NONE**: no checks (dangerous). - **\*_TRANSITIVE** variants (`BACKWARD_TRANSITIVE`, etc.): check compatibility against **every previous version**, not just the immediately prior one. Use these when consumers might be reading very old retained data. **BACKWARD (the default in Confluent)** is usually right for event streams because you want to roll out new consumers without coordinating producer deploys. ## Which concrete changes are safe (Avro) - **Safe (backward)**: add a field **with a default**; remove a field that **had a default**. - **Unsafe**: add a **required** field (no default); remove a **required** field; change a field's **type** incompatibly; **rename** a field (unless you add an Avro **alias**). - **Enums**: adding symbols can break old readers unless a default symbol is set. - **Protobuf** is more lenient (field numbers, not names, define identity; new fields are optional by design) — but **never reuse a field number** for a new meaning. ## Strategy for genuinely breaking changes When a change can't be made compatible: 1. **New event version / type**: introduce `OrderPlaced` **v2** (new subject or a versioned type) and **dual-publish** v1 and v2 during migration; retire v1 once all consumers move. 2. **New topic**: route v2 to a new topic and migrate consumers over. 3. Use the registry's **subject naming strategy** (TopicNameStrategy vs RecordNameStrategy/TopicRecordNameStrategy) deliberately if a topic carries multiple event types. ## Golden rules / pitfalls - **Always add defaults** to new fields. - **Never repurpose** an existing field's semantics — add a new one. - **Never reuse** Protobuf field numbers / never change Avro field meaning. - Pick **transitive** modes if old retained data must remain readable. - Treat the schema as a **published API**: review changes, keep them additive, and version explicitly for breaks. - Test evolution in CI by registering candidate schemas against the registry before deploy.
- You must add a brand-new required field with no sensible default. How do you ship it without breaking consumers?It's a breaking change under BACKWARD/FULL. Introduce a new event version (e.g. OrderPlaced v2) or a new topic, dual-publish v1 and v2 during the migration window, move consumers to v2, then retire v1. Don't force the required field onto the existing schema.
- What's the difference between BACKWARD and BACKWARD_TRANSITIVE, and when does it matter?BACKWARD checks the new schema only against the immediately previous version; BACKWARD_TRANSITIVE checks it against all prior versions. Transitive matters when consumers may still read very old retained records (e.g. long-retention or compacted topics), so the new reader must understand every historical writer schema.
saying these in an interview costs you the question
- Thinking you can rename or repurpose a field freely
- Adding a required field with no default and assuming it's safe
- Confusing BACKWARD (upgrade consumers first) with FORWARD
- Reusing a Protobuf field number for a new meaning
- Believing retained old records don't need to stay deserializable