How would you evolve the payload schemas of Debezium outbox events without breaking downstream consumers?
answer
- the outbox decouples but enforces nothing
- add optional fields, never repurpose one
- a breaking change becomes a new version
- publish both during the migration window
- one topic, several event types, one subject problem
basics
~20 sTreat the outbox payload as a published contract: additive, optional-only changes by default; an explicit event version for breaking ones, published alongside the old until consumers migrate; and a subject naming strategy that lets several event schemas share one routed topic.
solid answer
~50 sThe outbox buys you a decoupled contract — consumers see an event the service authored, not a table — but nothing enforces it. So put the discipline in yourself. Default to additive change: new fields optional with defaults, never remove or retype. For a genuinely breaking change, version the event rather than mutating it — carry the version in the `type` column (`OrderPlaced.v2`) or in a header, publish both versions from the same transaction for a deprecation window, then retire v1 once telemetry shows nobody reads it. A registry subtlety matters here: routing by aggregate type puts several event types on one topic, so the default topic-based subject strategy forces incompatible schemas into one subject. Use a record-name-based subject strategy so each event type versions independently. Govern it like an API: schema in the producer's repo, compatibility checked in their CI, an owner, and a published deprecation policy.
code
sql · 4 lines-- dual-publish inside the same local transaction as the business change
INSERT INTO outbox (id, aggregatetype, aggregateid, type, payload)
VALUES (gen_random_uuid(), 'Order', :orderId, 'OrderPlaced.v1', :legacyBody),
(gen_random_uuid(), 'Order', :orderId, 'OrderPlaced.v2', :newBody);go deeper
Understand that the event body a service writes into its outbox is read by other teams, so adding an optional field is safe while removing or renaming one is not.
Explain how a new event version can be carried in the event type column or a header, and why both versions are written in the same transaction during a migration window.
Be ready to run the migration end to end: dual publish, measure who still consumes the old version, retire it on a published schedule, and get the registry subject strategy right for a topic carrying several event types.
Own the policy. Decide whether an incompatible change stops the pipeline or flows into quarantine, put the compatibility gate in each producing team's build, and keep the platform team out of the per-change approval path.
## What the outbox does and does not give you Routing domain events out of an outbox table rather than publishing raw row images removes the accidental coupling: an `ALTER TABLE` on a business table is invisible to consumers, and the published shape is something a person designed. That is real. What it does not give you is any enforcement. The router ships whatever string the service wrote. A single deploy that changes how the service serialises its payload can break every consumer, and nothing between the service and the consumer will notice. The outbox moves the contract from accidental to deliberate; keeping it *safe* is an organisational job. ## Rule one: additive by default The cheapest evolution is the one consumers never notice. Add fields, make them optional, give them defaults, and never repurpose the meaning of an existing field. A consumer built against last quarter's schema keeps working; a consumer built today reads the new field. Ninety percent of real evolution fits inside this rule, and a team that internalises it rarely has to run the expensive machinery below. What breaks: removing a field, narrowing a type, renaming, changing units or semantics under a stable name. That last one is the nastiest because no automated check catches it — a `total` that quietly starts excluding tax passes every compatibility rule and corrupts every downstream number. ## Rule two: version the event, do not mutate it When a change genuinely cannot be additive, publish a new event version rather than changing the existing one. The natural carrier is the event type column the router already reads: `OrderPlaced` becomes `OrderPlaced.v2`, or a dedicated version column is projected into a header so consumers can dispatch without deserialising. The migration then has a shape: the service writes both v1 and v2 rows inside the same local transaction, so both remain exactly as consistent with the business change as they were before. Consumers migrate at their own pace. When telemetry — consumer group lag on the old type, or a simple counter — shows nothing reads v1, the producer drops it. That window is the deprecation policy, and it should be written down in advance rather than negotiated per incident. ## Rule three: get the registry subject right Here is the piece that catches teams out. Routing by aggregate type means one topic carries every event *type* for that aggregate — `OrderPlaced`, `OrderCancelled`, `OrderRefunded` all land on the orders topic. Under the default subject naming strategy, which derives the subject from the topic, all three schemas compete for one subject, and the registry's compatibility rule is applied *across unrelated event types*. Each new event type looks like a wildly incompatible change to the previous one. The fix is a record-name-based subject strategy, so the subject follows the schema's own name and each event type versions independently on the shared topic. Deciding this at design time is nearly free; discovering it after six months of registered garbage is not. ## Rule four: enforce at the producer The schema belongs in the producing service's repository, next to the code that writes the outbox row, and its compatibility should be checked in that service's CI against the currently registered version. That way an incompatible change fails a build in the team that made it, seconds after they made it — rather than failing a connector task at 3am in a team that did not. Pair it with ownership: a named owner per event type, a catalogue of which events exist, and a subscriber list so a producer can see who they would break. The technical mechanisms above are worthless without knowing who to talk to. ## Rule five: decide what a violation should do This is the judgment call, and it has no single answer. Strict enforcement means an incompatible change stops the pipeline: no bad data, but an outage, and on a source with finite log retention a long enough outage forces a re-snapshot. Permissive handling means the pipeline keeps flowing and the problem surfaces downstream as wrong numbers. For a system of record — payments, ledgers — stopping is usually right. For behavioural telemetry, flowing with quarantine of the bad records is usually right. State the reasoning; an interviewer at this level is testing whether you know the choice exists. ## What this looks like when it works Producers ship additive changes weekly without telling anyone. Breaking changes are rare, announced, dual-published, and retired on a published schedule. The registry holds independently-versioned subjects per event type. The producing team's CI is the gate. And the data platform team is not the bottleneck for any of it — which is the actual goal, because a governance model that requires a central team to approve every change fails at the second team that adopts it.
- Why does routing by aggregate type complicate registry subjects?Because one topic then carries several distinct event types. A topic-derived subject applies one compatibility lineage to all of them, so OrderCancelled looks like an incompatible mutation of OrderPlaced. A record-name-based subject strategy gives each event type its own lineage, which is what you actually want on a fan-in topic.
- How do you know when it is safe to retire the old event version?Evidence, not calendar alone. Watch consumer group lag and offsets on the deprecated type, keep a counter of deliveries actually processed, and publish the retirement date up front. Retire when the telemetry has shown zero consumption across a full business cycle — month-end jobs are the classic consumer nobody remembers.
- Where should the compatibility check live, and why not in the pipeline?In the producing service's CI, against the currently registered schema. A check in the pipeline fails after the change has shipped, in a different team's on-call rotation, with a running outage. A check in the producer's build fails before merge, in the team that can fix it, at zero cost.
- What breaks that no compatibility rule can catch?Semantic change under a stable name — a total that starts excluding tax, a status value whose meaning shifts, a timestamp that changes timezone convention. Every structural check passes and every downstream number is wrong. Only human review, documented field semantics and downstream assertions on value distributions catch this class.
The outbox payload is a published API. You add optional parameters freely, you never quietly change what an existing one means, and a breaking change ships as a new version served alongside the old until the callers move.
saying these in an interview costs you the question
- Assumes the outbox pattern alone protects consumers from payload changes
- Changes the meaning of an existing field instead of adding a new one
- Lets several event types share one topic-derived registry subject
- Plans a big-bang cutover with no dual-publish window
- Puts a central data team in the approval path for every producer change