skip to content

When evolving the schema of an event type that's already been written to an append-only event store, which kinds of changes are generally safe to make without breaking existing consumers or replay, and which are dangerous?

level: middleimportance: must knowfreq 65%

answer

  1. additive + optional = safe
  2. rename/retype/split = needs upcaster
  3. backward vs forward compatibility are different guarantees
  4. unknown-field tolerance in the serializer matters
  5. rolling deploys expose forward-compat gaps

basics

~20 s

Adding a new optional field is usually safe because old readers can ignore it and new readers can default it. Renaming, removing, retyping, or changing the meaning of an existing field is dangerous because old events won't have the new shape and something will misread them.

solid answer

~40 s

Safe changes are additive and optional: adding a new field with a sensible default for events that predate it, adding a new event type, or widening a type in a way old readers still parse. These preserve backward compatibility (new code can read old events) and forward compatibility (old code can still read new events by ignoring unrecognized fields), as long as the serialization format tolerates unknown fields. Dangerous changes are anything that changes the meaning or shape of an existing field: renaming, changing its type, removing a required field, splitting one field into two, or changing units/semantics without a version bump. Those require an explicit versioned upcaster or full migration, because old data can't satisfy the new contract on its own.

go deeper

for a junior

Should be able to name adding an optional field as safe and renaming/removing a field as risky, even without the backward/forward compatibility vocabulary.

for a middle

Should correctly use backward vs forward compatibility terms and give a concrete example of each category, such as add-with-default versus rename/retype/split.

for a senior

Should connect the rules to rolling-deployment realities and serialization-format specifics like Avro-style default requirements, and explain why silent corruption is worse than a loud failure.

for a principal

Should be able to set organization-wide policy: when additive-only discipline is enough, when to mandate an explicit versioning/upcasting process, and how to handle the deprecation window for fields to eventually drop.

## Two audiences read the same log An event-sourced system has two audiences reading the same log over time: the current version of your own code replaying history, and potentially other services or long-running consumers that may run an older or newer schema simultaneously during a rolling deploy. Schema evolution rules exist to let both audiences keep working without a synchronized stop-the-world upgrade. The two properties being protected are: | Property | What it guarantees | |---|---| | **backward compatibility** | new code can still correctly read old events | | **forward compatibility** | old code doesn't blow up when it encounters an event written by newer code, for example during a canary rollout where two consumer versions run side by side | A change is 'safe' precisely when it doesn't put either property at risk. ## The safe category The safe category is **additive-and-optional**. - **Adding a brand-new field to an event**, with a defined default for events written before it existed, is the textbook safe change: old events simply don't have it, and any reader that understands the new schema fills in the default; readers still on the old schema ignore it if the serialization format (JSON, Avro, Protobuf) tolerates unknown fields. - **Adding an entirely new event type** is similarly safe, since it's opt-in. - **Widening a numeric type** (say a 32-bit count to a 64-bit count) is usually safe if the serialization and runtime handle the widened range transparently. The common thread is that nothing that already existed changes meaning; you're only ever adding. ## The dangerous category The dangerous category is anything that changes the meaning, name, type, or cardinality of a field that already has historical data written under the old contract. - **Renaming** `amount` to `total` breaks every reader still looking for `amount`, because the bytes on disk still say `amount` — a classic mistake that looks harmless in code review since it 'obviously' still means the same thing, but breaks any consumer, including your own replay logic, that keys on the field name. - **Changing a field's type** (a string ID becoming a structured object, a float becoming a fixed-point decimal) is dangerous because old serialized values won't parse under the new type without an explicit conversion. - **Removing a field** some reader still depends on, or **making an optional field required**, breaks things for the same reason in reverse. - **Splitting one field into two** (like `total` into `subtotal` + `tax`) is dangerous because a naive deserializer has no way to know how to divide the old value — it needs an explicit rule, i.e. an upcaster. ## The trade-off with additive-only discipline The trade-off with sticking rigidly to 'only additive changes' is that it doesn't cover every real evolution need — sometimes the business genuinely needs to split a field or fix a wrong unit, and there's no additive way to do that. Staying purely additive keeps you safe with zero extra machinery, but you accumulate cruft (deprecated-but-still-present fields, awkward parallel fields) because you can never truly remove or rename anything without a breaking version bump. Doing a breaking change properly (bump version, write an upcaster, deprecate the old field over a compatibility window) costs effort and discipline but keeps the schema clean long-term. Teams that skip that discipline and just rename fields in place accumulate silent data corruption that isn't caught until someone notices a report is wrong months later. ## Failure modes - **In production, the most common failure mode is a 'safe-looking' rename or retype that ships without a version bump;** old events already in the store don't change, so nothing crashes on deploy — instead, every replay of pre-change events either throws a deserialization error (loud, best case) or silently defaults to null/zero for the renamed field (quiet, worst case, compounding through every downstream projection). - **Another common failure is rolling deploys:** if service A starts writing the new event shape before an independent consumer B has deployed code that understands it, B either crashes on the unexpected shape or silently drops data it can't parse — exactly the forward-compatibility risk additive-only changes are designed to avoid. A concrete real-world illustration is Avro's schema evolution rules, used heavily with Kafka: - adding a field is safe **only if it has a default**; - removing a field is safe **only if it had a default in the old schema**; - renaming requires an **explicit alias recorded in the schema itself**, precisely because 'just rename it' silently breaks both old readers and old data with no error at all.

  • Why is a rolling deployment specifically relevant to forward compatibility, even if the event schema change itself is additive?
    During a rolling deploy, two versions of the same consumer can run at once against the same stream. If the new writer starts emitting events with a new required field before every consumer has upgraded, still-running old consumers may fail to parse or silently drop that field, which is a forward-compatibility failure — old code encountering new data.
  • How does Avro's handling of field removal differ from just deleting a field in your event class?
    Avro treats a field removal as safe only if that field had a default value in the writer's schema, because a reader on the newer schema without that field can then apply the default when reading an old event that still has it. Just deleting a field from your language-level class with no schema-level default breaks both directions.
  • Is changing an event field's unit, for example an amount from dollars to cents, an additive change?
    No — even though the field's name and type stay the same, its meaning changes, so old stored values would be silently misinterpreted by 100x under the new semantics. This needs a version bump and an explicit upcaster that multiplies old values, not a silent in-place reinterpretation.

Like adding a new optional line to a printed form — anyone can leave it blank and old form-readers don't even notice it — versus renaming an existing line on the form, after which every filing cabinet full of old forms is now mislabeled and nobody filing new ones knows which drawer to use.

saying these in an interview costs you the question

  • Calls a field rename 'safe' because the meaning is obviously the same
  • Doesn't distinguish backward compatibility from forward compatibility
  • Assumes JSON's flexibility alone guarantees compatibility with no default-value discipline
  • Ignores rolling-deploy windows where old and new consumers run simultaneously
  • Thinks changing a field's unit or scale is a no-op if the type stays the same

context