An aggregate's apply function was written when its ProductAdded event had fields (sku, quantity). Two years later the event is redesigned to (sku, quantity, warehouseId), and old events already in the stream only have the two original fields. What has to happen for replay to still correctly rehydrate aggregates whose streams contain the old-shaped events, and why can't you just edit the old events in place?
answer
- events are immutable, never rewritten
- upcaster: old shape -> new shape, read-time only
- chain of upcasters across versions
- apply only ever sees current shape
- forgetting an upcaster = replay crash on old data
basics
~20 sYou can't rewrite old events - they're an immutable historical record - so instead you add a translation step, often called upcasting, that transforms an old-shaped event into the new shape (e.g., filling in a default warehouseId) right before it's folded, so the apply function only ever has to deal with one, current shape.
solid answer
~50 sBecause events are immutable, append-only facts, you never rewrite history to fix a schema change - doing so would falsify the audit trail and risk breaking any other consumer that already processed the old shape. Instead, you introduce an upcasting layer between 'read raw event from the store' and 'pass it to apply': for each old event version, a small transform function maps it to the current (or next) version's shape - e.g., ProductAddedV1(sku, quantity) becomes ProductAddedV2(sku, quantity, warehouseId: DEFAULT_WAREHOUSE) via a documented, defensible default. Upcasters are chained if there have been multiple schema versions, so a V1 event might go V1 to V2 to V3 before reaching the current apply function, which then only ever needs to know about the latest shape. This keeps the apply function simple and lets replay work correctly across the entire history, including the oldest events, without touching stored data.
go deeper
Should understand, at a high level, that old events can't be edited and that something has to translate them before they're used.
Should be able to describe the upcaster concept - a transform from old shape to new shape applied at read time - even if not fluent in chaining multiple versions.
Should be able to design a chained-upcaster strategy for multiple schema revisions and reason about default-value risk and the trade-off against a one-time migration.
Should reason about this as a long-term system-design and organizational-process question: versioning conventions, how upcaster debt is tracked and tested over years, and when a migration project is worth the compliance risk versus indefinite upcaster accumulation.
## Why you cannot rewrite history Event sourcing's core promise is that the event log is the permanent, immutable source of truth, which creates a real tension the moment a team needs to change an event's shape: you cannot simply go back and edit `ProductAdded` events already sitting in the store to add a `warehouseId` field — - because that would mean silently rewriting history, falsifying what was actually recorded at the time, potentially in violation of audit requirements that were part of the reason to choose event sourcing in the first place; - and because other systems may already have consumed and acted on the old-shaped events, so a retroactive edit wouldn't even be consistently visible everywhere it matters. So the stream stays exactly as it was written: some `ProductAdded` events have only (sku, quantity), and any events appended after the schema change have (sku, quantity, warehouseId). Replay still needs to fold all of them into one consistent `apply` function without that function having to special-case every historical version it might ever encounter. ## The upcasting layer The standard solution is an **upcasting layer** sitting between "read the raw stored event" and "call apply." An upcaster is a small, pure transform function tied to a specific event type and version: given an old-shaped event, it returns the next version's shape, filling in any new fields with a sensible, documented default or a value derived from other data available at that point (here, perhaps a `DEFAULT_WAREHOUSE` constant representing "the only warehouse that existed before multi-warehouse support shipped"). If there have been several schema revisions over the aggregate's lifetime, upcasters are **chained**: 1. A V1 event passes through the V1-to-V2 upcaster. 2. Its output passes through the V2-to-V3 upcaster. 3. And so on, until it reaches the shape the current `apply` function actually expects. This means apply itself never needs to know history exists — it only ever sees the latest, current shape of every event type — which keeps the state-transition logic simple and lets it evolve independently of how many schema versions preceded it. ## Why schemas change at all This exists because software evolves and event schemas are no exception: - business requirements change (multi-warehouse support gets added); - bugs in the original event design get discovered (a field was missing that turned out to matter); - or naming/structure just needs cleanup as understanding improves. A system that can't tolerate any event schema change without a painful migration would make event sourcing far less attractive for long-lived aggregates, which is exactly where event sourcing's audit-trail benefits matter most. Upcasting decouples "how data is stored historically" from "what shape current code wants to work with," the same way a database migration framework decouples "how a table was created three years ago" from "what the ORM model expects now" — except that instead of an `ALTER TABLE` that changes stored data in place, upcasting is applied at read time, every time, leaving storage untouched. ## The trade-off The trade-off is ongoing maintenance cost: every schema change adds one more upcaster to the chain, forever, because you can never delete an upcaster as long as any event of that old version might still exist in any stream that could be replayed — which, for compliance-heavy systems, is effectively forever. Over years, a heavily-evolved event type can accumulate a long chain of upcasters: - each one needing its own tests and its own documented rationale for whatever default or derivation it applies; - and each one being genuinely painful to reason about years later when the person who wrote it has moved on and the "why DEFAULT_WAREHOUSE" context is only in a code comment (if it's anywhere at all). Some teams manage this by periodically taking a snapshot-and-migrate approach for aggregates whose full history is no longer operationally needed — but that has to be weighed carefully against exactly the audit/compliance requirements that motivated using event sourcing. ## Failure modes Failure modes here are common and expensive because they surface as broad, replay-triggered outages rather than isolated bugs. - **The most frequent** is forgetting to add an upcaster at all when an event shape changes, so any code path that replays old data — a projection rebuild, a debugging tool, a rarely-used aggregate that hasn't been touched in months — starts throwing deserialization errors the moment it hits a pre-change event, often discovered only in production, long after the schema change shipped and passed all tests (which only exercised new-shaped events). - **A second** is an upcaster that fills in a default that's subtly wrong for some historical events — e.g., assuming `DEFAULT_WAREHOUSE` is correct for all old events when a handful of them were actually for a since-decommissioned second warehouse that predates the "multi-warehouse support" feature the team remembers adding — producing state that's plausible-looking but factually incorrect, which is worse than an outright crash because nobody notices. - **A third, architecturally deeper problem** is chaining so many upcasters over the years that the read path for old aggregates becomes measurably slower and harder to reason about than for new ones, which is itself a signal that the team should consider a deliberate migration project rather than open-ended upcaster accumulation. Frameworks like **Axon Framework** and **NEventStore** ship explicit upcasting/event-versioning support for exactly this reason — it's considered a first-class, expected part of running event sourcing in production for more than a year or two, not an edge case.
- Why can't you just make the apply function itself handle both old and new event shapes with an if/else on which fields are present?You can for one schema change, but it doesn't scale - every subsequent revision adds another branch, and apply ends up permanently coupled to the full history of every shape the event type has ever had, defeating the goal of keeping current business logic simple. Upcasting isolates that version-translation concern outside apply so business logic only ever deals with the latest shape.
- What happens if an upcaster's chosen default value turns out to be factually wrong for some subset of historical events?Replay produces state that looks valid but is quietly incorrect for those aggregates, which is often worse than a crash because nothing signals the problem - it can go unnoticed until a downstream report or audit surfaces an inconsistency. This is why upcaster defaults need explicit review and documentation, and ideally a data audit to check the assumption before shipping.
- Is it ever acceptable to actually mutate old stored events instead of upcasting?In systems with strict compliance/audit requirements, essentially no - it would falsify the historical record and could violate the same regulations that motivated using event sourcing. Some teams do run one-time, carefully audited migration projects for non-regulated data to reduce the upcaster chain, but that's a deliberate, tracked exception, not a routine practice, and it should preserve an archived original where retention rules require it.
Like a translator standing between an old foreign-language letter and a modern reader who only speaks the current language: the original letter is never altered, but it's translated on the fly every time someone needs to read it, and if the letter went through several past translations, each hop happens in sequence before it reaches the reader.
saying these in an interview costs you the question
- Suggests editing/rewriting old stored events in place to fix the schema
- Doesn't distinguish between 'stored shape' and 'shape the apply function expects'
- Has apply directly branch on multiple historical event shapes indefinitely instead of using an upcasting layer
- Assumes schema changes are rare enough not to need a strategy for them
- Doesn't consider that upcaster defaults could be wrong for some historical data