What is the 'copy-and-replace' approach to migrating an event stream, how does it differ from upcasting-on-read, and when would you choose it despite it being more invasive?
answer
- one-time rewrite vs forever-on-read translation
- copy-and-replace can use external context, upcasters can't
- pays down accumulated upcaster-chain cost
- needs dual-write/backfill + verification during cutover
- blue-green-style flip for the event log
basics
~20 sUpcasting translates old events every time they're read, keeping the original bytes forever. Copy-and-replace instead writes a whole new, corrected copy of the event stream once, then switches everyone to read from that new copy — more work up front, but no more translation cost on every future read.
solid answer
~40 sUpcasting-on-read keeps original events untouched and transforms them at deserialization time, on every replay, forever — cheap to introduce but the transformation cost and code never go away, and it's limited to changes computable from the event's own data. Copy-and-replace instead runs a one-time batch job that reads the entire old stream, applies whatever transformation is needed, including ones that need external context like a reference table, and writes a brand-new stream, then cuts consumers over once verified. It's more invasive — needing a migration window or dual-write/backfill strategy, extra storage during transition, and a verification pass — but it pays off when the transformation isn't a pure per-event function, when accumulated upcaster chains have gotten too deep, or when you want to physically restructure the stream.
go deeper
Should understand at a high level that some schema changes need a one-time rewrite rather than an on-the-fly translation, without needing to design the cutover.
Should be able to describe the basic mechanics of copy-and-replace, reading the old stream, transforming, writing a new stream, and contrast it with upcasting's translate-on-every-read approach.
Should reason about cutover risk, including dual-write, verification, and event loss during migration, and be able to justify choosing one approach over the other for a specific transformation.
Should set migration policy across many streams and services, weigh accumulated upcaster-chain cost against migration risk at organizational scale, and design the safe rollback/verification strategy for a copy-and-replace on a business-critical stream.
## Two ends of one spectrum Upcasting-on-read and copy-and-replace are the two dominant strategies for handling event schema evolution, sitting at opposite ends of a spectrum: | Strategy | The end it sits at | |---|---| | **Upcasting-on-read** | 'do nothing to stored data, translate every time you read it' | | **Copy-and-replace** | 'rewrite the stored data once, then read it plainly forever after.' | **Upcasting-on-read** keeps every historical event byte-for-byte as originally written and inserts a chain of pure transformation functions between storage and the application whenever an old-versioned event is deserialized. **Copy-and-replace** instead treats the migration as a batch ETL job: read the entire existing stream from start to finish, run each event through whatever transformation logic is needed — which, unlike an upcaster, is free to consult external systems or reference data — and append the transformed events into a brand-new stream, typically under a new stream identifier, leaving the original stream intact as an immutable historical artifact. ## What each one can transform The mechanical difference matters because it changes what kinds of transformations are possible. An upcaster is constrained to be a pure function of the event's own payload, because it has to run correctly and identically every single time that event is read, potentially years apart. Copy-and-replace has no such constraint: because it runs once, as a deliberate, supervised operation, it can - join against a currency-rate table as of the event's timestamp, - call a lookup service, - or even have a human review edge cases before the new stream is finalized. This is precisely why copy-and-replace is the right tool exactly where upcasting is not: transformations that need context not recoverable from the event itself. ## Why it exists as a distinct pattern Why this exists as a distinct pattern, rather than always just adding more upcasters: an upcaster chain has a real, compounding cost a one-time rewrite eliminates. Every event ever written keeps needing translation on every future read, so after a system has been through several schema revisions, replaying the oldest events means running them through several chained transformations before domain logic sees them. Copy-and-replace lets a team 'pay down' that accumulated complexity: once the new stream exists and is verified, upcasters for versions with no live stream left can be retired, and future replay reads the current shape directly with zero translation cost. ## The trade-off The trade-off is that copy-and-replace is a bigger operation with real risk on both sides. **On the cost side:** - it requires a maintenance window or a careful dual-write/backfill/cutover dance to avoid losing events written during the migration; - it doubles storage during the transition; - it needs a verification step, often replaying both old-with-upcasters and new-plain and diffing the resulting state; - and unlike upcasting, which is naturally incremental and easy to roll back, a bad copy-and-replace can produce a corrupted new stream that's hard to detect until already in production use. **On the benefit side**, once done, you get a flat, single-version stream going forward with no ongoing translation tax, and the freedom to also restructure the stream's granularity — for example collapsing years of fine-grained events into fewer, richer ones. ## Failure modes at cutover In production, failure modes cluster around the cutover moment and verification gaps. - **If writers aren't paused or dual-written correctly during the batch rewrite,** events written mid-migration can be silently dropped from the new stream, leaving new-stream replay subtly incomplete versus the old stream — the event-sourcing equivalent of a lost-update race condition, easy to miss because both streams still 'work,' just with different content. - **A rewrite that isn't verified against replaying the original stream** can encode the same kind of silent misinterpretation error a bad upcaster would, except now it's baked permanently into the new canonical stream instead of being a fixable transformation function. A well-known real-world pattern for de-risking this, used in systems built around EventStoreDB or Kafka-backed event sourcing, is to write the new stream alongside the old, run consumers against both in parallel for a bake-in period, diff the resulting projections, and only decommission the old stream and its upcasters once confidence is high — essentially a **blue-green deployment applied to the event log itself.**
- Why can't an upcaster do the kind of transformation copy-and-replace can, such as looking up a historical exchange rate?An upcaster must be a pure, deterministic function of the event payload alone, because it's expected to run identically every time that event is deserialized, potentially years apart, with no external dependency that could change or become unavailable. A lookup against a live or historical reference table introduces exactly that kind of dependency, fine for a one-time supervised batch job but unsafe as a function expected to behave identically forever.
- What's the biggest operational risk during a copy-and-replace cutover, and how is it typically mitigated?The biggest risk is losing or duplicating events written during the migration window, since writers may still be appending to the old stream while the batch job is reading it. Teams typically mitigate this with a dual-write period, a verification pass diffing projections built from each stream, and only decommissioning the old stream once the new one is proven complete.
- If copy-and-replace can eliminate the ongoing cost of an upcaster chain, why not always do it instead of upcasting?Because it's a much bigger, riskier operation — it needs a migration window or dual-write machinery, doubles storage temporarily, and requires careful verification, whereas an upcaster is a small, easily-tested, easily-rolled-back function shipped incrementally. Upcasting is the low-cost default; copy-and-replace is reserved for cases upcasting genuinely can't handle or when accumulated chain cost has become a real problem.
Upcasting is like translating a foreign-language letter fresh every time someone asks to read it. Copy-and-replace is like commissioning one official translation of the whole archive, filing it as the new master copy, and retiring the need to translate on demand — worth the upfront project once the archive gets big enough or the translation needs a human judgment call a phrasebook can't make.
saying these in an interview costs you the question
- Thinks copy-and-replace and upcasting are interchangeable with no trade-off
- Doesn't mention the risk of events written during the migration window being lost
- Assumes copy-and-replace is 'free' just because it happens once
- Can't explain why an upcaster can't consult external data but a migration job can
- Forgets that the old stream and upcasters should stay available until the new stream is verified