A live endpoint must replace a field with an incompatible one while old callers keep sending the old shape. What does an expand-and-contract migration do in each phase?
answer
- one breaking edit, several compatible ones
- add, write both, read new, backfill
- the fallback branch needs a counter
- stop writing is the first breaking gate
- retire the old identifier at the end
basics
~20 sExpand-and-contract ships one incompatible change as a sequence of compatible ones: add the new form beside the old, write both, move readers to the new with a fallback, backfill what already exists, stop writing the old, then remove it.
solid answer
~50 sThe trick is that at no moment does a deployed reader meet bytes it cannot handle. **Expand**: add the new field alongside the old; nothing reads it yet, and old readers ignore it. **Dual write**: every writer populates both forms on every write, so the two are kept equivalent. **Read migration**: readers switch to the new form but fall back to the old when it is absent, which is still the case for records written before the dual-write rollout. **Backfill**: rewrite existing data so the new form is present everywhere, and watch a counter on the fallback branch until it stays at zero. **Contract**: writers stop populating the old form - this is the first step that can break a reader nobody upgraded, so it is gated on that counter and on the caller inventory. **Remove**: delete the old field and retire its identifier so it is never reused.
code
pseudocode · 6 lines# phase 3: prefer the new form, fall back to the old
function read_amount(record):
if record.has("amount_minor"): # new form: integer minor units
return record.amount_minor
increment(counter.fallback_reads) # must stay at zero before contracting
return round(record.amount * 100) # old form: decimal major unitsgo deeper
Recall that a breaking change is shipped as several small steps, and that for a while the payload carries both the old and the new form of the same value.
Explain the phase order and why each one is individually compatible: adding, dual writing, reading with a fallback, backfilling, then removing in two separate steps.
Demonstrate the gates: the counter on the fallback branch, a writer inventory that includes batch jobs, and the fact that halting the dual write is the first step that can break someone.
Weigh the carrying cost - every phase is deployed, reviewed and on-call time - against simply running two versions, and decide which changes are worth this and which are not worth making at all.
## Why a single incompatible edit cannot be deployed In any running system, writers and readers are upgraded at different moments: during a rollout both versions of the code are live at once, and data written yesterday is still being read today. So a change that is incompatible in one step has no safe ordering - deploy writers first and old readers meet bytes they cannot interpret; deploy readers first and they meet old bytes they no longer expect. **Expand-and-contract** (also described as parallel change) dissolves the problem by decomposing one incompatible edit into a chain of edits that are each individually compatible in both directions. The cost is time and discipline, not cleverness. ## The phases 1. **Expand.** Introduce the new form next to the old one. The payload now carries both shapes, or is capable of carrying both. Nothing writes the new one yet and nothing reads it; the only requirement is that a reader which knows nothing about it ignores it. 2. **Dual write.** Every writer populates *both* forms on every write. This is the phase where correctness is won or lost: the two must be genuinely equivalent, which means one canonical source and a deterministic derivation, not two independent code paths that can drift. 3. **Read migration.** Readers prefer the new form and fall back to the old when the new one is missing. The fallback is not defensive decoration - it is load-bearing, because records written before phase 2 and messages still in flight from an un-upgraded writer carry only the old form. 4. **Backfill.** Rewrite data at rest so the new form is present on every record. After this, the fallback branch should stop firing. Instrument it: a counter on the fallback path is the evidence for the next decision. 5. **Contract (stop writing the old form).** Writers stop populating the old field. This is the **first step that can break a reader nobody upgraded** - everything before it was reversible and invisible to laggards - so it is gated on the fallback counter sitting at zero and on knowing who still reads. 6. **Remove.** Delete the old field from the contract, and retire its identifier so a future change cannot reuse it for something else. ## What holds at each step | Phase | Old reader over the bytes | New reader over the bytes | Reversible? | |---|---|---|---| | Expand | Ignores the added form | Falls back to the old form | Yes | | Dual write | Reads the old form as before | Reads the new form | Yes | | Read migration | Unaffected | New form, fallback when absent | Yes | | Backfill | Unaffected | Fallback stops firing | Yes | | Contract | **Breaks** if it still exists | Fine | Only by resuming the dual write | | Remove | Breaks | Fine | No | ## The places it goes wrong - **Skipping the backfill.** Readers work perfectly on new traffic and fail on anything older, which surfaces weeks later as a report or a replay that cannot be read. - **A fallback with no counter.** Without the counter, "is anyone still hitting the old path?" is answered by opinion, and the contraction is either done too early (an outage) or never (the fallback outlives everyone who understands it). - **Two writers, one forgotten.** Dual write must cover every producer, including batch jobs, admin tooling and replays. The one that runs monthly is the one that is missed. - **Deriving the new value from the old at read time instead of writing it.** That is a shim, not a migration: it never lets you remove anything, because the old form stays load-bearing forever. - **Semantic, not structural, change.** If the new field means something subtly different - different units, different rounding, different time base - dual write has to state the conversion explicitly, and a lossy direction means the old form cannot be reconstructed. Decide that before phase 2, not during phase 5. ## What it buys Every phase is deployable on its own schedule, and every phase except the last two can be rolled back by redeploying the previous build. Where ecosystems differ is only in how the intermediate state is expressed - some contracts let both forms coexist naturally, others require the new form to be optional with a default - but the phase order is the same everywhere, because it follows from mixed-version deployment rather than from any format.
- Which phase is the last one you can roll back cheaply, and why?The backfill. Up to and including it, both forms are present on every record, so redeploying the previous build of any reader or writer still works. Once writers stop populating the old form, a rollback means resuming dual write and backfilling again before the old readers recover.
- The new field means the same quantity in different units. What does that add to the dual-write phase?An explicit, single-direction conversion with a decision about rounding, plus a check on the lossy direction. If minor units cannot reproduce the original decimal exactly, the old form stops being reconstructible from the new one, so the backfill must be treated as a one-way data change and verified on a sample.
- Why is a read-time shim that derives the new shape from the old not a migration?Because nothing ever gets removed. The old form stays the source of truth, every reader keeps a dependency on it, and the shim becomes permanent infrastructure. A migration is defined by the ability to reach the contract step; a shim is designed to avoid it.
saying these in an interview costs you the question
- Deploys writers and readers together and calls that a migration
- Treats the read fallback as optional defensive code
- Skips the backfill and breaks only on historical data
- Removes the old field in the same release that stops writing it
- Forgets batch jobs and replays when enabling dual write
- Assumes the old form can always be rebuilt from the new one