During a rolling deploy, newer instances write a changed value format into a byte-opaque store while older instances still read those entries — what breaks, and how do you prevent it?
answer
- two versions run at once
- the store converts nothing
- readers first, writers second
- a tolerant reader is a lossy writer
- lifetimes and the next write bound it
basics
~20 sAn older instance fails on an entry written in a shape it does not know, or worse, misreads it. Deploy in two phases: readers that accept both shapes first, writers of the new shape second, old branch retired last.
solid answer
~50 sA gradual deployment means both versions run at once, sharing one keyspace. Because the store keeps bytes and validates nothing, a new-shape entry is written happily and only breaks when an older instance reads it — as a deserialization error if you are lucky, as a silently misread field if you are not. Make the change **two deploys, not one**: ship readers that tolerate both shapes and ignore unknown fields, then the writers of the new shape, then remove the old branch. Additive changes are cheap; renames, unit changes and type changes are not — express each as an addition now and a removal later. The window is bounded by the entries' own lifetimes plus how soon each is next rewritten, which is why this costs far less on a volatile tier than in a durable store.
go deeper
The point to hold on to: during a gradual deploy two versions run at the same time and share the same entries, and the store does nothing to reconcile them. A value written by the new version may be read by the old one.
Explain the phase order and why it is that way — tolerant readers everywhere first, new-shape writers second, old decoding path removed last — and classify changes into additive ones, which are safe, and renames or unit changes, which are not.
Show the failure you have actually seen: not the loud deserialization error but the tolerant reader that rewrote a whole value and dropped the field it did not recognise. Then bound the window with the entries' lifetimes and their next write, noting the entries that have neither.
Make this a fleet rule rather than a per-change decision: who may change the format, what the mandatory phase gap is, and how you know the old shape is gone. Weigh it against putting the new shape under its own namespace, which trades duplicated entries for a clean cut.
## The window a gradual rollout opens A rolling deployment is, by design, a period in which two versions of your service run simultaneously. If they share nothing, that is fine. But they share this keyspace, and on a store that hands back exactly the bytes it was given, the store will not tell either of them that they now disagree. So there is a window in which: - a **newer instance** writes an entry in the new shape; - an **older instance** reads that entry and must make sense of it; - and, symmetrically, an older instance keeps writing the old shape for the newer one to read. This is the one part of format evolution that belongs to the store: the entries an older process must still read after a newer one has written them. ## What goes wrong, in order of how bad it is 1. **A deserialization error.** The reader fails loudly on the entry. Unpleasant but honest — you get a stack trace and a bounded blast radius. 2. **A silently misread field.** The value still deserializes, but a field changed meaning or unit and the reader believes it. This is the dangerous one, because nothing errors and the wrong value propagates. 3. **A dropped field on rewrite.** An older instance reads a new-shape entry, tolerantly ignores the field it does not know, edits something else, and writes the whole value back — without the unknown field. The newer instance's data is gone, destroyed by a read-modify-write round trip that looked harmless. The third is the failure specific to byte-opaque storage, and it is worth saying out loud in an interview: when every partial change is a whole-value rewrite, a tolerant reader is also a lossy writer. ## The two-phase rollout 1. **Deploy readers that accept both shapes.** They understand the new fields, ignore what they do not recognise, and default what is absent. No instance writes the new shape yet. Wait until the whole fleet is on this version. 2. **Deploy the writers.** Now new-shape entries appear, and every instance can read them, because step one already finished. 3. **Retire the old branch**, but not before every entry written before the flip has aged out or been rewritten. Until then the old-shape decoding path is load-bearing. If the reader from step one must also *write* — which it will, given the read-modify-write round trip — it has to preserve fields it does not understand, or step three inherits a data-loss problem. Round-tripping unknown fields is the usual answer. ## What each kind of change costs an older reader | Change to the format | What an older reader sees | Safe in one deploy? | |---|---|---| | Add an optional field | an unknown field it can ignore | yes, if readers are tolerant | | Remove a field | a required field now absent | no — stop reading it first, remove later | | Rename a field | the old name gone, a new one unknown | no — write both, then drop the old | | Change a field's type or unit | a plausible value with the wrong meaning | no — this is the silent one; use a new field name | | Change the serialization itself | bytes it cannot parse at all | no — needs a version marker and both decoders | The pattern underneath the table is one rule: **express every change as an addition now and a removal later**, with a deploy in between. ## What bounds the window Nothing in the store converts stored values, so the old shape stops circulating only when the entries carrying it are gone. Two things do that: - **the lifetimes on those entries** — once every entry written before the flip has passed its deadline, no old-shape entry remains under a lifetime; - **the next write to each key** — any entry rewritten after the flip is converted by that write. Entries with no lifetime that nothing rewrites are the exception, and they will sit in the old shape indefinitely. That is worth checking before you plan step three around an assumption that everything ages out. ## Why this is cheaper here than in a durable store This is the consolation of a volatile tier: the entries are not the record of truth. There is no migration to write, no backfill, no historical data in the old shape to keep decoding forever. In the worst case you can decide that entries written before the flip are not worth reading, and let them be dropped — though "drop the sessions" and "drop the rendered fragments" are very different decisions, and only one of them is free. One structural alternative is worth naming and setting aside: some teams give the new shape a different key namespace so the two never meet, which turns the whole problem into a key-naming decision rather than a value-format one, with its own cost in duplicated entries during the transition. A store that can address named fields narrows all of this but does not close it. The server then knows the names, so a renamed field is at least visible rather than silently absent, and a reader that changes one field is no longer forced to rewrite the whole value and lose the rest. It still knows nothing about units or meanings, so the silent misread survives.
- Why is a tolerant reader dangerous if it also writes?Because every partial change on a byte-opaque store is a whole-value rewrite. A reader that ignores a field it does not know, edits another field and writes the entry back has just deleted the unknown field. The remedy is to round-trip unrecognised fields through the deserializer untouched, not to make the reader strict.
- How would you change a field's unit — say a duration stored in seconds becoming milliseconds?Never in place. A value in the new unit is still a plausible value in the old one, so an older reader accepts it and is wrong by a factor of a thousand with no error anywhere. Add a differently named field, move readers over, then remove the original once no writer produces it and no entry still carries it.
- The rollback case: you revert the writers. Does anything break?Only if the reverted version cannot read the new shape. Since phase one made every instance tolerant before any new-shape entry existed, a revert of the writers leaves readers that still cope. That is the second reason for the phase order — it makes the deployment reversible, not just forward-safe.
saying these in an interview costs you the question
- Assumes the store converts stored values to the newest format.
- Ships the format change as one deploy across the whole fleet.
- Changes a field's unit or type in place without renaming it.
- Forgets that a tolerant reader rewriting the entry drops unknown fields.
- Believes the rollout window itself bounds how long old entries survive.