skip to content

A live tier's key convention must gain an owning prefix while callers keep reading and writing — how do you run that change, and how do you know the old shape is dead?

level: seniorimportance: should knowfreq 46%

answer

  1. one builder, first of all
  2. readers before writers
  3. both shapes exist for a while
  4. deletes are worse than updates
  5. instrument the fallback, then wait

basics

~20 s

Deploy readers that try the new key and fall back to the old, then flip writers, then let entries with a lifetime age out or backfill the rest. The old shape is dead when the instrumented fallback stays at zero.

solid answer

~50 s

Treat it as a rolling change with readers ahead of writers. First get every key built by **one function**, because you cannot migrate strings scattered across call sites. Then deploy readers that try the new key and fall back to the old on a miss — that release is safe on its own and changes nothing. Then flip writers to the new shape. From that moment new entries are only in the new shape and old ones are only in the old, so an entry that must stay current is written in **both** shapes, or only in the new one while the old copy is treated as dead. Entries carrying a lifetime migrate by attrition and need no backfill; entries without one need a deliberate pass. Instrument the fallback read so you can watch it fall to zero — that counter, not a code search, is what tells you the old shape is finished.

go deeper

for a junior

The takeaway is that changing a key means addressing a different entry, not renaming an existing one, so a live change needs both shapes to work at once for a while.

for a middle

Be able to sequence it: one key-building function, readers that fall back, then writers, then cleanup. Explain why readers must go first and what a reader from the previous release would see if they did not.

for a senior

This is the level's question. Own the window where both shapes exist — the dual-write failure gap, the delete that resurrects through the fallback, and the instrumented counter plus a wait longer than the longest lifetime before you remove anything.

for a principal

Ask what the convention change costs every other team on the tier, and design the next convention so it can be extended without a second migration — a stable owner segment first, room to append on the right.

## Why a convention is hard to change once it is carrying traffic The key is the only address the store understands. Change the string and you are not renaming a column — you are pointing at a different entry, one that does not exist yet. Meanwhile callers are reading and writing continuously, several versions of the application are live at once during any rolling deployment, and the store will happily let every shape coexist because it enforces nothing. There is no migration tool here for the same reason there is no schema. ## Step zero: get the key into one place Before anything else, every key must be produced by a single builder — owner, entity, identifier in, string out. If keys are concatenated inline at forty call sites, the migration is an archaeology exercise and you will miss some. This step is often most of the work, and it is worth doing even if the convention never changes again. ## The order of deployment 1. **Readers that understand both.** Deploy a release that reads the new key and, on a miss, reads the old one. Writers are untouched, so this release is inert in production: every read misses the new shape and falls back. Let it bake. 2. **Writers to the new shape.** Flip writing. New entries now appear under the new key only; existing entries still answer through the fallback. 3. **Deal with what is left.** Either let entries age out, or run a deliberate pass that reads an old entry and writes it under the new key. 4. **Remove the fallback**, then remove the old entries, then remove the builder's old branch — in that order, so that nothing is deleted while any reader can still want it. Doing this in the other order — writers first — means a reader from the previous release looks for an old key that nobody writes any more, and reports a miss for data that exists. ## The window in which both shapes exist This is the part candidates skip, and it is where the incident comes from. Between steps two and four, one logical thing has two possible addresses, and they can disagree: - an update written only in the new shape leaves a stale value under the old key, which an older reader still trusts; - an update written in **both** shapes costs two writes with no atomicity between them, so a failure between them leaves the shapes disagreeing; - a deletion is worse than an update: delete only the new key and the old one resurrects the value through the fallback. Pick one discipline and state it: either **write both and delete both**, accepting the failure window between the two operations, or **write only the new shape and treat old entries as read-only leftovers**, which is only safe once no caller can still write the old shape. Deleting must follow whichever rule you chose. ## How you know the old shape is dead Instrument the fallback. Increment a counter every time a read misses the new key and finds an old one, and label it by the calling site if you can. Then: - watch the counter fall as traffic naturally rewrites entries; - when it reaches zero, keep waiting — an entry read once a quarter has not been observed yet; - for a keyspace whose entries carry a lifetime, the honest waiting period is **longer than the longest lifetime in play** after the last non-zero reading; - for entries with no lifetime, waiting proves nothing, and you need the deliberate pass instead. What you should not do is confirm it by reading the code. A code search tells you nothing about what is in the store, and this class of store never had a schema to tell you either. ## What varies between stores, and why it matters to the plan - **Moving an entry server-side.** Some stores offer a single server-side step that gives an entry a new name; others do not, and the move becomes read-then-write in the client. Do not build the plan on the step existing. - **Where the keyspace is split across nodes**, the old and new key may sit on different nodes, so any step that wants to touch both at once may not be expressible. Which node a key lands on is a separate subject; the consequence for you is simply that a migration step should touch one key at a time. - **Finding the leftovers** is its own problem. Enumerating what remains under the old shape on a serving store is a traversal question, and it is the reason attrition through lifetimes is so attractive when it is available. ## The senior signal A candidate who has only read about this describes the new convention. One who has run it describes readers before writers, the window where both shapes exist, the rule for deletes inside that window, and the counter that gave them permission to delete the old branch.

  • During the window, why is a delete more dangerous than an update?
    An update that touches only one shape leaves a stale value under the other, which is wrong but plausible. A delete that touches only the new key leaves the old entry intact, and the reader's fallback then resurrects a value the application believes it removed — so the data comes back from the dead rather than merely going stale. Whatever discipline you choose for writing must be applied to deleting too.
  • The fallback counter has read zero for a week. Is it safe to delete the old entries?
    Only if you know the longest interval at which anything reads them. A week of zero proves nothing about an entry read once a month, and for entries carrying a lifetime the honest wait is longer than the longest lifetime in play. Where entries have no lifetime, waiting is not evidence at all and you need a deliberate pass over what remains instead.
  • Could you skip the dual shape by writing new keys and just accepting misses on old data?
    Sometimes, and it is the cheapest plan when the tier holds only regenerable entries with short lifetimes — every miss is repopulated and the old shape drains away. It is unacceptable where an entry is the only copy of something, such as state the application cannot reconstruct, because a miss there is a loss rather than a delay.

saying these in an interview costs you the question

  • Assumes the store can rename a whole keyspace in one operation.
  • Flips writers to the new shape before any reader understands it.
  • Dual-writes without a rule for what a delete must touch.
  • Calls the migration finished when the code stops writing the old shape.
  • Believes entries carrying a lifetime still need a backfill pass.
  • Confirms the old shape is gone by searching the source, not the traffic.