When a new encoder version re-encodes every stored player embedding, why write them to a separate key namespace and cut over at once rather than overwriting in place?
answer
- build it where nothing reads it
- verify before anyone consumes it
- one pin flip, not a rolling blend
- old vectors survive for rollback
- dual-write new players during the overlap
basics
~20 sA separate namespace keeps the two coordinate systems apart, so the switch is one pinned-version flip instead of a nightly-growing mixture, and the superseded vectors survive for rollback and for re-running old evaluations. The cost is double storage during the overlap.
solid answer
~40 sRe-encoding in place makes the population a moving blend of two incompatible spaces for as long as the pass runs, and destroys the old vectors as it goes. Writing into a second namespace makes the change atomic from the consumer's point of view: the new vectors accumulate where nothing reads them, you verify coverage and uniform stamps, retrain and gate the value model against the new space, then deploy the model version pinned to the new encoder version and namespace in a single step. Rollback becomes redeploying the previous model version, because its vectors are still there. You pay for it in storage - both copies are live during the overlap - and in a dual-write for players who install while the cutover is in flight.
code
pseudocode · 11 linesfunction embeddingFor(player, model):
ns = model.pinnedNamespace # "emb.v4" before cutover, "emb.v5" after
row = onlineStore.get(ns, player.key)
if row is null: # installed since the last encode pass
return MISSING # model reads a missing marker, not zeros
if row.encoderVersion != model.pinnedEncoderVersion:
reject("encoder mismatch in " + ns) # a namespace was built or written wrong
return row.vectorgo deeper
The idea to take away: build the new copy somewhere nothing reads it, check it, then switch. That pattern is older than machine learning and applies here for the same reason.
Explain why the consumer reads through a pinned namespace, and what the two live copies cost in storage while the cutover is in flight.
Walk the order of operations and catch the step most people miss - dual-writing newly installed players - then say when the old namespace gets deleted and who decides.
Frame it as buying optionality: the second copy is the price of a reversible adoption, and the retention window is an explicit trade of storage spend against how long rollback stays cheap.
## What "cut over at once" actually means It does not mean writing eighty million rows in one transaction. It means that from the consuming model's point of view the change happens at one instant, because the model reads a namespace it is pinned to, and the pin moves once. Before the cutover, the value model version in production pins encoder version 4 and reads namespace `emb.v4`. The encode pass writes version-5 vectors into `emb.v5`, which nothing reads. When the new namespace is complete and verified, a single deploy puts live a model version pinned to encoder 5 and `emb.v5`. There is never a moment when one model version reads a blend of the two. ## What the second namespace buys - **No mixed population.** The failure where error tracks refresh coverage cannot occur, because a partially built namespace is never read. - **Rollback is a redeploy.** The superseded vectors still exist, so reverting is putting the previous model version back, not a multi-night encode pass. - **Old evaluations stay re-runnable.** A regression cohort scored last month can be scored again against the exact inputs it used. - **The two spaces can be compared.** Scoring the same cohort against both namespaces measures what the new encoder is actually worth, before any user sees it. - **Verification is possible at all.** You can count rows, confirm every stamp says version 5 and confirm the width, *before* committing. ## What it costs | cost | size on this system | notes | |---|---|---| | storage during the overlap | 164 GB -> 328 GB | 512 coordinates at four bytes is 2,048 bytes per player, times eighty million, per namespace | | encode compute | one full pass over eighty million histories | unavoidable for any adoption of a new encoder | | dual-write for new installs | one extra write per new player | for the length of the overlap window only | | operational steps | a verify gate and a deletion date | the deletion reclaims the second copy | ## The order of operations 1. Create the new namespace and run the encode pass into it; nothing reads it yet. 2. During the pass, the job that encodes newly installed players writes into **both** namespaces, each with its own encoder version. Skip this and the old namespace silently loses everyone who installed since the pass began, which quietly breaks the rollback you are paying for. 3. Verify: row count against the population, every stamp reporting version 5, the width matching what the new consumer expects. 4. Retrain the value model on the new space and gate it offline against the version-4 model on a cohort with matured outcomes. 5. Deploy the model version pinned to encoder 5 and `emb.v5`. This is the cutover. 6. Keep the old namespace for a stated retention window, then delete it and reclaim the storage. The window is a decision, not an oversight: it is how long you are willing to keep rollback cheap. ## When rolling in place is defensible The rule is narrower than it sounds. Overwriting in place is fine when **no consumer was fitted to the old vectors** - a first launch, or a feature nobody reads yet - and it is fine for writing a *new* player's first vector with the current encoder, which is not a re-encode at all. What makes it dangerous is specifically the combination of a live consumer, a space change, and no surviving copy of what the consumer was trained on. It is also worth naming the thing a namespace does *not* fix: it does not make the new encoder better, and it does not remove the need to retrain every consuming model against the new space. It buys atomicity and reversibility, which is what turns a risky adoption into a routine one.
- What happens to players who install while the encode pass is still running?The job that encodes new players must write them into both namespaces during the overlap, each row stamped with that namespace's encoder version. Writing only into the new one leaves the old namespace missing every recent install, so a rollback lands on a store with holes exactly where traffic is growing fastest.
- How is a rollback performed after the cutover?Redeploy the previous model version, which pins the previous encoder version and the old namespace; both the artifact and its vectors still exist. That is only true inside the retention window, which is why the deletion date for the superseded namespace is a decision someone makes deliberately rather than a cleanup job nobody owns.
- Does a second namespace remove the need to retrain the consuming model?No. The new vectors live in a different space, so a model fitted to the old one cannot read them regardless of where they are stored. The namespace makes the adoption atomic and reversible; retraining the consumer against the new space is still the work.
saying these in an interview costs you the question
- Calls a rolling in-place refresh safe because each row stays fresh
- Assumes a rollback is possible after the old vectors were overwritten
- Forgets that new installs need writing into both namespaces during the overlap
- Thinks a new namespace alone lets the old model read the new vectors
- Keeps superseded namespaces forever with no deletion date or owner