A retrained sequence encoder is refreshing stored player embeddings in place, five million a night; the value model throws no errors but its daily error climbs all week. What is happening?
answer
- no errors, worse numbers
- two coordinate systems, one population
- the error tracks refresh coverage
- bucket the metric by encoder version
- in-place overwrite leaves no rollback
basics
~20 sThe rolling refresh is mixing two coordinate systems in one population: re-encoded players now sit in a space the value model was never fitted to. Nothing errors, and the error grows in step with refresh coverage.
solid answer
~40 sEach night's batch replaces version-4 vectors with version-5 vectors of identical width, so the population becomes a moving mixture of two incompatible spaces. The value model applies weights learned for the old space to whichever players have already been touched, and the blended error rises with coverage: after seven nights, 35 of 80 million players are refreshed - about 44% - and mean absolute error on the nightly backtest cohort has gone from 4.10 to 5.33. The diagnosis takes one query: bucket the cohort's error by the `encoderVersion` stamped on each player's row. The untouched bucket is unchanged at 4.10; the refreshed bucket sits near 6.9. Stop the job before going any further.
go deeper
The thing to hold on to: a model can be fed perfectly valid, perfectly fresh numbers and still be wrong, because validity says nothing about whether the numbers mean what the model expects.
Explain why a same-width re-encode passes every check on the read path, and why the error rises in step with how much of the population has been touched.
Name the diagnosis - bucket the quality metric by the encoder version stamped on each row - and say what you do first, which is stop the job before the mixture grows.
The interesting question is what made an unrecoverable rollout possible: an in-place mutation of a shared stored input, with no second namespace and no cost attached to overwriting history.
## The shape of the symptom Eighty million players each hold one stored history embedding in the online feature store. A retrained sequence encoder was shipped and an encode job is walking the population at five million players a night, overwriting each row in place - sixteen nights to full coverage. The quality signal is a nightly regression run: a fixed cohort of players whose thirty-day outcomes have already matured is re-scored every night with whatever the online store currently holds, and mean absolute error (MAE) is compared against those known outcomes. Over the week that MAE goes 4.10 -> 5.33. Meanwhile the service is healthy: no exceptions, no change in fetch latency, no null-rate movement, feature freshness green on every row. ## Why nothing errored A refreshed vector is indistinguishable from an old one to every guard on the path: - it has the same width, so a shape check passes; - it has the same scalar type and a similar numeric range; - the row exists, so the null check passes; - it was written last night, so it is the *freshest* data in the store - a staleness check would celebrate it; - the model returns a number for every request, so the error rate stays flat. Each individual read is perfect. What is wrong is a property of the *population*: half of it is expressed in a coordinate system the consuming model has never seen. Nothing on the serving path measures that. ## The arithmetic, which is also the tell Five million of eighty million is 6.25 points of coverage per night. After seven nights, 35 million players are refreshed: 43.75%. If the untouched bucket still scores at MAE 4.10 and the refreshed bucket scores at 6.9, the blended number is - 0.5625 x 4.10 + 0.4375 x 6.9 = **5.33** which is what the dashboard shows. The signature of this failure is that the metric tracks **coverage**, not the calendar and not traffic: it moves by roughly the same step every night, and it would keep moving if the job ran twice as fast. ## The query that settles it Bucket the regression run by the encoder version stamped on each player's row: ```sql SELECT r.encoder_version, COUNT(*) AS players, AVG(ABS(s.predicted_ltv - c.matured_ltv)) AS mae FROM nightly_backtest_scores s JOIN backtest_cohort c ON c.entity_key = s.entity_key JOIN embedding_rows r ON r.entity_key = s.entity_key WHERE s.run_date = DATE '2026-09-14' GROUP BY r.encoder_version; ``` Two rows come back, one good and one bad, and the bad one is the version that is still being written. That is a five-minute diagnosis if the rows carry a stamp and an unanswerable question if they do not. ## The fix, and two that are not 1. **Stop the encode job now.** Every further batch enlarges the mis-matched half. 2. **Re-encode into a separate namespace instead**, verify it, retrain the value model against the new space, gate it offline, then flip the model's pinned version and namespace together in one deploy. 3. **Accept that rollback is now expensive.** The job overwrote rows in place, so the version-4 vectors for 35 million players no longer exist anywhere in the online store. Restoring them means a fresh encode pass with the old encoder - the cost a second namespace would have avoided entirely. Two tempting non-fixes: - **"Let the refresh finish, it will settle."** It settles at the *wrong* level. At 100% coverage every input is from the new space and the MAE sits near 6.9 - consistent, and worse than the 5.33 that prompted the page. Consistency is not correctness. - **"Retrain the consumer now."** Retraining mid-refresh fits the model to a mixture whose composition changes every night, so the model is stale against its own inputs the moment the next batch lands. ## What the incident also cost The nightly regression run is no longer reproducible. Its labels are fixed, but its *inputs* were mutated underneath it, so last week's number and this week's number were computed on different data and the series compares nothing. Keeping each encoder version in its own namespace preserves the ability to re-run any past evaluation against the inputs it actually used. The general lesson is wider than embeddings: any in-place mutation of a stored model input whose meaning is defined somewhere else - by encoder weights, by a feature definition - produces exactly this shape of incident. No error, a metric that tracks rollout progress, and no way back.
- Once the refresh reaches 100% coverage, does the value model's error return to its old level?No. It settles near the refreshed bucket's number, around 6.9, because every input now comes from a space the model was never fitted to. The mixture becomes uniform, which makes the metric stop moving, but uniformly wrong is worse than half wrong. Only retraining the consumer on the new space, or restoring the old vectors, brings the error down.
- Why is rolling back harder than simply stopping the job?The job overwrote rows in place, so the previous vectors for the 35 million touched players are gone. Rolling back means re-encoding those players with the old encoder version - a full pass of encode compute and several nights of wall-clock - before the model is reading a single consistent space again.
- What did the in-place overwrite cost the nightly regression run itself?Its reproducibility. The cohort's labels are fixed but its stored inputs changed underneath, so the week's metric series mixes two input distributions and no past run can be reproduced. A per-version namespace keeps every evaluation re-runnable against the inputs it actually consumed.
saying these in an interview costs you the question
- Blames stale features when the vectors were written last night
- Says letting the refresh finish will restore the previous accuracy
- Hunts for an exception or a null-rate spike that a space change never produces
- Calls it a change in player behaviour rather than a changed input space
- Retrains the consuming model mid-refresh against a mixture that moves nightly