A promoted weekly autocomplete refresh degrades the live acceptance rate overnight. What must the automatic rollback restore?
answer
- the pointer, not a redeployment
- promote a bundle, roll back a bundle
- keep the previous version loaded
- cache keyed by model version
- quarantine, or the loop re-promotes it
basics
~20 sEverything the promotion changed, not just the model file: the registry's live pointer, the feature definitions, index snapshot and cutoffs promoted with it, and cached suggestions keyed to the bad version. The previous version stays loaded so the swap is a pointer move.
solid answer
~40 sTreat what was promoted as a **bundle** — the model artifact, the feature definitions it expects, the suggestion index snapshot it was built against, and the serving configuration tuned for it — versioned together, because the rollback unit has to equal the promotion unit. Rolling back then means pointing the registry's live version at the previous bundle, which the fleet has kept loaded, so the swap needs no restart and in-flight requests finish on the version they started with. Two things are routinely forgotten. Cached suggestions produced by the bad version must be keyed by model version or purged, or the rolled-back system keeps serving the regression from cache. And the reverted version must be **quarantined**, or next week's unattended gate rebuilds and promotes the same regression.
code
json · 22 lines{
"model": "query-suggester",
"live": "2026-w36",
"rolled_back_from": "2026-w37",
"versions": {
"2026-w37": {
"artifact": "sha256:1f3a9c...",
"feature_defs": "v13",
"index_snapshot": "2026-09-14",
"config": { "max_suggestions": 8, "min_prefix_len": 2, "score_cutoff": 0.31 },
"state": "quarantined",
"quarantine_reason": "acceptance rate -3.1% vs 2026-w36 over a 6h window"
},
"2026-w36": {
"artifact": "sha256:9c0417...",
"feature_defs": "v12",
"index_snapshot": "2026-09-07",
"config": { "max_suggestions": 8, "min_prefix_len": 2, "score_cutoff": 0.28 },
"state": "live"
}
}
}go deeper
Know that rolling back means pointing serving at the previous model version, and that the previous artifact must still exist for that to be possible at all.
Explain that a promotion may have changed more than a model file — feature definitions, an index snapshot, cutoffs — so they version together and revert together.
Cover the live mechanics: both bundles loaded so no restart is needed, a cache keyed by model version, and the reverted candidate quarantined against re-promotion.
Decide how far to trust automation: what an unattended reversal may do alone, what it must page a human for, and how long previous bundles stay warm and retrievable.
## What actually changed at promotion The instinct is that a promotion swapped a model file, so a rollback swaps it back. For a suggester that is usually wrong, because a weekly refresh can carry several coupled changes at once: - the **model artifact** itself; - the **feature definitions** it consumes — a new normalisation of the prefix, a recency feature the previous model never read; - the **suggestion index snapshot** the candidate was built against; - the **serving configuration** — how many suggestions to return, the minimum prefix length, a score cutoff tuned to this candidate's score distribution. If those are promoted separately, a rollback can restore the old artifact while leaving the new feature definitions and the new cutoff live. The previous model then runs on inputs and thresholds it was never evaluated with: a third state nobody tested, and often worse than either of the two intended ones. **Promote a bundle, roll back a bundle.** One version identifier covering all four, recorded in the model registry, is what makes an unattended reversal safe. ## The rollback unit, surface by surface | surface | restored by | what is missed if only the artifact is swapped | |---|---|---| | model artifact | the registry's live pointer moves back | nothing | | feature definitions | pinned by the bundle version | the old model reads features it was not trained on | | index snapshot | the previous snapshot is retained on disk | old model against a new index shape | | serving config and cutoffs | versioned with the bundle | thresholds tuned for the reverted candidate stay live | | cached suggestions | cache keyed by model version, or purged | the regression keeps being served from cache | | the gate's own state | the version is quarantined | next week's refresh re-promotes the same regression | ## What serves during the overlap The swap should be a **pointer read, not a deployment**. The serving fleet keeps the last promoted bundle and its predecessor loaded; each request reads the live version once at the start and uses that version end to end, so no request mixes the new model with the old cutoff. In-flight requests drain on the version they began with, and no process restarts. The practical consequence is that the reversal completes in the time it takes the pointer change to propagate rather than the time it takes to redeploy a fleet — the difference between a seconds-long regression and a twenty-minute one. The prediction cache deserves its own sentence. Autocomplete serving usually caches suggestions for hot prefixes. If the cache key does not include the model version, then after the pointer moves back the cache keeps handing out the bad version's suggestions until its entries expire, and the graphs stay red long enough for someone to conclude the rollback did not work. Keying by version makes the old entries unreachable immediately; a purge does the same thing more bluntly, at the cost of a cold-cache latency spike. ## Arming the reversal The rollback is armed by the same guardrails the gate stated before promotion — for a suggester, the acceptance rate against its pre-promotion baseline, suggest latency, the share of prefixes returning no suggestions, and the serving error rate. Three parameters make an unattended trigger workable: a comparison baseline, a fixed observation window, and a minimum volume before the comparison may fire. A quality regression that raises no errors is exactly the case this exists for, because nothing else in the stack notices a suggestion list that is merely worse. ## Stopping the loop from re-promoting it This is the part that separates a standing gate from a one-off launch. The pipeline that produced the bad bundle has not changed, and it runs again next week. If the cause is still present — a feature that silently went empty, a corrupt index snapshot — the next refresh can produce the same artifact, clear the same frozen backtest and promote the same regression. So the reversal writes state: the version is marked quarantined with the reason and the measurement that tripped it, and promotion stays blocked until a human clears it or a stated condition is met. Without that, an automatic gate becomes an automatic oscillator. ## The failures teams actually hit 1. **The previous artifact was garbage-collected** by a retention rule, so the reversal has no target and becomes a rebuild. 2. **The rollback path was never exercised**, and the first real attempt discovers the previous bundle no longer starts against the current feature-definition version. 3. **The cache outlives the swap**, extending the incident by however long its entries live. 4. **The regression is recorded nowhere**, so a month later the same candidate shape is promoted again and nobody connects the two events. Rehearsing the pointer move on an ordinary day — promote, revert, confirm that the live version and the cache both followed — is what turns all four from discoveries into non-events.
- How long should the previous suggester bundle stay loaded and serving-ready?At least as long as the guardrail needs to see a regression: a signal measured over a six-hour window cannot justify discarding the previous bundle after one. Teams typically keep the last one or two promoted bundles warm on the fleet and older artifacts retrievable, so a same-day reversal is a pointer move and an older one is a redeploy.
- Why quarantine the reverted version rather than simply leaving it unpromoted?Because the pipeline that produced it has not changed and runs again next week. If the cause is still present — an empty feature, a bad index snapshot — the next refresh can rebuild the same artifact, clear the same frozen backtest and promote the same regression. Quarantine records the version and the reason, and blocks promotion until someone clears it.
- Acceptance drops but suggest latency and error rate are unchanged. Can the reversal still be automatic?Yes, provided the guardrail is stated as a comparison against the pre-promotion baseline over a fixed window, with a minimum volume before it may fire. A silent quality regression that raises no errors is precisely what a model-version rollback is for; requiring an error signal would make the gate blind to the commonest case.
saying these in an interview costs you the question
- Rolling back the model file and leaving the new feature definitions live
- Restarting the serving fleet to change which model version answers
- Rebuilding the previous model from its snapshot instead of keeping the artifact
- Letting cached suggestions from the bad version expire on their own
- Allowing the next scheduled refresh to re-promote the version just reverted
- Assuming the rollback path works because it has never been needed