In a weekly-refreshed query autocomplete suggester, why must each candidate beat the live model before it replaces it?
answer
- the incumbent is the baseline
- newer data, not a better model
- a refresh yields a candidate, not a release
- one live version, one registry pointer
- the rule runs without a human
basics
~20 sNewer training data does not guarantee a better suggester, and a refresh can be worse without failing. The model already serving is the baseline every candidate must clear on a stated rule, so promotion is earned rather than automatic.
solid answer
~40 sRetraining runs on a cadence and its output is a *candidate*, not a release. The model already serving prefixes — the champion — is the only honest baseline, because promotion changes what users get relative to what they get today. A weekly refit can be worse for unremarkable reasons: a skewed log week, a logging change that silently emptied a field, an artifact that is slower than the one it replaces. None of those fail the pipeline. Since nobody reads each weekly candidate, the comparison has to be written down as a rule — a metric, a margin over the champion, veto conditions — and evaluated by the pipeline itself. Until a candidate clears that rule, the model registry's live pointer keeps naming the champion.
go deeper
Recall the two roles: the champion is serving users, the challenger is this week's artifact, and only a stated comparison moves one into the other's place.
Explain what makes a refresh worse without failing — a hole in a logged field, a skewed week, a slower artifact — and why the comparison is written as configuration.
Show the operational side: a single registry pointer, the previous version kept ready, and a rejected candidate that is recorded and counted rather than silently discarded.
Argue what an automatic gate buys the organisation — refreshes that ship without a meeting — and what it costs, since only what was written down in advance is ever checked.
## The two roles A query autocomplete suggester turns a half-typed prefix into a short list of completions. It is refit on a cadence — weekly, say — because the query log it learns from moves: new product names, new spellings, a news event that invents a phrase overnight. Each run of the retraining pipeline produces an **artifact**: a trained model plus the data snapshot, feature definitions and configuration it was built from. That artifact is the **challenger**. The model answering live prefix requests right now is the **champion**. Champion-challenger promotion is the rule that decides, with no human in the room, whether this week's challenger takes the champion's place. The load-bearing word is *replaces*. There is exactly one live suggester behind the search box, and a model registry records which version that is. Promotion is not "ship the new model" — it is "move the registry's live pointer, and keep the ability to move it back". ## Why newer data does not mean a better model Retraining is an industrial process whose output nobody reads. Plenty of ordinary things make a refresh worse than the model it would replace, and none of them fail the pipeline: - **A skewed input week** — an outage, a bot wave or a campaign distorts the query distribution the candidate learned from. - **A silent upstream break** — a logging change stops populating a field, and the training job succeeds on a hole. - **A shorter memory** — a refresh weighted to recent traffic forgets rare but real prefixes the champion still handles. - **A non-quality regression** — the candidate is larger and slower, so suggestions land after the user has finished typing. No error is raised; the box is just worse. - **Plain run-to-run variance** — two refits of the same pipeline are not identical, and one of them is a little worse. Each of these produces a green pipeline run and a bad artifact. That is precisely the case a promotion gate exists to catch. ## The champion is the baseline, not an absolute bar It is tempting to gate on a fixed number: promote anything scoring above some threshold. That bar is hard to justify in the abstract and it drifts out of date, because the only change promotion makes in the world is *this model instead of that one*. Comparing candidate against champion on identical data also cancels much of what the two have in common — the same prefixes, the same seasonality, the same catalogue — so what remains is closer to the difference the swap would actually cause. The cost of a relative rule is worth stating plainly: it never asserts that the champion is good. If the world moved and the live suggester is now poor, a gate demanding "better than the champion" will happily promote a slightly-less-poor challenger and say nothing about the underlying decay. The quality of the live model is a monitoring question; promotion is a comparison question. ## What the standing rule has to state | element | what it fixes | for a suggester | |---|---|---| | the comparison set | both models scored on identical sessions | a frozen window of logged sessions, held out of training | | the promotion metric | what "better" counts as | share of replayed sessions whose submitted query appeared in the suggestions | | the margin | how far ahead is far enough | a stated improvement over the champion on that same window | | the veto conditions | what disqualifies regardless of the metric | suggest latency, empty-suggestion rate, error rate on live prefixes | | the pointer | what promotion actually changes | the model registry's live version for this model | | the reversal | what happens when it was wrong | an automatic model-version rollback | ## The loop, end to end 1. The scheduled refresh runs and registers its artifact as a challenger; the registry's live pointer still names the champion. 2. The gate replays the frozen window for both models and compares the two scores against the stated margin. 3. A candidate that clears it is exercised against live prefix traffic with its output discarded, to show it runs inside its latency and error budget. 4. If nothing vetoes it, the live pointer moves to the challenger, which becomes the champion. 5. Live guardrails keep watching; a regression moves the pointer back automatically and quarantines that version. ## Where a standing gate differs from a one-off launch A new suggester architecture, hand-built once a year, gets a human: someone designs the comparison, watches the traffic and decides. The standing gate is the opposite case — it reaches the same kind of decision every week with nobody watching, which is why every part of it is configuration rather than judgment. The price of that automation is that anything the human would have noticed has to be written down in advance, and anything not written down is not checked. ## What an interviewer is listening for The weak answer treats retraining as deployment: "the pipeline runs weekly and pushes the new model". The answer that lands names three moving parts — a candidate that is not live, a comparison against the model that is, and a pointer that can move both ways — and then admits the limit: the gate cannot tell you the champion is good, only whether the challenger is better.
- The weekly refit has finished and the artifact is registered. What is its state before the gate runs?It is a challenger: a versioned artifact in the model registry, linked to the data snapshot, feature definitions and configuration that produced it, and serving no user traffic. The registry's live pointer still names the champion. Its only job at that point is to be scored — first against the champion on the frozen backtest window, and only then against live traffic.
- Does the champion ever have to re-prove itself?It never has to win anything, but the same live guardrails measure it continuously. If the champion decays because the world moved, the promotion gate will not fix that: no challenger is required to be good, only better. A champion that is quietly getting worse is a monitoring and refresh-trigger problem, which is why a gate that keeps rejecting deserves an alarm of its own.
A dish does not go on the menu because the recipe is newer; it replaces the one already there only if a tasting says it is better. The old recipe stays in the book, because that is what you go back to.
saying these in an interview costs you the question
- Assuming a fresh refit is better because its training data is newer
- Promoting every weekly candidate and watching the dashboards afterwards
- Treating a green training run as evidence about model quality
- Gating on a fixed absolute score instead of the live model's score
- Expecting an engineer to review each weekly candidate by hand