Your weekly autocomplete gate has promoted no candidate in two months while every dashboard is green — what do you check?
answer
- silence is not health
- read the gate's rejection reasons
- a margin above the achievable gain
- a veto tripping on the serving tier
- alarm on time since last promotion
basics
~20 sCheck the gate, not the models. The usual causes are a promotion margin above what a weekly refresh can achieve, a veto condition tripping for a reason outside the model, or a frozen backtest window that has aged away from live prefixes.
solid answer
~40 sSilence from an automatic gate looks exactly like health: the champion keeps serving, no rollback fires, and nothing is wired to alarm on a promotion that never happened. Start with the gate's own record — every refresh should log which condition it failed. The common causes are a promotion margin larger than the improvement a weekly refit can produce, so every candidate is "not better enough"; a veto tripping on something outside the model, such as a suggest-latency regression caused by an infrastructure change, which rejects candidates indefinitely; a frozen backtest window that has drifted from live prefixes, so candidates better on today's traffic score worse on it; and a broken comparison, where the champion is rescored with the challenger's configuration. The durable fix is an alarm on the gate's own outcome.
go deeper
Know that an automatic gate can reject every candidate and still look healthy, because nothing alarms on a promotion that simply never happened.
Explain the checks in order: the recorded rejection reason first, then the promotion margin, the veto conditions, and the age of the frozen backtest window.
Show the instrumentation: rejection reasons counted by cause, time since last promotion alarmed, and the challenger-champion gap tracked against the margin over time.
Own the trade-off: a high margin forfeits real but modest gains, a low one buys churn and reversals, and someone must name which the product can afford.
## Why a stuck gate is invisible Every alarm around a retraining loop watches something that moved. The live suggester's guardrails watch the champion, which is behaving. The rollback path watches a promotion, and there has been none. The pipeline's job monitor watches for failures, and the weekly refresh succeeds every time — it produces an artifact, the gate scores it, the gate says no. Nothing in that sequence is an error, so nothing pages. Meanwhile the live model is months old, which is the exact condition the whole loop exists to prevent. The first move is therefore not to look at a model. It is to read the gate's decision record for the last eight refreshes and ask *which condition failed*, because the four causes below have entirely different fixes. ## The four things to check 1. **The promotion margin is unreachable.** The margin says how far ahead a challenger must be. If it was picked as a round number rather than from what a weekly refit realistically gains, every candidate lands under it. The symptom is distinctive: challenger scores consistently a little *above* the champion's and consistently below the bar. 2. **A veto condition is tripping on something that is not the model.** Guardrails such as suggest latency or error rate are measured on infrastructure the model shares. A slower host generation, a noisy neighbour, a change in how the latency is measured — any of these can push the measurement past a fixed veto and reject candidates forever, because no retraining run can fix a serving-tier regression. 3. **The frozen backtest window has aged out.** Its prefixes and their completions are from the past. A candidate trained on recent traffic is better at today's prefixes — new product names, new spellings — and those sessions are not in the window at all, so its improvement is invisible there while every change to old behaviour counts against it. An aged window inverts the gate: it starts rejecting exactly the candidates that would win. 4. **The comparison itself is broken.** The champion's score has to be recomputed on the same window with its own configuration. If a config change leaks so the champion is rescored with the challenger's cutoff, or if the champion's score is a cached number from months ago, the difference being tested is not the one anyone intended. | what the gate log shows | likely cause | fix | |---|---|---| | challenger ahead, under the margin, every week | margin set above the achievable weekly gain | re-derive the margin from observed refit-to-refit improvement | | rejected on a veto, the metric never evaluated | a guardrail tripping on the serving tier | fix or re-baseline the guardrail; it is not a model problem | | challenger behind on the window, ahead on live-shaped checks | the frozen window has aged | re-cut the window from recent held-out sessions, as a recorded event | | the champion's score moves with no promotion | a broken or stale comparison | rescore both models in the same run with their own configurations | ## The margin is an operating decision A margin is not a safety device that gets safer as it grows. Set high, it forfeits every improvement that is real but modest, and those are most of them — the loop's value is the accumulation of small weekly gains. Set low, it buys churn: more promotions, more reversals, more cache invalidation, more versions to reason about during an incident. Someone has to name which the product can afford and write the number down with its justification, so the next person does not raise it after one bad week. Raising the margin in response to a single regression is the commonest way a healthy gate becomes a stuck one. ## Monitor the gate, not only the model The instrumentation that prevents this is small and almost never present: - **Time since the last promotion**, alarmed against the refresh cadence. Eight consecutive rejections on a weekly loop is an incident even though nothing is broken. - **Rejection reasons counted by cause**, so the alarm arrives as a diagnosis rather than a mystery. - **The gap between challenger and champion over time**, which renders an unreachable margin as a flat line just below the bar. - **The age of the frozen window**, reviewed on a stated cycle rather than when someone happens to remember. ## The opposite failure A gate tuned only to avoid being stuck develops the symmetrical problem: a margin so small that ordinary run-to-run variation clears it, promoting a new bundle every week and reverting a good share of them. The pattern to watch is a rising ratio of rollbacks to promotions. Both failures come from the same underlying mistake — a promotion rule chosen by feel and then never revisited — and both are why the gate's configuration deserves the same review as the model it governs.
- Which single signal would have caught this two months earlier?Time since the last promotion, alarmed against the refresh cadence. On a weekly loop, eight consecutive rejections is an incident even though no component failed, and the alarm belongs to the gate rather than the model. Counting rejections by their stated reason turns that alarm into a diagnosis instead of the start of an investigation.
- How can an aged frozen window reject a candidate that would genuinely win on live traffic?The window's prefixes are from the past. A candidate trained on recent traffic is better at today's prefixes — new names, new spellings — and none of those sessions appear in the window, so its real improvement is unmeasurable there. Meanwhile any change to how it handles older prefixes shows up in full, and counts against it.
saying these in an interview costs you the question
- Assuming no rollback and no alarm means the retraining loop is healthy
- Reading a stuck gate as evidence the live suggester is good enough
- Raising the promotion margin whenever one bad candidate slips through
- Never recording why a candidate was rejected
- Treating a frozen backtest window as permanently representative
- Blaming the model when the veto is measuring the serving tier