skip to content

Why does an autocomplete promotion gate score every weekly candidate on the same frozen, time-ordered window of logged sessions?

level: middleimportance: should knowfreq 52%

answer

  1. the same sessions every week
  2. comparable, not just representative
  3. carved out of every training set
  4. replay as of the session's timestamp
  5. the window ends before the training cut

basics

~20 s

Freezing the window makes weekly scores comparable: champion and challenger are replayed over identical sessions, so a change in the number means a change in the model and not an easier week. Time-ordering keeps each replay honest.

solid answer

~40 s

The gate replays prefixes from logged sessions: for each session the suggester is asked what it would have offered at that keystroke, and the result is checked against the query the user actually submitted. Two properties turn that measurement into a gate. **Frozen** — the same sessions every refresh, carved out of every candidate's training data, so this week's score is comparable with last week's and no candidate is scored on sessions it memorised. **Time-ordered** — each session is replayed as of its own timestamp, so evidence that only existed afterwards, such as a query that trended the next day, is not credited. The window sits in the past, ending before the candidate's training cut, because a refresh trained through yesterday has no labelled traffic after it to score on.

code

pseudocode · 17 lines
pseudocode
// the window is fixed once, and excluded from every candidate's training data
backtest = sessions where WINDOW_START <= session.time < WINDOW_END
assert candidate.training_cut >= WINDOW_END
assert backtest is disjoint from candidate.training_sessions

function replay(model, sessions):
    hits = 0
    for each s in sessions:
        // as_of forbids evidence dated after this keystroke
        suggestions = model.suggest(prefix = s.prefix, as_of = s.time)
        if s.submitted_query in suggestions:
            hits = hits + 1
    return hits / count(sessions)

champion_score   = replay(registry.live_model, backtest)
challenger_score = replay(candidate, backtest)
passes_backtest  = challenger_score >= champion_score + PROMOTION_MARGIN

go deeper

for a junior

Know what a replay is: the model is asked what it would have suggested at that keystroke, and judged against the query the user actually submitted in that session.

for a middle

Explain the two properties separately — frozen gives comparability across refreshes, held out keeps the score honest — and say why they are different requirements.

for a senior

Talk about maintenance: spotting a window that has drifted from live prefixes, re-cutting it deliberately, and accepting that the re-cut breaks comparability with every earlier score.

for a principal

Weigh a permanently held-out window against the training data it removes, and decide what comparability across unattended refreshes is worth to the business.

## What a replay actually measures A backtest for an autocomplete suggester is a **replay**, not a test-set score in the usual sense. The gate takes a logged session, reconstructs the prefix the user had typed at some keystroke, asks the model for the suggestions it would have produced at that instant, and compares the list with the query the user finally submitted in that session. The promotion metric is built from that comparison — for example the share of replayed sessions whose submitted query appeared in the suggestions, or a rank-sensitive version of the same idea. Because this is the number that decides an unattended promotion every week, it has to hold two properties that are easy to confuse with each other. ## Frozen: comparability across refreshes **Frozen** means the window is a fixed set of sessions, chosen once, replayed identically by every refresh. - A stated promotion margin — "promote when the challenger is ahead by at least this much" — only means something if the thing it is measured on does not move. - Week to week, a score change then attributes to the model rather than to whichever fortnight happened to be replayed. - The champion's score is recomputed on the same window with its own configuration each time, so the pair being compared is genuinely a pair. - A rejected candidate's score is meaningfully filed next to last month's rejected candidate. The cost is that a frozen window **ages**. Its prefixes were typed in the past, and the products, spellings and events they lead to belong to that past. The longer it stays frozen, the less it resembles the traffic a promoted model will actually serve. ## Held out: why a window inside the training data is worthless Frozen is not the same requirement as **held out**. The window's sessions must be excluded from every candidate's training set — carved out of the timeline permanently, not merely skipped once. Otherwise the gate rewards memorisation: a candidate trained on those exact sessions reproduces the submitted query without having learned anything transferable, and beats a champion whose own training predates them. | window design | what it buys | what it costs | |---|---|---| | frozen, held out of training | comparable, honest scores across refreshes | those sessions never train a model; the window ages | | rolling recent window | always resembles live traffic | this week's score is incomparable with last week's | | frozen but inside training data | nothing usable | flatters whichever candidate memorised it | | several frozen windows by segment | catches a regression confined to one traffic slice | more gate configuration, more ways to get stuck | ## Time-ordered: the direction of the leak Within the replay, the suggester may only be credited with evidence that existed **as of the replayed session's timestamp**. A query that first trended the following week must not appear among the suggestions for a prefix typed before it existed — if it does, the replay measures hindsight, and the score cannot be reproduced in production where no such evidence is available at request time. This is also why the window ends before the training cut rather than after it. A refresh trained through yesterday has nothing after its cut to be scored on: those labelled sessions do not exist yet. The held-out window therefore lives in the past, and the as-of rule is what stops the model's later knowledge from flowing backwards into the replay. It bounds that advantage rather than eliminating it — a challenger trained through a later cut has seen more of the world than the champion it is compared against, which is one reason an honest gate treats a narrow backtest win with suspicion. ## What the window cannot tell you - It cannot see **suggest latency**, so a slower candidate scores exactly like a fast one. - It cannot see users who **abandoned** the search box without submitting anything. - Its sessions were logged while the champion was serving, so the suggestions that shaped them were the champion's; a replayed challenger is judged on behaviour it never produced. That is why a backtest win is a filter and not a verdict: clearing it earns the candidate live exposure under the gate's guardrails, and nothing more. ## Maintaining the window itself Eventually the window is too old to be useful and has to be re-cut from more recent held-out sessions. That is a deliberate, recorded event, because it breaks comparability with every earlier score: the margin has to be re-justified on the new window, and rejected-candidate history from before the swap is not directly comparable with history after it. Teams usually re-cut on a long, stated cycle rather than whenever a candidate is rejected — re-cutting in response to a rejection is how a gate quietly becomes whatever the latest candidate happens to be good at.

  • The frozen window says the challenger wins. Why is it still not promoted on that number alone?
    A replay cannot see suggest latency, cannot see the users who abandoned the box, and scores sessions whose suggestions the champion itself produced. So the gate treats the backtest as a filter: a candidate that clears it is then exercised against live prefix traffic under the stated guardrails, and only a candidate that survives both moves the registry's live pointer.
  • What breaks if the window is rolled forward to the most recent two weeks at every refresh?
    Comparability. Each weekly score is then computed on a different population, so a rise can mean an easier fortnight rather than a better suggester, and a fixed promotion margin no longer means the same thing week to week. It also risks scoring a candidate on sessions inside its own training range, which flatters it against a champion trained earlier.
  • How would you notice that the frozen window has aged out?
    Compare the window's prefix distribution with live traffic on a schedule: the share of live prefixes whose completions never occur in the window at all is a blunt but effective signal. A steadily rising rejection rate with no change to the pipeline points the same way, and both argue for re-cutting the window as a recorded event.

saying these in an interview costs you the question

  • Scoring each week's candidate on whatever the last week of traffic happened to be
  • Evaluating a candidate on sessions that sat inside its own training data
  • Letting a replay credit a query that only became popular after that session
  • Treating the frozen backtest score as a prediction of the live acceptance rate
  • Believing a frozen window stays representative of live prefixes indefinitely
  • Re-cutting the window whenever a candidate is rejected