Why does a drift test whose reference window rolls forward with recent traffic miss the slow shift it should catch?
answer
- the baseline moves with the data
- measures change, not distance travelled
- the ruler stretches to fit
- step changes still break it
- frozen instead fires on seasonality
basics
~20 sA rolling reference re-estimates normal from data that already contains the drift, so each week's comparison is against last week's already-moved distribution. Gradual movement never accumulates in the statistic, though a step change still trips it.
solid answer
~50 sThe rolling window adopts the drift it was built to catch. Each run compares the detection window against a reference cut from recent traffic, which has already absorbed the previous weeks' movement, so a shift of a percent or two a week keeps the statistic near zero while the distance from the training distribution grows all season. It still catches **step** changes — an overnight sensor swap, a units change — because those break the recent past too. The opposite policy, a reference frozen at the training snapshot, does accumulate gradual movement, but on a seasonal system it eventually fires on legitimate phenology. The usable design is a frozen reference **aligned by calendar position** (week 27 against week 27 of the training seasons), with a short rolling reference alongside as a separate step-change test carrying its own threshold.
code
pseudocode · 17 linesevery week:
live = model_inputs_scored_since(now - 7 days)
frozen = reference_for(serving_model_version, calendar_week(now))
rolling = model_inputs_scored_between(now - 5 weeks, now - 1 week)
for f in model_input_features:
d_train = psi(live[f], frozen[f])
d_recent = psi(live[f], rolling[f])
if d_train > threshold_train[f]:
report('drifted from training', f, d_train)
if d_recent > threshold_step[f]:
report('step change', f, d_recent)
# gradual drift: d_recent stays near zero every week, d_train climbs
# overnight sensor swap: both cross on the first run after the swapgo deeper
Recall that the reference window is chosen, not given, and that a baseline recomputed from recent traffic will report calm while the inputs slide away from training.
Explain the mechanism: a rolling baseline re-estimates normal at the same rate the world moves, so it measures change between runs rather than distance from training.
Show the design you would actually ship: a seasonally aligned frozen reference for displacement plus a short rolling one for step changes, each with its own threshold and label.
Weigh the cost of nuisance firings against blindness. A calm dashboard nobody trusts and a noisy one nobody reads fail the same way, and the window policy is where that balance is set.
## The reference window is a policy, not a detail A shift test compares a **detection window** (what production just scored) against a **reference window**. The detection window is given. The reference is chosen, and the two common choices fail in opposite directions: - **frozen** — cut from the rows the model was trained on and never moved; - **rolling** — recomputed each run from the last N weeks of production traffic. For a national crop-yield forecast scored per field each week from overhead imagery and weather, the difference decides which failures the team ever hears about. ## Why a rolling reference goes quiet on exactly the drift you fear Suppose an atmospheric-correction change rolls out gradually across the imagery fleet, or an instrument degrades, moving a band's distribution about two percent a week. With a four-week rolling reference, every run compares this week against a baseline that already contains the last four weeks of that movement. The measured distance stays small and flat. Meanwhile the **cumulative** distance from the distribution the model was fitted on grows without bound, and by harvest the model is scoring inputs unlike anything it saw in training — with a dashboard that has been calm all season. The mechanism is simple: the rolling window re-baselines the definition of normal at the same rate the world changes. It measures *acceleration*, not *displacement*. What it does catch, and catches well, is a **step** change: a sensor swapped overnight, a units change in a weather feed, one region's ingest silently switching source. Those break the recent past as violently as they break the training distribution, so a rolling reference sees them immediately — often before a frozen reference has produced its next scheduled run. That is why a rolling reference is a good *second* test, not a bad idea. ## Why a frozen reference alarms on the calendar A frozen training-time reference accumulates displacement, which is what you want, but on a seasonal system it has a standing nuisance problem. Spectral inputs for a crop move enormously between week 6 and week 22 of a growing season **by design** — that phenological signal is the model's whole basis for forecasting. A reference pooled across a whole training season compares an early-season detection window against a season-wide average and fires with nothing wrong. Fire it every spring and the test stops being read. ## The resolution: freeze, but align 1. **Capture the reference from the training rows**, keyed by model version, so the test answers "how far is the live input from what this model learned?" 2. **Cut it by calendar or phenological position**, so week 27 of the live season is compared against week 27 of the training seasons — a seasonal-naive alignment. Only non-seasonal movement then trips the test. 3. **Run a short rolling reference alongside**, with its own threshold, reported as a distinct test, so a step change is caught within a run or two instead of waiting for displacement to build. | reference policy | the question it answers | what it goes blind to | its nuisance source | |---|---|---|---| | frozen at training | how far the live input is from what the model learned | nothing gradual, displacement accumulates | seasonality, and a change in the population mix | | rolling recent window | did something break in the last few weeks | gradual drift, which it re-baselines away | ordinary week-to-week variation | | frozen, seasonally aligned | non-seasonal movement since training | a shift that mimics normal seasonal movement | a season unlike any training season | ## Consequences to state out loud in a design round - The two tests need **different thresholds** and different labels in the output; collapsing them into one number makes the result uninterpretable. - When the model is retrained, the frozen reference is **recaptured** from the new training rows and published with the new model version, and thresholds derived against the old baseline do not carry over unchanged. - A seasonally aligned reference needs enough past seasons to have a comparable week; in the first season you have a pooled baseline and should say so rather than pretend the alignment exists. - No reference policy tells you whether the shift *hurt* the forecast. It tells you the input moved and by how much; what that cost is a separate reading on a separate timeline. The interviewer is listening for one sentence: the reference window is a design choice with two distinct failure modes, and you know which one you chose and why.
- Is there a case for running both a frozen and a rolling reference at once?Yes, and it is the usual answer. They ask different questions: the frozen one measures displacement from what the model learned, which is the retrain conversation; the rolling one measures whether something broke lately, which is the upstream-incident conversation. Give them separate thresholds and report them as separate tests, because one number blending both is uninterpretable.
- What happens to the frozen reference when the model is retrained on the newest season?It is recaptured from the new training rows and published with the new model version, because that model's notion of normal is now a different distribution. Thresholds derived against the old baseline do not transfer unchanged, since the statistic's scale depends on both sides of the comparison; re-derive them and keep the old reference for reproducing past firings.
- How do you stop a frozen reference from alarming every spring on legitimate phenology?Align the comparison by calendar or phenological position rather than pooling the training season: compare this week's detection window against the same week of the training seasons. Only non-seasonal movement then crosses the threshold. In a first season with no comparable week, use the pooled baseline and label the test as provisional rather than tuning the threshold until it goes quiet.
A rolling reference is a tape measure that restretches to whatever you last measured, so it always reads no change. A frozen one keeps its markings but reads the tide as a leak.
saying these in an interview costs you the question
- Picks a rolling reference because it needs no stored baseline
- Says a rolling reference detects nothing at all, missing step changes
- Treats seasonal movement in the inputs as drift to be alarmed on
- Reuses the training-time threshold after recapturing the reference window
- Blends the frozen and rolling results into one drift number
- Assumes a quiet drift statistic proves the forecast is still accurate