skip to content

questions

6

Why does a drift test whose reference window rolls forward with recent traffic miss the slow shift it should catch?

level: middleimportance: must knowfreq 66%

answer

  1. the baseline moves with the data
  2. measures change, not distance travelled
  3. the ruler stretches to fit
  4. step changes still break it
  5. frozen instead fires on seasonality

basics

~20 s

A rolling reference re-estimates normal from data that already contains the drift, so each week's comparison is against last week's already-moved distribution. Gradual movement never accumulates in the statistic, though a step change still trips it.

solid answer

~50 s

The rolling window adopts the drift it was built to catch. Each run compares the detection window against a reference cut from recent traffic, which has already absorbed the previous weeks' movement, so a shift of a percent or two a week keeps the statistic near zero while the distance from the training distribution grows all season. It still catches **step** changes — an overnight sensor swap, a units change — because those break the recent past too. The opposite policy, a reference frozen at the training snapshot, does accumulate gradual movement, but on a seasonal system it eventually fires on legitimate phenology. The usable design is a frozen reference **aligned by calendar position** (week 27 against week 27 of the training seasons), with a short rolling reference alongside as a separate step-change test carrying its own threshold.

code

pseudocode · 17 lines
pseudocode
every week:
    live = model_inputs_scored_since(now - 7 days)

    frozen  = reference_for(serving_model_version, calendar_week(now))
    rolling = model_inputs_scored_between(now - 5 weeks, now - 1 week)

    for f in model_input_features:
        d_train  = psi(live[f], frozen[f])
        d_recent = psi(live[f], rolling[f])

        if d_train > threshold_train[f]:
            report('drifted from training', f, d_train)
        if d_recent > threshold_step[f]:
            report('step change', f, d_recent)

# gradual drift: d_recent stays near zero every week, d_train climbs
# overnight sensor swap: both cross on the first run after the swap

go deeper

for a junior

Recall that the reference window is chosen, not given, and that a baseline recomputed from recent traffic will report calm while the inputs slide away from training.

for a middle

Explain the mechanism: a rolling baseline re-estimates normal at the same rate the world moves, so it measures change between runs rather than distance from training.

for a senior

Show the design you would actually ship: a seasonally aligned frozen reference for displacement plus a short rolling one for step changes, each with its own threshold and label.

for a principal

Weigh the cost of nuisance firings against blindness. A calm dashboard nobody trusts and a noisy one nobody reads fail the same way, and the window policy is where that balance is set.

## The reference window is a policy, not a detail A shift test compares a **detection window** (what production just scored) against a **reference window**. The detection window is given. The reference is chosen, and the two common choices fail in opposite directions: - **frozen** — cut from the rows the model was trained on and never moved; - **rolling** — recomputed each run from the last N weeks of production traffic. For a national crop-yield forecast scored per field each week from overhead imagery and weather, the difference decides which failures the team ever hears about. ## Why a rolling reference goes quiet on exactly the drift you fear Suppose an atmospheric-correction change rolls out gradually across the imagery fleet, or an instrument degrades, moving a band's distribution about two percent a week. With a four-week rolling reference, every run compares this week against a baseline that already contains the last four weeks of that movement. The measured distance stays small and flat. Meanwhile the **cumulative** distance from the distribution the model was fitted on grows without bound, and by harvest the model is scoring inputs unlike anything it saw in training — with a dashboard that has been calm all season. The mechanism is simple: the rolling window re-baselines the definition of normal at the same rate the world changes. It measures *acceleration*, not *displacement*. What it does catch, and catches well, is a **step** change: a sensor swapped overnight, a units change in a weather feed, one region's ingest silently switching source. Those break the recent past as violently as they break the training distribution, so a rolling reference sees them immediately — often before a frozen reference has produced its next scheduled run. That is why a rolling reference is a good *second* test, not a bad idea. ## Why a frozen reference alarms on the calendar A frozen training-time reference accumulates displacement, which is what you want, but on a seasonal system it has a standing nuisance problem. Spectral inputs for a crop move enormously between week 6 and week 22 of a growing season **by design** — that phenological signal is the model's whole basis for forecasting. A reference pooled across a whole training season compares an early-season detection window against a season-wide average and fires with nothing wrong. Fire it every spring and the test stops being read. ## The resolution: freeze, but align 1. **Capture the reference from the training rows**, keyed by model version, so the test answers "how far is the live input from what this model learned?" 2. **Cut it by calendar or phenological position**, so week 27 of the live season is compared against week 27 of the training seasons — a seasonal-naive alignment. Only non-seasonal movement then trips the test. 3. **Run a short rolling reference alongside**, with its own threshold, reported as a distinct test, so a step change is caught within a run or two instead of waiting for displacement to build. | reference policy | the question it answers | what it goes blind to | its nuisance source | |---|---|---|---| | frozen at training | how far the live input is from what the model learned | nothing gradual, displacement accumulates | seasonality, and a change in the population mix | | rolling recent window | did something break in the last few weeks | gradual drift, which it re-baselines away | ordinary week-to-week variation | | frozen, seasonally aligned | non-seasonal movement since training | a shift that mimics normal seasonal movement | a season unlike any training season | ## Consequences to state out loud in a design round - The two tests need **different thresholds** and different labels in the output; collapsing them into one number makes the result uninterpretable. - When the model is retrained, the frozen reference is **recaptured** from the new training rows and published with the new model version, and thresholds derived against the old baseline do not carry over unchanged. - A seasonally aligned reference needs enough past seasons to have a comparable week; in the first season you have a pooled baseline and should say so rather than pretend the alignment exists. - No reference policy tells you whether the shift *hurt* the forecast. It tells you the input moved and by how much; what that cost is a separate reading on a separate timeline. The interviewer is listening for one sentence: the reference window is a design choice with two distinct failure modes, and you know which one you chose and why.

  • Is there a case for running both a frozen and a rolling reference at once?
    Yes, and it is the usual answer. They ask different questions: the frozen one measures displacement from what the model learned, which is the retrain conversation; the rolling one measures whether something broke lately, which is the upstream-incident conversation. Give them separate thresholds and report them as separate tests, because one number blending both is uninterpretable.
  • What happens to the frozen reference when the model is retrained on the newest season?
    It is recaptured from the new training rows and published with the new model version, because that model's notion of normal is now a different distribution. Thresholds derived against the old baseline do not transfer unchanged, since the statistic's scale depends on both sides of the comparison; re-derive them and keep the old reference for reproducing past firings.
  • How do you stop a frozen reference from alarming every spring on legitimate phenology?
    Align the comparison by calendar or phenological position rather than pooling the training season: compare this week's detection window against the same week of the training seasons. Only non-seasonal movement then crosses the threshold. In a first season with no comparable week, use the pooled baseline and label the test as provisional rather than tuning the threshold until it goes quiet.

A rolling reference is a tape measure that restretches to whatever you last measured, so it always reads no change. A frozen one keeps its markings but reads the tide as a leak.

saying these in an interview costs you the question

  • Picks a rolling reference because it needs no stored baseline
  • Says a rolling reference detects nothing at all, missing step changes
  • Treats seasonal movement in the inputs as drift to be alarmed on
  • Reuses the training-time threshold after recapturing the reference window
  • Blends the frozen and rolling results into one drift number
  • Assumes a quiet drift statistic proves the forecast is still accurate
open as a page

Where should the alert threshold on a drift statistic like PSI come from, if not the conventional 0.1 and 0.25 bands?

level: seniorimportance: must knowfreq 57%

basics

~20 s

From a backtest on closed past seasons: replay the statistic over the same windows and pick the value that separated seasons where forecast error actually rose from seasons where it did not. The 0.1 and 0.25 bands are convention inherited from another domain.

open as a page

In a weekly input-drift check on a deployed crop-yield model, what must have been stored in advance?

level: juniorimportance: should knowfreq 50%

basics

~20 s

A reference: a sample of the training inputs or their per-feature summaries, produced by the same feature code as production and pinned to the model version. The live detection window is the only side production supplies for free.

open as a page

A crop-yield model is scored weekly from imagery with a five-day revisit, so how often should its input shift test run?

level: middleimportance: should knowfreq 43%

basics

~20 s

At the rate the inputs genuinely refresh and the team could act on a result, which here is weekly, triggered by the feature partition landing rather than by a clock. Testing daily re-tests the same imagery and multiplies correlated firings.

open as a page

Why can a national input drift test read flat while one agro-climatic region's features have clearly moved?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Pooling weights every region by its share of rows, so a region holding 4% of scored fields can move its whole distribution and contribute only a few hundredths to the national statistic — under any sensible threshold.

open as a page

A KS test over two million scored field-weeks flags a shifted input band at p<0.001, so why is that not yet an incident?

level: seniorimportance: should knowfreq 48%

basics

~20 s

At that row count almost any difference is significant, so the p-value reports that the movement is not sampling noise and says nothing about its size. Significance is not harm: read the effect size, then how much the model uses that input.

open as a page