skip to content

In a weekly input-drift check on a deployed crop-yield model, what must have been stored in advance?

level: juniorimportance: should knowfreq 50%

answer

  1. a comparison needs two sides
  2. only one side arrives free
  3. one side is captured, not observed
  4. same feature path on both
  5. pinned to the model version

basics

~20 s

A reference: a sample of the training inputs or their per-feature summaries, produced by the same feature code as production and pinned to the model version. The live detection window is the only side production supplies for free.

solid answer

~40 s

A shift test compares two populations, and only one of them shows up by itself. The **detection window** is whatever the serving path scored in the last week. The **reference window** has to have been captured earlier and persisted on purpose, from the rows the model was actually fitted on. It is stored either as a retained sample or as per-feature summaries (fixed quantile bin edges plus counts, category counts), and it must be produced by the same feature-computation path as serving, otherwise implementation differences read as drift. Because a different model version was trained on a different distribution, the reference belongs to a model version and is resolved by it, so a model-version rollback brings back the matching reference instead of judging the restored model against the newer one.

go deeper

for a junior

Recall that a shift test compares a live window against a stored one, and that the stored one has to be captured by the training pipeline and kept with the model version.

for a middle

Explain why summaries versus a retained sample is a real choice, and why building the baseline with a different feature implementation manufactures drift that is not there.

for a senior

Show how the baseline is versioned and resolved at run time, so a model-version rollback restores its reference too and past firings stay reproducible.

for a principal

Weigh the retention, privacy and storage cost of keeping row samples for every model version against the ability to ask a new question of a past season without recapturing it.

## The test has two sides, and only one of them arrives by itself A shift test on a live model answers one question: **do the inputs the model is scoring now look like the inputs it was fitted on?** That comparison needs two populations. One of them, the **detection window** (the rows scored in the last day or week), is produced by the serving path and costs nothing to obtain. The other, the **reference window** (the *baseline*), does not exist in production at all. It has to have been captured earlier and stored on purpose, or the job has nothing to compare against and produces no verdict. Everything else in this subject follows from that asymmetry. The reference is an artifact the pipeline produces and versions, exactly like a model file or a feature definition. ## What the stored reference has to contain For a national crop-yield forecast scored per field each week from multispectral overhead imagery and weather, the reference is cut from the same rows the training job consumed: 1. **The features the model actually sees** — the assembled model inputs (band indices, weather aggregates, static soil attributes), not the raw imagery tiles upstream of them. A movement in raw pixels that the feature code normalises away is not a movement in the model's input. 2. **A form the statistic can consume** — a retained sample of rows, or per-feature summaries: fixed quantile bin edges and bin counts for a numeric feature, category counts for a categorical one. 3. **The metadata describing the cut** — which seasons, which calendar weeks, how many fields, which regions. Without it nobody can later say what *normal* meant. | stored form | what it supports | what it costs | |---|---|---| | per-feature bin counts | PSI, chi-squared, a KS check on the binned form | very little, but the bins are fixed at capture time | | a retained row sample | any statistic added later, plus per-slice and multivariate cuts | storage, and a retention decision | Bins are enough for the checks you know you want today. A sample is what lets you ask a new question next season instead of waiting a year to recapture one. Many systems keep both: summaries for the routine job, a modest stratified sample for investigation. ## The same feature code on both sides The reference must come out of the **same feature-computation path that serving uses**. If the baseline was built by one implementation and the live window is read from another, every difference between the two implementations arrives as drift that is not there — and worse, real drift then hides inside noise the team has learned to ignore. That is why the reference is normally captured **by the training pipeline itself**, from the exact rows the model was fitted on, rather than reconstructed months later by a monitoring job running its own query. ## Pinning the reference to a model version A reference belongs to a model version, not to a service. When a champion trained on an older set of seasons is replaced by one trained on newer data, the new model's notion of normal is a different distribution; judging it against the old baseline reports drift from the first minute it serves. So: - publish the reference **inside the released model bundle**, or beside it in the model registry keyed by model version; - have the job resolve the reference for the model version currently serving, rather than reading a global default; - **never overwrite a reference in place** — a model-version rollback has to bring its baseline back with it, and last month's firing has to stay reproducible. ## What goes wrong when this is loose - The job quietly uses recent production traffic as its own reference and reports calm forever. - Two model versions serve different regions against one shared baseline, so one of them looks permanently drifted. - The baseline is refreshed at every retrain with no record of which one produced which firing, and "was this band already moving in the spring?" becomes unanswerable. - The baseline is cut **after** the training data cut, so it describes a population the model never saw. A junior engineer is not expected to choose a window policy or a threshold. They are expected to know that the comparison has a stored side, that it is versioned with the model, and that both sides come out of the same feature code.

  • Does the stored reference have to be raw rows, or are summaries enough?
    Summaries are enough for the checks you already run: fixed quantile bin edges with counts support PSI and chi-squared, and a binned KS. A retained sample costs storage but buys optionality — adding a new statistic, cutting by a slice you had not thought of, or running a multivariate check later, without waiting a full season to recapture a baseline.
  • Where should the reference live so the job always uses the right one?
    Treat it as part of the released model bundle, or store it in the model registry keyed by model version, and have the drift job resolve it from the version currently serving. That way a model-version rollback restores the matching baseline automatically, and no run can silently compare a restored model against a newer model's notion of normal.

saying these in an interview costs you the question

  • Thinks a drift test can run on live data with nothing stored
  • Builds the reference with a different feature implementation than serving uses
  • Uses one global baseline while several model versions serve traffic
  • Overwrites the stored reference in place, so old firings cannot be reproduced
  • Takes the baseline from raw upstream data rather than assembled model inputs