A crop-yield model is scored weekly from imagery with a five-day revisit, so how often should its input shift test run?
answer
- no new observation, no new evidence
- match the scoring cycle
- gate on data arrival
- smallest slice sets the window
- comparisons equal features times slices
basics
~20 sAt the rate the inputs genuinely refresh and the team could act on a result, which here is weekly, triggered by the feature partition landing rather than by a clock. Testing daily re-tests the same imagery and multiplies correlated firings.
solid answer
~50 sTwo things set the cadence: how fast the inputs actually change, and how fast anyone could respond. With a five-day revisit and weekly scoring, a daily run compares largely the same pixels against the same baseline, so it adds correlated firings and no earlier warning. Run the test **once per scoring cycle**, and trigger it on the feature partition being complete rather than on a wall-clock time, or a run will read a half-written partition and report a shift that is really a missing region. Window size follows from cadence: the detection window must hold enough rows for the **smallest slice** you intend to judge, not for the national total, so a small region may need a longer window or a minimum-row guard. Longer windows give a steadier statistic and slower detection; that is the trade being made.
go deeper
Recall that a shift test needs new data to have arrived: on inputs that refresh every few days, testing hourly re-examines the same observations.
Explain the two constraints that set cadence, the revisit rate of the inputs and the response time, and why window size is driven by the smallest slice you test.
Show the operational guards: trigger on partition completeness, skip and record rather than testing partial data, and keep detection and reference windows disjoint.
Frame cadence as a budget: comparisons per run times runs per week is the firing volume the team will live with, and it is set before any threshold is chosen.
## Cadence follows the data and the response, not the calendar Two questions set how often a shift test runs: 1. **How fast can the input actually change?** A national crop-yield forecast draws on overhead imagery with roughly a five-day revisit and on daily weather aggregates. Half the feature vector cannot move between Monday and Tuesday because no new observation exists. 2. **How fast could anyone act on the answer?** If the response to a confirmed input shift is to investigate an upstream feed and, at most, to schedule a retrain, a result produced six times between two scoring runs changes nothing that a single result would not. Running the test once per scoring cycle — weekly, here — matches both. Running it daily compares largely the same imagery against the same baseline six extra times, which does not detect anything earlier but does produce strings of correlated firings that read like six problems. ## Trigger on arrival, not on a clock A drift job scheduled at a fixed hour will eventually run while the feature pipeline is still writing. Half the regions are present, the national input distribution looks nothing like the baseline, and the test reports a large shift that is really a partially-written partition. The fix is operational, not statistical: - gate the run on a **completeness signal** from the feature pipeline (expected partitions present, expected row count within tolerance); - if completeness fails, **skip and report skipped** rather than testing the rows that happen to be there; - record which data version the run consumed, so a later reader can tell a real shift from an incomplete one. ## Window size follows from the smallest slice you intend to judge Cadence sets the detection window's span; the span has to leave enough rows in every cut you plan to test. With roughly 1.2 million fields scored nationally each week, a national test has more rows than any statistic needs. But once the test is cut by agro-climatic region, a small region may contribute only a few thousand fields, and a statistic computed on a couple of hundred rows swings wildly from run to run and manufactures firings. - Set a **minimum-row guard** and report "insufficient rows" rather than a number. - Where a slice is genuinely small but genuinely matters, give **that slice** a longer detection window (four weeks instead of one) and accept slower detection there. - Do not let the detection window overlap the reference window; if it does, the two sides share rows and the statistic is biased toward calm. | cadence | detection speed | statistic stability | firing volume | |---|---|---|---| | daily, on weekly-refreshed inputs | no faster in practice | worse, fewer new rows per run | much higher, and correlated | | once per scoring cycle | matched to when the input can change | good on national and mid-size slices | proportionate | | monthly | too slow to act within a season | very stable | low, but a break can run for weeks | ## Cadence multiplies with the slice and feature count The number of comparisons a run produces is features times slices. With 40 model inputs cut by 8 regions, one run is **320 tests**. At a 5% per-test false-positive rate that is about **16 red cells every run** before anything has actually shifted. Move that from weekly to daily and the same 16 arrive every day, which is how a drift dashboard becomes permanently red and stops being read. Three levers, used together: 1. run at the cadence the data justifies; 2. control the false-discovery rate across the run rather than applying a per-test cutoff independently (a Benjamini-Hochberg procedure at a stated rate, or a Bonferroni cut of 0.05/320 = 0.00016 if you would rather be conservative); 3. threshold on **effect size** — the magnitude of PSI or a distance — rather than on a p-value, since with hundreds of thousands of rows per slice almost any difference is significant. ## What this sounds like in a design round Good answers name the constraint before the number: the revisit interval, the scoring cycle, the response time, the smallest slice. Weak answers pick a cadence because it is a round number, or reach for the fastest possible schedule on the theory that more monitoring is more safety. On a weekly scoring pipeline it is not: it is the same evidence, resampled, with a higher chance that whoever is reading has stopped looking.
- Cutting 40 features by 8 regions gives 320 comparisons per run — what does that do to alert volume?At a 5% per-test false-positive rate it produces about 16 firings per run with nothing actually wrong, and running daily delivers those 16 every day. Control the false-discovery rate across the run rather than testing each cell independently, threshold on effect size rather than on a p-value, and rank surviving firings by how much the model uses the feature before anyone opens them.
- Should the job run on a fixed schedule or on the feature data landing?On the data landing, gated by a completeness check. A clock-triggered run can read a partition that is still being written, see a national distribution missing whole regions, and report a large shift that is really an incomplete input. If completeness fails, skip the run and record it as skipped rather than testing whatever rows happen to be present.
saying these in an interview costs you the question
- Runs the test hourly on inputs that refresh every five days
- Sizes the detection window from the national row count only
- Lets the detection window overlap the stored reference window
- Starts the job on a clock without checking the input partition is complete
- Reports a statistic on a slice holding a couple of hundred rows