Why can a national input drift test read flat while one agro-climatic region's features have clearly moved?
answer
- the aggregate is a weighted average
- a small share moves the total little
- green nationally, broken locally
- watch the mix, not only the values
- each axis multiplies the comparisons
basics
~20 sPooling weights every region by its share of rows, so a region holding 4% of scored fields can move its whole distribution and contribute only a few hundredths to the national statistic — under any sensible threshold.
solid answer
~50 sA pooled statistic is a weighted average over the population it was computed on. If one agro-climatic region holds about 4% of the 1.2 million fields scored each week, a sensor change that moves that region's inputs decisively shifts only 4% of the national mass; the pooled PSI lands in the low hundredths while the region's own statistic is an order of magnitude larger. The fix is to run the test **per slice**, on the few axes along which the data-generating process genuinely differs — region, imagery platform, crop — with a minimum-row guard so thin slices report insufficient rows rather than a jumpy number. Watch the converse too: onboarding many new fields in one region changes the **mix** and moves the national statistic with no region having shifted at all, so test the mix itself separately.
code
pseudocode · 13 linesMIN_ROWS = 2000
report('national', psi(live.all, reference.all))
for slice in slices_where_process_differs: # region, platform, crop
rows = live.where(slice)
if count(rows) < MIN_ROWS:
report(slice, 'insufficient rows', count(rows))
continue
report(slice, psi(rows, reference.where(slice)))
# did the population itself get reweighted, with no slice moving?
report('mix', psi(row_share_by_slice(live), row_share_by_slice(reference)))go deeper
Recall that a statistic computed over everything is a weighted average, so a small group can change completely without moving the overall number much.
Explain the arithmetic of dilution with a share and a magnitude, and why the same slicing that reveals it also multiplies the run's comparison count.
Show the full design: slice on axes where the process differs, guard thin slices, test the composition separately, and rank slices by magnitude in the output.
Choose how many axes the organisation will monitor, since each one buys detection in a concentrated failure and costs attention across every quiet run.
## Pooling is a weighted average, and the weight is the population share A drift statistic computed over all scored rows compares two pooled distributions. A subpopulation's movement enters that comparison in proportion to its share of the rows, which is why aggregate monitoring is systematically blind to concentrated failures. Take a national crop-yield forecast scoring roughly 1.2 million fields each week across eight agro-climatic regions. One region holds about 48,000 fields, near 4% of the total. Its imagery source changes and its values for one band move decisively — enough that the region's own PSI against its baseline lands around 0.6. Nationally, that is 4% of the mass moving out of one bin and into another. Working the PSI contribution through for bins that held roughly a tenth of the population each, a 4-point-of-share move contributes on the order of **0.03** — below any watch band, invisible on the national panel, and stable enough that nobody looks twice. The region has a broken input feed and the dashboard is green. ## Slice on the axes where the process genuinely differs The cure is to cut the test, but not to cut it everywhere. Each slicing axis multiplies the number of comparisons per run, so the axes have to earn their place. Cut on the dimensions along which the **data-generating process** differs: - **region or agro-climatic zone** — different weather, phenology, soils, and often different ingest paths; - **imagery platform or sensor generation** — the most common source of a step change confined to a subset; - **crop type** — different phenology and different feature behaviour entirely; - occasionally **irrigation regime or field size band**, where the model behaves differently across them. Do not cut on every categorical column available. Forty inputs across eight regions is already 320 comparisons per run; adding two more axes multiplies that into a panel nobody can read, and the false-firing count grows with it. ## The converse trap: the mix moves, no slice does Slicing also protects against the opposite error. Suppose 120,000 newly onboarded fields appear in one region, taking the scored population to about 1.32 million. No region's own input distribution has changed at all — but the national population is now weighted differently, roughly nine points of share have moved between regions whose distributions genuinely differ, and the pooled statistic jumps by about as much as the real regional break did. That is a **mix shift**, and responding to it as if the inputs had drifted sends the team looking for an upstream fault that does not exist. Two defences: 1. test the **composition itself** — compare the share of rows by region between the reference and the detection window, as its own reported number; 2. read the per-slice results before the pooled one, since a mix shift shows as a moved national statistic with every slice quiet. ## Guard the thin slices Slicing creates small populations, and a statistic on a few hundred rows swings enough to fire on its own: - set a **minimum-row guard** and report "insufficient rows" rather than a number; - give a small but important slice a **longer detection window** — four weeks instead of one — and accept slower detection there; - where the test does run on a thin slice, this is the one place a significance test genuinely helps, because sampling variation is the live alternative explanation. | what you compute | what it catches | what it misses | |---|---|---| | one pooled national statistic | broad movement affecting most of the population | a concentrated break in a small slice | | per-slice statistics | a break confined to one region, platform or crop | nothing, but it multiplies comparisons and firings | | a composition test on the mix | a reweighted population masquerading as drift | movement inside a slice whose share stayed constant | All three are cheap. The design decision is which axes to cut and how to present the result so a reader sees the concentrated case without wading through hundreds of cells — usually by ranking slices by magnitude and showing the pooled number as context rather than as the headline.
- How many slicing axes should the test cut on?Few, and only those along which the data-generating process genuinely differs — region, imagery platform, crop. Each axis multiplies comparisons per run: forty inputs by eight regions is already 320, and two more axes turn the panel into a wall of cells with a matching rise in nuisance firings. Slicing is a targeting decision, not a default to apply to every categorical column.
- The national statistic jumped but every region's own statistic is quiet. What happened?The population mix changed rather than the inputs. Newly onboarded fields concentrated in one region reweight the pooled comparison, and regions whose distributions legitimately differ now contribute in different proportions. Confirm it by testing the share of rows per slice between the reference and detection windows; a moved composition with quiet slices is a reporting fact, not an upstream fault.
saying these in an interview costs you the question
- Monitors only the pooled national statistic and calls concentrated breaks rare
- Slices on every available categorical column without counting the comparisons
- Reports a drift statistic on a slice holding a couple of hundred rows
- Reads a mix shift as evidence that the inputs drifted
- Compares a slice's live window against the pooled national baseline