Where should the alert threshold on a drift statistic like PSI come from, if not the conventional 0.1 and 0.25 bands?
answer
- a convention is not a measurement
- borrowed from another domain
- anchor it to something that mattered
- replay closed seasons, sweep the value
- per feature, and re-derived after recapture
basics
~20 sFrom a backtest on closed past seasons: replay the statistic over the same windows and pick the value that separated seasons where forecast error actually rose from seasons where it did not. The 0.1 and 0.25 bands are convention inherited from another domain.
solid answer
~40 sTreat the familiar bands as a starting prior, not a finding — they are conventional cutoffs from population-stability work in an unrelated domain and know nothing about overhead imagery. The defensible threshold is derived: take past seasons that are already closed, recompute the statistic on the same reference and detection windows the live job would have used, and choose the value that fired in the seasons whose matured forecast error rose and stayed quiet in the ones where it did not. Publish the hit and nuisance count that value buys, because a threshold is a purchase. Set it **per feature** — bands with wide natural variability and bands the model leans on deserve different cutoffs — and re-derive it whenever the reference window is recaptured, since the statistic's scale depends on both sides.
go deeper
Recall that the conventional PSI bands are a rule of thumb borrowed from another domain, not a value computed from your data, and should be labelled as provisional.
Explain how a threshold is derived by replaying closed seasons and sweeping the cutoff, and why it belongs per feature rather than once per model.
Show that you would publish what the threshold costs in nuisance firings, and re-derive it whenever the reference window or binning changes.
Own the trade the number encodes: how many disturbed seasons the organisation will accept to catch the one that hurt, and who is allowed to move it.
## Where the familiar numbers actually come from Ask where a 0.1 or 0.25 PSI cutoff came from and the honest answer is usually "it was in the first article I read." Those bands are a **convention** from population-stability practice in scorecard work: below roughly 0.1 treat the population as stable, above roughly 0.25 treat it as materially changed. They are not critical values derived from a null distribution, they carry no sample size, and they have never seen a multispectral band over a growing season. Used as a starting point and labelled as provisional, they are fine. Presented as a finding, they are a fabricated number in an operations document. ## Deriving a threshold from closed seasons The material for a real threshold is history in which the outcome is already in the ground. For a national crop-yield forecast with six completed seasons of stored inputs and matured yield error: 1. **Replay** each past season week by week, computing the same statistic over the same reference and detection windows the live job would have used. No peeking forward: each week's statistic uses only what would have been available that week. 2. **Label** the seasons: in which ones did the matured forecast error actually rise beyond the tolerance the forecast promises? 3. **Sweep** a candidate threshold across its range and count, for each value, how many of the harmful seasons it would have flagged and how many quiet seasons it would have disturbed. 4. **Choose** the value that separates them best, and write down what that choice costs — for example: "fires in both seasons where error rose, plus three nuisance firings across the other four." 5. **Keep the sweep**, because next season adds a labelled point and the curve can be redrawn rather than re-argued. The key property is that the threshold is now anchored to something that mattered, rather than to a number's reputation. ## Per feature, not one number for the model A single global cutoff treats all inputs as equally volatile and equally important, and they are not: - a band whose values swing widely between years under normal weather needs a **higher** cutoff, or it will fire every season; - a stable input — a static soil attribute, a well-behaved derived index — can carry a **tighter** cutoff, since any real movement in it is suspicious; - an input the model leans on heavily deserves a tighter cutoff than one it barely uses, because the same displacement has a different consequence. | threshold source | what it is anchored to | what it is worth | |---|---|---| | published convention | practice in another domain | a labelled starting point, nothing more | | the statistic's own null distribution | sampling noise at your row count | with hundreds of thousands of rows, nearly everything clears it | | backtest against closed seasons | seasons where forecast error truly rose | defensible, and improves with every season | | the quietest value that stops complaints | the team's tolerance for noise | guarantees only that the test is silent | That last row is the trap worth naming: tuning a threshold upward until the dashboard goes green is how a drift test is decommissioned without anyone deciding to decommission it. ## When you have no history A first season has no labelled precedent, and inventing one is worse than admitting it. Start from convention, **mark the threshold provisional in the output itself**, and make every firing produce an adjudication — real upstream change, seasonal artifact, or noise — recorded against the statistic's value at the time. One season of that is the dataset the sweep needs. In the meantime, prefer a wider detection window so the statistic is stable, and prefer effect-size magnitude over any p-value, which at these row counts saturates immediately. ## Re-derive when either side of the comparison changes The statistic measures a distance between two populations, so its scale depends on both. When the model is retrained and the reference is recaptured, thresholds derived against the old baseline no longer mean what they meant. The same applies when bin edges are re-cut, when a slice's definition changes, or when the detection window's span changes. Carry the threshold's provenance — which reference version, which windows, which backtest — next to the number, or the next engineer inherits a cutoff nobody can defend and will do the only thing available to them: raise it until the noise stops.
- What should be published alongside the chosen threshold?Its provenance and its price: which reference version and window spans the backtest used, how many of the harmful seasons that value would have flagged, and how many quiet seasons it would have disturbed. A threshold is a purchase of detections at a cost in nuisance firings, and a number without that record cannot be defended or revised by whoever inherits it.
- There is no history at all — this is the model's first season. What then?Start from the convention, mark it provisional in the output, and widen the detection window so the statistic is stable. Require every firing to be adjudicated and recorded against the statistic's value: real upstream change, seasonal artifact, or noise. One season of those adjudications is exactly the labelled dataset the sweep needs next year.
- Does a threshold survive the reference window being recaptured at retrain?No. The statistic measures a distance between two populations, so recapturing the baseline changes its scale, and a cutoff tuned against the old reference means something different against the new one. Re-run the sweep after recapture, and keep the old reference and its threshold so past firings remain reproducible rather than being reinterpreted under the new number.
saying these in an interview costs you the question
- Cites the 0.1 and 0.25 bands as a derived statistical result
- Uses a p-value as the cutoff on windows of hundreds of thousands of rows
- Applies one global threshold across every model input
- Raises the threshold until the dashboard stops firing
- Keeps the old cutoff after the reference window is recaptured
- Tunes the threshold on the current season's still-open outcome