skip to content

A KS test over two million scored field-weeks flags a shifted input band at p<0.001, so why is that not yet an incident?

level: seniorimportance: should knowfreq 48%

answer

  1. significance scales with row count
  2. big enough to see, small enough to ignore
  3. magnitude first, not the p-value
  4. weight the feature by model influence
  5. real, large and still harmless

basics

~20 s

At that row count almost any difference is significant, so the p-value reports that the movement is not sampling noise and says nothing about its size. Significance is not harm: read the effect size, then how much the model uses that input.

solid answer

~40 s

A p-value shrinks with sample size, so on two million rows a displacement far too small to matter clears any conventional cutoff. Triage in three readings. First the **effect size** — the magnitude of PSI or a distance statistic, which does not inflate with rows. Second, **how much the model leans on that input**, a weight you already know from the training run rather than one you derive during an incident: the same displacement in a barely-used band and in a dominant one are different events. Third, whether it **propagated** at all to the model's output. A small movement in a lightly-weighted band with no visible propagation is a note with an owner, not an incident. This is why drift tests are thresholded on magnitude rather than on significance.

go deeper

for a junior

Recall that a p-value gets smaller as the number of rows grows, so on very large windows significance stops distinguishing a big shift from a trivial one.

for a middle

Explain why drift thresholds are set on effect size rather than on significance, and how influence on the model's output ranks which firings are worth opening.

for a senior

Show a triage order someone else can execute, and the discipline of recording a firing as real-and-harmless rather than silently closing it.

for a principal

Decide what the organisation treats as harm, since significance, magnitude and propagation are all available and none of them is the business consequence.

## A p-value is a statement about row count as much as about the shift A two-sample test asks whether the difference between two populations is larger than sampling noise would explain. Its answer depends on both the size of the difference and the number of rows. With two million scored field-weeks in a detection window, the noise floor is tiny, so a displacement of no operational consequence clears p < 0.001 comfortably. Run that test on every input every week and it will report significance almost permanently. This is why drift monitoring is thresholded on **effect size** — the magnitude of PSI, a distance between distributions — rather than on a significance verdict. Magnitude does not inflate as rows accumulate, so the same number means the same thing in a large region and a small one. ## Three readings, in order 1. **How big is it?** Read the magnitude against the threshold derived for that feature. A statistically certain movement of negligible size is the common case, not the exception. 2. **Does the model care?** Rank the shifted input by the influence it has on the model's output — a number that comes from the training run and sits in the model's documentation, not something to work out during triage. A band the model barely weights can move a long way with little consequence; a dominant input moving slightly can matter more. 3. **Did it reach the output?** If the input moved materially and the model's own output distribution over the same fields is unchanged, either the model is insensitive there or the movement cancelled against something else. If the output moved too, the case is stronger and the investigation has a direction. | reading | what it answers | what it cannot tell you | |---|---|---| | significance | is the movement bigger than sampling noise | anything about size, because it scales with rows | | effect size | how far the distribution actually moved | whether that distance matters to this model | | model influence | how much this input drives the output | whether the movement is benign in cause | | output propagation | did the shift reach the prediction at all | whether the prediction is now wrong | Notice what the last column keeps saying. None of these readings establishes that the forecast got worse. That question needs outcomes, and on a growing season those arrive at harvest, on a different timeline from this week's test. ## A shift can be real, large and still harmless Several perfectly real movements deserve a record rather than a response: - a band recalibrated upstream in a way the feature code normalises before the model sees it; - a legitimate change in conditions the model was trained to handle, within the range it saw; - a movement confined to a feature that a later model version stopped using but the pipeline still computes. Calling each of those an incident spends the same attention as a real break, and after a few of them nobody opens the next one. ## And a shift in an unused input can still be the visible edge of something The inverse trap is dismissing every low-influence firing. Upstream breaks rarely arrive one feature at a time: an imagery platform swap, a units change, a reprojection, a change of atmospheric correction usually touches a family of derived inputs at once, and the first one to cross a threshold may be the least important member of that family. So triage the **cause**, not only the feature. "Which inputs share the upstream source that moved?" is a better second question than "how much does the model weight this one?" ## What a design round is testing here The interviewer wants to hear the distinction between **statistically significant** and **operationally harmful** stated cleanly, and then a triage order that is executable at three in the afternoon by someone who did not build the model. Weak answers treat a small p-value as a verdict, or swing the other way and argue that drift tests are useless because everything is significant. The workable position sits between: thresholds on magnitude, ranked by influence, checked for propagation, with a written outcome for every firing — including "real, and fine", which is a result worth recording rather than a dismissal.

  • The effect size is large but the model barely uses that input. Is that closable?
    Close it as a record with a named owner rather than as noise. Upstream changes arrive in families — a platform swap or a correction change usually moves several derived inputs at once — so the useful next question is which other inputs share the source that moved, not only how much the model weights this one. A large movement with a known cause and no propagation is a result, not a dismissal.
  • If magnitude is what you threshold on, why compute a significance test at all?
    It earns its place on thin slices. In a small region with a few thousand rows, a large-looking magnitude may be ordinary sampling variation, and the test is what separates the two. On national windows of hundreds of thousands of rows it saturates and adds nothing, so the sensible design applies it where rows are scarce and thresholds on magnitude everywhere else.

A scale sensitive enough to weigh a signature will report that every envelope in the mailroom differs. It is telling the truth; it is just not telling you which parcel to open.

saying these in an interview costs you the question

  • Treats a small p-value as proof the shift matters operationally
  • Concludes drift tests are useless because everything is significant at scale
  • Assumes a shifted input means the forecast is already wrong
  • Ignores every firing on a low-influence feature without checking the cause
  • Derives feature influence during triage instead of taking it from the training run