skip to content

A weather-feature shift test trips daily while served forecast error is unchanged - should that test page the on-call?

level: middleimportance: must knowfreq 57%

answer

  1. inputs moved is not harm landed
  2. one pages, the other explains
  3. seasons move features every year
  4. many features, many chance trips
  5. dead feed is the narrow exception

basics

~20 s

No. The pager belongs to the served forecast error, the effect the operator feels. A feature whose distribution moved without moving error is a diagnostic: attach it to the quality page as context, and let it file a ticket at most.

solid answer

~40 s

A shift test firing means an input distribution moved; a quality alarm firing means the published forecast got worse. Only the second one has a cost, and only the second one has a response. Weather-driven load features move every season, so a feature-level test on a forecaster is expected to trip routinely, and with dozens of features tested daily some will trip by chance alone. Page on the served error, and carry the recent shift results **on that page** so the responder opens it with a ranked first hypothesis. The shift test earns a ticket when it predicts harm that has not landed yet, and it earns a page only in one narrow case: when the input path is provably dead, since served error cannot reveal that until actuals mature hours later.

go deeper

for a junior

Hold on to the distinction: a test saying the weather data changed is not the same as the forecast getting worse, and only the second one costs the operator anything.

for a middle

Explain why a weather-driven forecaster trips feature-level tests routinely, and why testing many features every day produces trips even when nothing is wrong.

for a senior

Show the wiring: quality on the pager, shift results attached as ranked context, and one narrow freshness rule that pages because label latency makes the quality alarm blind.

for a principal

The judgment call is how much blindness you accept between a failure and the first matured actual, and how many cause-side rules you will fund to shrink it.

## Two different firings, two different meanings This category contains two things that both get called "an alarm firing", and conflating them is the classic design-round error: - a **shift test tripping** - a statistic comparing a feature's recent distribution against a reference says the distribution moved; - a **model-quality alarm paging** - the published forecast's error against settled actuals crossed the rule that was agreed to be worth a human. The first is a statement about inputs. The second is a statement about harm. A load forecaster's inputs are temperature, humidity, calendar effects and recent load, all of which move constantly and legitimately: the distribution of temperature in July is not the distribution in March, and the model was built knowing that. Covariate shift with unchanged error is the normal operating state of this system, not an incident. ## Why a feature-level test is a poor pager source - **No coupling to harm.** The test does not know the model's sensitivity to that feature. A feature the model barely uses can move dramatically and cost nothing. - **Expected seasonal movement.** On a weather-driven forecaster, tripping is the default for months of the year, so the page rate tracks the calendar rather than the system's health. - **Multiplicity.** Dozens of features tested every day produce trips by chance at whatever rate the threshold implies, before any real shift exists. The page rate scales with the number of features, which is a property of the feature set, not of the forecast. - **No response.** "Temperature moved" does not name an action. "Intraday error is twice the baseline's for 90 minutes" names one. - **Statistical significance is not operational harm.** With enough intervals a trivially small distributional move is significant. The question the operator cares about is whether the megawatts changed. ## What the shift test is for instead It is the best enrichment a quality page can carry. When served error does cross, the responder wants to know within seconds whether an input moved at the same time, because that is the branch point between "the model has gone stale" and "something upstream broke". Wiring the recent test results into the page payload converts a noisy pager into a useful diagnostic. | signal | what firing means | tier | |---|---|---| | served megawatt error against matured actuals | the operator is paying now | page | | a feature's distribution moved, error flat | an input changed, harm unproven | context on the quality page; ticket if it persists | | a feature feed has stopped or frozen at a constant | harm is certain and not yet measurable | page, on an unambiguous rule | | a shift test that has tripped every summer for three years | the reference window is wrong for this feature | backlog item, not an alert | ## The one case where a cause signal is worth a page There is a real exception on an ML system, and it comes from label latency rather than from alerting philosophy. Served error is only computable after actuals settle, so between the failure and the maturation of the first affected interval the quality alarm is structurally blind. If an input feed has stopped - the same value repeating, or the feature missing entirely - the damage is certain, already being published, and invisible to the quality alarm for another hour or more. That justifies a page on the cause, and it should be written as a flat unambiguous rule ("no new value for N cycles") rather than as a distribution statistic, so the responder never has to interpret it at 03:00. That exception is narrow on purpose. It covers *the feed is dead*, not *the feed looks different*. ## The failure mode this prevents Teams that page per drifting feature end up in one of two places. Either the on-call learns to ignore the pager, and the real quality page is lost in the noise; or the team starts silencing the shift tests one by one, which throws away the diagnostic value that made them worth computing. Keeping the tests loud on a dashboard and silent on the pager keeps both properties: the responder still gets them, but only after something that costs money has fired.

  • Served error only becomes computable after actuals settle. When is a cause-side alarm worth a page then?
    When the input path is provably dead rather than merely different: a feed that has stopped, or a feature pinned at one repeated value. The forecast being published now is already wrong, and the quality alarm cannot show it until actuals mature. Write that rule as a flat freshness condition, not as a distribution statistic, so it is unambiguous at night and cannot fire on a seasonal move.
  • The shift test is overwhelmingly significant on a large sample but megawatt error is flat. What does that tell you?
    That there is strong evidence the distribution moved, and no evidence it hurts. Significance grows with sample size, so on months of half-hourly intervals even a negligible move clears any fixed significance bar. The operational question is effect on error, so keep the alarm on harm and treat the test's effect size as the thing worth reading, not its p-value.
  • Should the drift context on the page be the tests' verdicts or the raw statistics?
    Both, but ranked: the responder needs to see at a glance which inputs moved most and when they started, not a wall of pass/fail flags. A short ordered list of the largest recent moves with their onset times turns the page into a hypothesis list. The thresholds those verdicts came from belong to the test's own design, and they are not what the responder should be re-deriving at night.

saying these in an interview costs you the question

  • Treats any statistically significant input shift as model degradation
  • Pages one alert per drifting feature, so summer wakes the team
  • Assumes stable input distributions prove the model is still accurate
  • Waits for a shift test before believing a served-error page
  • Silences the shift tests entirely once they prove noisy
  • Sets one shift threshold for every feature regardless of its influence