A three-day heatwave doubles intraday load forecast error and pages the on-call nightly - how should the alarm be redefined?
answer
- hard days are hard for everyone
- run a control on the same intervals
- ratio, not absolute level
- keep one unusable-regardless ceiling
- burn-in counted in matured intervals
basics
~20 sAlarm on skill against a seasonal-naive baseline scored on the same intervals rather than on an absolute error level. A heatwave degrades the baseline too, so the ratio moves little, while a genuine model failure pushes it up.
solid answer
~40 sRun a non-learned baseline - load at the same half-hour one week earlier - permanently in parallel and score it on exactly the intervals the model is scored on. The alarm metric becomes the ratio of the model's error to the baseline's. On a hard day both rise together and the ratio barely moves; when the model itself degrades, the ratio climbs toward 1.0, meaning the model has stopped beating a rule anyone could write. Damp it with a smoothed error series, a burn-in of several consecutive matured settlement intervals, and a clear threshold lower than the fire threshold. Keep one absolute megawatt ceiling as a separate, higher severity tier, because the ratio is blind to the case where the model and the baseline are both unusable.
code
pseudocode · 21 lines// evaluated once per matured settlement interval
e_model = mean_abs_error(published_forecast[t], actual[t])
e_base = mean_abs_error(seasonal_naive[t], actual[t]) // same half-hour, one week earlier
ratio = ewma(e_model / max(e_base, floor_mw), alpha = 0.2) // half-life ~3 intervals
if ratio >= 0.90:
breaches = breaches + 1
else:
breaches = 0
if breaches >= 3 and not firing: // 3 half-hourly intervals = 90 minutes
firing = true
page(horizon = "intraday", ratio = ratio, model = served_model_version)
if firing and ratio <= 0.75: // clears lower than it fires
firing = false
resolve()
if e_model >= unusable_ceiling_mw: // separate tier, ratio-independent
page(severity = "critical", reason = "absolute megawatt ceiling")go deeper
The idea to keep: compare the model against a simple rule running at the same time, so a day that was hard for everyone does not look like a broken model.
Be able to define a seasonal-naive control for load, explain why the ratio moves little on a heatwave, and describe smoothing, burn-in and hysteresis.
Demonstrate the blind spot and cover it with an absolute ceiling tier, and count burn-in in matured intervals because the metric advances with data, not with the clock.
The tradeoff to own is detection delay against page precision: every interval of burn-in is exposure the operator is buying to avoid a false page, and that price should be chosen deliberately.
## Why an absolute threshold pages on hard days A fixed rule such as "page if intraday error exceeds 2%" encodes an assumption that the forecasting problem has constant difficulty. It does not. Extreme weather, holidays, a large industrial outage and a public event all make load genuinely harder to predict, and every forecaster on earth does worse on those days. The alarm then fires for a reason no responder can act on: the model is fine, the world is hard, and the only available levers make things worse. Teams usually respond by ratcheting the threshold up after each such night, which ends with an alarm that no longer fires for real failures either, or by silencing the summer, which is a page-shaped hole in the monitoring. ## The baseline as a live control The fix is to run a control that faces the same conditions. For load, the standard non-learned control is a **seasonal-naive (persistence) baseline**: the actual load at the same half-hour one week earlier, which carries both the daily and the weekly shape for free. It is trivial to compute, never needs retraining, and reads nothing from the model's feature path. Score it on exactly the intervals the model is scored on, then alarm on the ratio: - **ratio well below 1** - the model is adding real skill, which is the normal state; - **ratio approaching 1** - the model has stopped beating a rule a spreadsheet could implement; - **ratio above 1** - the model is actively worse than doing nothing clever, which is unambiguous. On heatwave nights the baseline is badly wrong too, so the ratio stays near its usual value and nobody is woken. When a feature pipeline breaks or a bad model version is promoted, the baseline is unaffected and the ratio jumps - which is exactly the separation the absolute threshold could not make. ## The blind spot, and the tier that covers it A ratio hides the case where both terms degrade together: a corrupted actuals feed, or a shock so large that neither producer is usable. The ratio can look healthy while the published forecast is useless. So keep **one absolute megawatt ceiling** as a separate, higher-severity rule with a level chosen to mean "unusable regardless of what anything else managed". Two rules, two meanings: | rule | fires on a heatwave? | catches a broken feature feed? | catches both producers failing? | |---|---|---|---| | absolute error threshold, tuned tight | yes, wrongly | yes | yes | | absolute error threshold, tuned loose | no | late or never | yes | | skill ratio against the seasonal-naive baseline | no | yes | no | | ratio plus a high absolute ceiling | no | yes | yes | ## Damping: burn-in, smoothing, hysteresis A single settlement interval's error is one noisy sample, and the ratio of two noisy numbers is noisier still. Three mechanisms, in order of how much delay they buy: 1. **Smooth the series** with an exponentially weighted moving average over the ratio, so one wild interval cannot fire the rule on its own. An alpha near 0.2 gives a half-life of roughly three intervals. 2. **Require burn-in**: N consecutive matured intervals over the fire level. Three half-hourly intervals costs 90 minutes of exposure before the page, which is the price of not waking someone for a blip - a tradeoff to state explicitly rather than to tune silently. 3. **Add hysteresis**: clear at a stricter level than you fire at, so a rule that fires at a ratio of 0.90 resolves only below 0.75 and cannot chatter on a boundary. Count burn-in in **settlement intervals, not wall-clock minutes**. The metric advances only when an actual matures, so a wall-clock duration can elapse with one new data point, or with none, and the rule then either fires on a single sample or never clears. ## What this does not buy The ratio says the model stopped being worth its keep. It never says why, and it is not a drift test: the reference windows, cadences and statistics that decide whether an input distribution moved are a separate mechanism whose output belongs on the page as context. Nor does the ratio replace the judgment of which absolute level counts as unusable - that number comes from the operator's tolerance, not from the data.
- Why count burn-in in settlement intervals rather than in wall-clock minutes?Because the metric only advances when an actual matures. A five-minute wall-clock duration may contain no new observation at all, so the rule effectively fires on a single sample, and during a maturation stall it can neither fire nor clear. Counting matured intervals ties the burn-in to evidence rather than to the passage of time, and makes the delay it costs explicit: three half-hourly intervals is 90 minutes.
- Which control belongs in the ratio - the previous model version or the seasonal-naive baseline?The non-learned baseline. It never changes, needs no retraining, and reads none of the model's features, so it stays a fixed yardstick and is unaffected by the failures you are trying to catch. The previous model version is a rollback target, not a control: it shares the same feature path, so an upstream break moves both terms of the ratio and hides the very fault you wanted to see.
- What does the skill ratio miss that an absolute ceiling catches?The case where both producers degrade together - a corrupted actuals feed, or a shock so large that neither the model nor the baseline is usable. The ratio then reads normally while the published forecast is worthless. That is why the ceiling exists as a separate, higher-severity rule, with a level set by what the operator can no longer dispatch on rather than by recent error statistics.
Judging a class by its raw exam average calls a hard paper a teaching failure; comparing it against how every other class did on the same paper separates the paper from the teaching. The comparison has the same blind spot as the skill ratio: if the whole cohort was taught badly, everyone scores low together and the ranking looks fine.
saying these in an interview costs you the question
- Raises the absolute threshold after every hard night until nothing fires
- Silences the alarm for the whole summer instead of rescoping it
- Uses the training-time validation error as the live control
- Pages on a single settlement interval's error spike
- Assumes a skill ratio also covers both producers failing together
- Clears the alarm at the same level it fires at