skip to content

For a grid operator's intraday load forecast, which number should the single paging model-quality alarm be tied to?

level: seniorimportance: must knowfreq 66%

answer

  1. protects a decision, not a model
  2. score what shipped, not the snapshot
  3. megawatts, not only percent
  4. one pager per forecast horizon
  5. frozen offline score cannot move

basics

~20 s

Tie it to the error of the forecast actually published, in megawatts against settled actuals, scored per forecast horizon. Not the offline validation loss, and not a per-feature shift statistic - neither one moves when the operator starts paying.

solid answer

~40 s

The alarm should watch the quantity the forecast exists to move: megawatts of error on the published forecast for the horizon whose commitment is still open, scored against actuals as each settlement interval settles. Three consequences follow. Score the **published** series, including intervals a fallback produced, because that is what the operator dispatched on. Express it in megawatts rather than only a percentage, because the same percentage error costs far more at system peak than overnight, and keep the sign, because buying short-notice reserve and wasting committed capacity are not priced the same. Evaluate only intervals whose actuals have matured, and alarm separately if maturation stops. Everything else - shift statistics, prediction-distribution summaries, the offline score - stays a diagnostic attached to the page rather than a second pager.

go deeper

for a junior

Remember the shape: the alarm watches the forecast that was actually sent out, compared against what the load turned out to be, not a score computed during training.

for a middle

Be able to explain why an offline validation number cannot detect a live problem, and why the published series and the model's own output are different series that can disagree.

for a senior

Show that you pick the unit and the horizon from the decision being protected, handle fallback-produced intervals, and add a staleness alarm so a blind metric cannot look healthy.

for a principal

The tradeoff to argue is how much of the business cost you fold into one number: a cost-weighted metric pages for the right reasons but is harder to explain and to defend when the cost model itself is wrong.

## What the alarm is actually protecting A short-term load forecaster does not exist to minimise a training loss. It exists so a grid operator commits the right generation and reserve for the intervals ahead. Every megawatt the forecast is wrong by turns into either balancing energy bought at short notice or committed capacity nobody needed. The paging alarm is the thing that wakes a human, so it should be tied to the number that moves when the operator starts paying. That single decision fixes three properties of the metric before any threshold is argued about: - it is computed on the forecast that was **published**, not on the model's raw output; - it is expressed in the unit the operator acts in - **megawatts** - on a horizon whose commitment is still open; - it is scored **per horizon**, because a 30-minute-ahead miss and a day-ahead miss are paid for by different decisions at different times. ## Measure the series that was served Between the model and the operator sits a publisher: clamping to plausible limits, rounding to the market's granularity, a fallback ladder that substitutes a previous forecast or a non-learned baseline when scoring fails, and the delivery itself. A quality alarm scored on the model's own output goes green through an outage that the fallback covered badly, because the model's numbers were never the ones dispatched on. So score the published series, whichever component produced each interval, and keep the model-only series beside it as a **diagnostic**. The pair is what tells a responder whether the model degraded or the plumbing around it did - which is exactly the first branch of the runbook. Suspending the alarm while a fallback serves is the wrong instinct: a fallback running for six hours is a quality problem, not an exemption. ## Choose the unit the operator acts in A percentage is a convenient headline and a poor pager. Mean absolute percentage error of 1.2% on a 30 GW system peak is about **360 MW**; the same 1.2% at a 15 GW overnight trough is **180 MW**. One of those is a real reserve decision and the other is noise, and a single percentage threshold cannot tell the responder which one fired. Megawatt error, or the cost of covering it, is the number the business actually feels; percentage error remains useful for comparing across days and against the baseline. Sign matters too. Under-forecasting means generation was not scheduled and has to be found at short notice; over-forecasting means capacity was committed and wasted. Those rarely cost the same, so a symmetric threshold is a choice that should be made deliberately rather than inherited from a training metric. | candidate number | what it is good for | why it is not the pager | |---|---|---| | offline validation score on a frozen snapshot | the promotion gate before the model went live | a fixed model on fixed data returns the same number forever | | a per-feature shift statistic | the first hypothesis once something has paged | moves with every season without harming the forecast | | a prediction-distribution summary | early warning while actuals are still immature | describes outputs, not quality | | percentage error on the published forecast | comparing days and tracking trend | hides how many megawatts a peak-hour miss costs | | megawatt error on the published forecast, per horizon | **the pager** | - | ## One pager, many diagnostics The reason to insist on one number is that a pager with five inputs has five ways to wake someone and no agreed meaning when it does. The working shape is: 1. **One paging alarm per horizon family**, tied to served error against matured actuals. 2. **Diagnostics attached to that page** - recent shift-test results, feature staleness, the model-only series, the fallback rate - so the responder opens the page with hypotheses already ranked. 3. **A separate staleness alarm** on the measurement path itself: if no settlement interval has matured for several cycles, the quality alarm is blind and silence means nothing. An alarm that cannot go red is not the same as a healthy system. ## What this alarm is not It is not a drift detector: it says the forecast got worse, never why. It is not an estimate of quality for intervals whose actuals have not arrived - the alarm evaluates matured intervals only, and estimating quality under late truth is a separate mechanism with its own error bars. And firing it does not itself decide to retrain; the refresh policy is a standing mechanism elsewhere, while the alarm's job ends at getting a human to a lever tonight.

  • The forecaster emits a probabilistic forecast. Should expected calibration error or the Brier score become the paging number?
    Only if the downstream decision consumes the distribution - if reserve is sized from a specific quantile, alarm on the coverage or pinball loss of that quantile, because that is the number being spent. Otherwise a calibration metric is a diagnostic that explains a page, not one that causes it: calibration can drift while the point forecast the operator dispatches on stays accurate enough.
  • Actuals for a settlement interval arrive with a lag. What is the alarm metric doing in the meantime?
    Nothing for those intervals. The alarm is evaluated only on intervals whose actuals have matured; the freshest ones are censored and excluded rather than scored as if the actual were zero. Because that leaves a blind window, pair it with a staleness alarm on the maturation path itself, so a stopped actuals feed pages instead of producing a reassuring green dashboard.
  • Why not one alarm on the average error across all horizons?
    Averaging mixes populations with different error scales and different decision deadlines, so a bad intraday night is diluted by good day-ahead intervals and the page arrives late or never. It also destroys the page's meaning: the responder cannot tell which commitment is at risk, and the mitigation differs by horizon. Keep one alarm per horizon family and let the dashboard do the aggregation.

saying these in an interview costs you the question

  • Alarms on a frozen validation score that live traffic cannot move
  • Scores the model's raw output and ignores what was published
  • Reports only percentage error, hiding megawatts at system peak
  • Assumes over-forecasting and under-forecasting cost the same
  • Averages every forecast horizon into one headline number
  • Reads a silent alarm as healthy when actuals stopped arriving