skip to content

What must a load forecaster's night runbook pre-decide so an on-call engineer can roll back the model version alone?

level: principalimportance: should knowfreq 44%

answer

  1. decide it before the night, not during
  2. trigger, target, authority
  3. previous version shares the broken feed
  4. roll back features and artifact together
  5. reversible action needs no approval

basics

~20 s

Three things: the trigger that counts as evidence, the target to fall back to - the previous model version or the non-learned baseline - and the authority, meaning the responder may switch without approval because the switch is reversible and expires.

solid answer

~40 s

At 03:00 nobody can distinguish a bad model promotion from a broken upstream feed, so the runbook has to pre-commit a response that is safe under both readings. It names the trigger in the alarm's own terms, names which fallback is the default under uncertainty - the non-learned baseline, since the previous model version consumes the same features and fails identically on an input outage - and states that the named on-call may act without approval inside a declared blast radius. It also says what a rollback must cover: the model artifact **plus** the feature definitions and transformation code that shipped with it, or the fault returns. Finally it sets an expiry and a provenance flag, so the fallback cannot become permanent by inertia and consumers can see which producer served each interval.

code

json · 17 lines
json
{
  "alarm": "intraday-forecast-skill",
  "horizon": "next-settlement-interval",
  "fired_after_intervals": 3,
  "skill_ratio": 0.97,
  "absolute_error_mw": 480,
  "served_model_version": "load-fc-2026-06-11",
  "promoted_at": "2026-06-11T09:00Z",
  "previous_model_version": "load-fc-2026-04-02",
  "input_side_signals": ["temperature_feed: stale 4 intervals"],
  "default_fallback": "seasonal_naive",
  "authorised_without_approval": true,
  "blast_radius": "intraday horizon only",
  "reverts_effective": "next settlement interval",
  "fallback_expires": "2026-07-14T09:00Z",
  "runbook": "rollback-intraday-forecaster"
}

go deeper

for a junior

Learn the shape of the answer: the safe action is decided in advance and written down, so the person woken at night executes a choice rather than inventing one.

for a middle

Be able to explain why a rollback must include the feature definitions and post-processing that shipped with the model, not only the artifact.

for a senior

Show the diagnostic asymmetry - the previous model version shares the broken feature path while a non-learned baseline does not - and make the fallback expire.

for a principal

The judgment is where to draw the authority line: what blast radius an unapproved, reversible switch may cover, and which irreversible actions justify waking a second person.

## Why the decision is made in advance The responder at 03:00 has one page, no colleagues, and no way to tell whether the model degraded or the data feeding it did. Diagnosis at that hour is slow and unreliable, and every interval spent on it is published at whatever quality caused the page. A night runbook is therefore not documentation of how the system works - it is a pre-made decision that is acceptable under every reading of the evidence the responder can actually gather. That means the runbook is written when people are rested and have the counterfactuals in front of them, and reviewed whenever the model, the fallback or the decision clock changes. ## Which fallback, and why the default matters | target | cures a bad promotion | cures a dead or frozen feature feed | accuracy cost on a normal night | time to take effect | |---|---|---|---|---| | previous model version | yes | no - it reads the same feature path | small | next interval | | seasonal-naive baseline | yes | yes - it reads only past actuals | significant | next interval | | do nothing, keep serving | no | no | none, but the error continues | - | The asymmetry is the whole point. A version rollback assumes the fault was introduced by the promotion; if the real cause is upstream, the previous version consumes the same poisoned inputs and fails the same way, and the responder has spent the night's one action for nothing. The non-learned baseline is immune to the feature path, which makes it the safer default under uncertainty even though it is less accurate when everything is healthy. The runbook should state the default explicitly, and state the evidence that would justify the other choice - typically a promotion within the last few days with no accompanying input-side signal. ## What a rollback has to cover A model version is rarely just weights. Rolling back partially is a common and expensive mistake, so the runbook enumerates the set: 1. **The model artifact** actually loaded by the serving path. 2. **The feature definitions and transformation code** the candidate introduced - a rolled-back artifact reading newly-shaped features can be worse than either version on its own. 3. **Any post-processing** that shipped with the promotion: clamps, calibration adjustments, rounding to market granularity. 4. **The alarm's own thresholds**, if they were retuned to the new version's error profile, or the restored producer will look broken against a rule written for its successor. 5. **The provenance flag** on every published forecast, so downstream consumers and the morning review can tell which producer served which interval. ## The authority line The expensive failure is not a wrong rollback; it is an hour spent looking for someone to approve a reversible action. The runbook should therefore say, in plain terms, that the named on-call may switch producers without approval, within a declared blast radius - which horizons, which regions, which intervals. Reversibility is what earns that: if the switch takes effect next interval and can be undone next interval, the cost of being wrong is bounded and small, while the cost of hesitating accrues every interval. The converse also belongs in the runbook. Where an action is irreversible - anything that touches an already-submitted market position - a second person is justified, and the runbook names who and how to reach them rather than leaving the responder to invent an escalation path at 03:00. ## What the night forbids Just as important as the permitted action is the list of things the responder must not attempt: - **retraining or refitting anything** - the refresh policy is a standing mechanism with its own gates, not a night-time improvisation; - **editing the alarm's thresholds to clear the page** - that destroys the record of what fired and why, and the change escapes the review a promotion would get; - **rolling back one component and not the rest**, producing a configuration nobody has ever tested; - **deep root-cause analysis** before mitigating - the forecast is being published continuously, and diagnosis is a morning activity once the bleeding has stopped. ## Making the fallback temporary by construction A fallback that outlives the night becomes the system by inertia, and the operator ends up running a seasonal-naive forecaster for a month without deciding to. So the switch carries an **expiry** and a scheduled morning review that must either restore the model, extend the fallback deliberately, or open a defect. The fallback's own quality sits on the same alarm, which means the degraded accuracy stays visible instead of becoming the new normal against which everything is judged.

  • The cause turns out to be a frozen weather feed. Which fallback was the right one, and why?
    The non-learned seasonal-naive baseline, because it reads only past actuals. The previous model version consumes the same frozen feature and produces the same kind of wrong answer, so a version rollback would have looked like an action while changing nothing. This asymmetry is exactly why the runbook's default under uncertainty is the baseline, with a version rollback reserved for cases where a recent promotion is the leading hypothesis.
  • What stops a night-time fallback from quietly becoming the permanent production forecaster?
    An expiry on the switch plus a morning review that must make an explicit choice: restore the model, extend the fallback with a stated reason, or open a defect. Keeping the fallback's own output on the same quality alarm helps too, because its weaker accuracy stays visible rather than becoming the baseline everything is compared against. Without both, inertia decides.
  • Should the responder be allowed to adjust the alarm's thresholds during the incident to stop the paging?
    No. Editing the rule that fired removes the evidence of what happened and silences the same condition for everyone afterwards, usually without review. If the alarm is genuinely mistuned, that is a change reviewed like any other promotion, after the incident. The night-time lever for noise is acknowledging the page or applying a scoped, expiring suppression that is recorded, not rewriting the rule.

saying these in an interview costs you the question

  • Leaves the choice of fallback to the responder's judgment at 03:00
  • Rolls back the artifact but leaves new feature definitions live
  • Requires approval for a reversible switch that costs less than waiting
  • Expects the responder to diagnose model versus data before mitigating
  • Assumes a previous model version survives an upstream feature outage
  • Leaves the fallback running with no expiry and no review