skip to content

A page configured as 'average error ratio over the last hour exceeds 14.4 times the budget rate' fires during a ten-minute total outage of a 99.9% service. The outage is mitigated, but the page keeps firing for most of the following hour. Explain why, and how the standard fix works.

level: seniorimportance: should knowfreq 32%

answer

  1. a sliding window keeps its memory
  2. the bad traffic has not aged out
  3. reset time follows the short window
  4. clears in five minutes, not fifty-nine

basics

~20 s

The outage stays inside the trailing one-hour average until it rolls out of the window, so the condition remains true long after the burn stops. The fix is to AND the long window with a short one, typically a twelfth of its length, so the alert clears roughly five minutes after recovery.

solid answer

~50 s

A sliding-window average has memory. Ten minutes at 100% errors inside a 60-minute window averages to about 16.7%, and the threshold is only 1.44%, so the condition stays true until nearly all of the outage has aged out — the average only drops below 1.44% when under a minute of bad traffic remains in the window, meaning roughly 59 minutes of continued firing. The standard fix is the second window: require both the 1-hour average and a 5-minute average to exceed the same threshold, and fire only when both do. Once traffic recovers, the 5-minute window drains within five minutes and the alert resolves, while detection speed is unaffected because a live burn trips both windows almost immediately. Reset time then tracks the short window, not the long one — which is the whole reason the pair exists.

go deeper

for a junior

Understand that an alert on an average over the last hour keeps firing until the bad data ages out of that hour, and that pairing it with a short window is what makes it clear quickly.

for a middle

Be able to compute it: ten minutes of total failure in a sixty-minute window is a 16.7% average against a 1.44% threshold, so it stays true for nearly the full hour. Explain that the conjunction with a short window is what governs reset.

for a senior

Show the operational judgment — why a page that keeps firing after mitigation destroys trust and gets reflexively silenced, why detection speed is unaffected by the pairing, and what you would do about a short window that flaps on bursty traffic.

for a principal

Own the second-order effect: alert reset behaviour is a trust property of the whole paging system, and an estate full of alerts that outlive their incidents produces responders who acknowledge by reflex. Treat reset time as a reviewable property of an alerting standard.

## The mechanism: a sliding window remembers An alert on "the average bad-event ratio over the last hour" is evaluated against a window that slides forward in time. The bad traffic does not leave that window when the incident ends; it leaves when it becomes older than the window length. Until then it keeps contributing to the average. Work the numbers for the scenario. The threshold is `14.4 * 0.001 = 1.44%` bad events. The outage contributes ten minutes at 100% errors. Assume steady request volume, so the one-hour average immediately after mitigation is roughly `10/60 = 16.7%` — more than ten times over the line. As minutes pass, less of the outage remains inside the window. The average falls below 1.44% only when the surviving bad minutes fall under `0.0144 * 60 ≈ 0.86` minutes, i.e. about 52 seconds' worth. That happens roughly 59 minutes after mitigation. The responder has fixed the problem and the pager still says it is broken. ## Why this is worse than an annoyance Three concrete costs. First, the alert stops carrying information: nobody looking at it can tell whether the burn is live or historical, which matters most during handoffs and escalations. Second, if the alerting system re-notifies on an interval, the responder is paged repeatedly for an outage they already fixed. Third — and this is the damaging one — people learn to silence or acknowledge that alert reflexively, and the next real burn arrives at a suppressed alert. ## The fix: require recency as well as significance The multi-window design ANDs a second condition onto the same threshold, evaluated over a much shorter window — conventionally a twelfth of the long one, so 5 minutes against 1 hour, or 30 minutes against 6 hours: ``` fire when avg_bad_ratio(1h) > 1.44% AND avg_bad_ratio(5m) > 1.44% ``` After mitigation the five-minute window is emptied of bad traffic within five minutes, the conjunction becomes false, and the alert resolves — even though the one-hour term is still overwhelmingly true. **Reset time is governed by the short window.** Detection is not slowed. During a real burn the short window crosses the threshold first, not last, because it is not diluting the current traffic with the previous fifty-five quiet minutes. The conjunction is therefore satisfied at essentially the moment the long window crosses, which is what makes the pairing free. ## Choosing the short window The twelfth is a heuristic balancing two failure modes. Too short — say 30 seconds — and the short term flickers across the threshold on ordinary jitter while the long term stays true, so the conjunction flaps: fire, resolve, fire again, minutes apart. That is a distinct pathology from the one you set out to fix and it trains people to distrust the alert just as effectively. Too long — say 30 minutes against a 1-hour window — and you have barely improved reset time; the alert still hangs around for half an hour after recovery. If the short window still flaps at a twelfth, the usual levers are a small pending or "for" duration in whichever alerting engine evaluates the rule, and a minimum bad-event count for low-volume services, rather than lengthening the short window back out. ## Related recovery behaviour worth knowing A partial mitigation produces a genuinely ambiguous state: the short window drops below threshold while the long window remains far above it. The alert resolves, which is correct — the burn has stopped — but the budget damage is done and the incident is not over. This is why burn-rate pages are a signal to respond, not the record of the incident: closing the alert is not the same as closing the incident, and the remaining budget is what tells you how much room you have left before the next one. Similarly, an intermittent burn that stops and restarts every few minutes will fire, resolve, and fire again, because that is literally what is happening to the service. The right response is to treat the repeated firing as the signal it is, not to lengthen the short window until the pattern is hidden.

  • Does adding the short window slow down how fast the alert fires?
    No, and that is why the pairing costs nothing. During a live burn the short window crosses the threshold before the long one does, because it is not averaging the current traffic against fifty-five quiet minutes. The conjunction is satisfied essentially the moment the long window crosses, so detection time is set by the long window exactly as it was, while reset time drops to the short window.
  • Someone proposes fixing the stuck alert by shortening the long window to five minutes instead. What is wrong with that?
    It throws away precision to buy reset time. A five-minute window pages on a burn that has only cost about 0.17% of a 30-day budget, so ordinary transient blips start waking people. The long window exists to guarantee the event was significant; the short window exists to guarantee it is current. Collapsing them into one window means you can have one property or the other, never both.
  • The 1h/5m pair now flaps every few minutes on a service with bursty traffic. What would you change?
    First check whether the service genuinely is burning intermittently — if so the alert is right and the service is the problem. If it is measurement noise, add a short pending duration so the condition must hold for a minute or two before notifying, and on low-volume services require a minimum absolute count of bad events alongside the ratio. Lengthening the short window is the last resort, since it reintroduces the stuck-alert behaviour.

saying these in an interview costs you the question

  • Assuming the alert is stuck or the metric is broken
  • Silencing the page instead of fixing the window pair
  • Shortening the long window to make it clear faster
  • Thinking the short window is there for faster detection
  • Treating alert resolution as the end of the incident

context