skip to content

You are designing SLO-based paging for a service with a 99.9% availability SLO over 30 days, using multi-window multi-burn-rate alerts. Which burn-rate and window pairs would you choose, and what job does each of the two windows in a pair do?

level: seniorimportance: must knowfreq 58%

answer

  1. two tiers, and two windows in each
  2. long window is sensitivity, short is recency
  3. fourteen point four over an hour
  4. two percent of budget in one hour
  5. short window is a twelfth of the long

basics

~20 s

Use two paging tiers: 14.4x averaged over 1 hour for fast burns and 6x over 6 hours for slower ones, each ANDed with a short window (5 minutes and 30 minutes) of the same threshold. The long window sets sensitivity; the short one proves the burn is still happening.

solid answer

~50 s

I would run a fast tier and a slow tier. The fast page fires when the burn rate exceeds 14.4x averaged over the last hour, which is exactly 2% of a 30-day budget consumed in that hour. The slow page fires at 6x over 6 hours, which is 5% of the budget. Each threshold is evaluated over two windows at once: the long window and a short window roughly a twelfth of its length — 5 minutes and 30 minutes respectively — and the alert only fires when both exceed the threshold. The long window supplies precision: it stops a brief blip that costs almost no budget from paging anyone. The short window supplies recency: it confirms the burn is ongoing right now, and it is what lets the alert clear within minutes of mitigation instead of hanging around for the rest of the long window. Below those, a 1x-over-3-days tier catches slow drains and is not worth a page.

code

python · 11 lines
python
WINDOW_HOURS = 30 * 24  # 720

def multiplier(budget_fraction, long_window_hours):
    return budget_fraction * WINDOW_HOURS / long_window_hours

for frac, hours in [(0.02, 1), (0.05, 6), (0.10, 72)]:
    short = hours / 12
    print(f"{multiplier(frac, hours):5.2f}x  long={hours}h  short={short:.2f}h")
# 14.40x  long=1h   short=0.08h
#  6.00x  long=6h   short=0.50h
#  1.00x  long=72h  short=6.00h

go deeper

for a junior

Know that SLO alerts are configured as a burn-rate threshold plus a time window, and that the canonical fast page is 14.4x over one hour. Be able to say why a single fixed error percentage is not the same thing.

for a middle

Be ready to derive the multipliers rather than recite them: budget spent equals rate times window over the SLO window, so 2% in an hour on a 30-day SLO gives 14.4x. Explain what each window in a pair is for.

for a senior

Demonstrate you have tuned these in production — computing detection time for a real outage shape, choosing tiers for the burn rates your service actually exhibits, and knowing that the low-traffic case needs an extra guard before any of this holds.

for a principal

Own the policy behind the numbers: the multipliers encode a decision about how much budget an organisation will spend before interrupting a human. Be prepared to argue for a different budget-spend policy and show how the whole table shifts with it.

## Why a single threshold cannot work A burn-rate alert has to satisfy two contradictory demands. It must fire fast enough that a total outage does not eat the month's budget before anyone notices, and it must not fire on every transient blip that costs a rounding error's worth of budget. One threshold over one window cannot do both: a short window is fast but noisy, a long window is precise but slow. The multi-window multi-burn-rate design resolves this by running **several thresholds at several sensitivities**, and by pairing each with a second, shorter window. ## The canonical table For a 30-day SLO window, the widely cited configuration is: | Burn rate | Budget consumed | Long window | Short window | Response | |---|---|---|---|---| | 14.4 | 2% | 1 hour | 5 minutes | page | | 6 | 5% | 6 hours | 30 minutes | page | | 1 | 10% | 3 days | 6 hours | ticket | Every number here is derived, not chosen by taste. Budget consumed is `rate * window / 720 hours`. Pick "2% of the budget in one hour is worth waking someone" and the rate falls out: `0.02 * 720 / 1 = 14.4`. Pick "5% in six hours" and you get `0.05 * 720 / 6 = 6`. Pick "10% in three days" and you get `0.10 * 720 / 72 = 1`. Change the policy — say 3% in an hour instead of 2% — and the multiplier changes to 21.6 with it. ## What the long window buys The long window is the precision knob. Averaging over an hour means a 30-second glitch is diluted 120-fold and never reaches the threshold, so it never pages. That is the desired behaviour: an event that costs 0.05% of the budget is not an emergency no matter how ugly the graph looks. Lengthen the window and you gain precision but lose detection speed; shorten it and you gain speed at the cost of noise. Critically, a long window does **not** mean slow detection of severe events. The alert fires as soon as the *average* over the window crosses the threshold. During a total outage of a 99.9% service the instantaneous ratio is 100%, and the one-hour average crosses `14.4 * 0.001 = 1.44%` after about 1.44% of the hour — roughly 52 seconds. The window length caps how long a *mild* burn must persist; a catastrophic one trips it almost immediately. ## What the short window buys Without a second condition, an alert on a one-hour average keeps firing long after the burn has stopped, because the outage stays inside the trailing window. Requiring the short window to be over the threshold too means the alert reflects the *present*. Two consequences follow: - **Reset time collapses** from roughly the long window to roughly the short window. The 1h/5m pair clears about five minutes after the burn stops rather than fifty-nine. - **The alert becomes about a live condition**, so an incident responder looking at it knows whether the thing is still burning or already mitigated. The conventional short window is a twelfth of the long one, which is where 5 minutes, 30 minutes and 6 hours come from. It is a heuristic, not a law: too short and the pair flaps on ordinary jitter, too long and reset time creeps back up. ## Why two paging tiers rather than one The 14.4x tier is deliberately deaf to anything below a serious burn. A service quietly running at 8x is not caught by it — yet 8x drains the full budget in under four days. The 6x/6h tier exists exactly for that band: severe enough to matter, too gradual to trip the fast page. Together the tiers cover a wide range of burn shapes; the 1x/3d tier catches the slow drain that is real but does not justify a night-time interruption. ## What this does not fix Multi-window alerting is arithmetic on a ratio, and it inherits that ratio's weaknesses. On a low-traffic service a single failed request can push the short window past any threshold, so a minimum-event guard or a longer window is usually needed. It also says nothing about *what* is broken — it tells you the user-visible contract is being violated at a given pace, and the diagnosis happens elsewhere. And if the SLI is measured somewhere that shares the failure mode of the service, the alert can go quiet exactly when it is needed most.

  • How long would it take the 14.4x over 1 hour alert to fire during a complete outage of that 99.9% service?
    About 52 seconds. The condition is on the one-hour average bad-event ratio crossing 14.4 × 0.001 = 1.44%. With the instantaneous ratio at 100%, the average reaches 1.44% once 1.44% of the hour has elapsed, which is roughly 52 seconds. This is the point people miss: a long averaging window does not mean slow detection of severe burns, only of mild ones.
  • Why is the 6x tier over six hours rather than simply lowering the one-hour threshold to 6x?
    Because 6x over one hour is only 0.83% of the budget — far too little to justify a page, and common enough that you would be woken constantly. Stretching the window to six hours means the alert requires the burn to persist, so by the time it fires 5% of the budget really is gone. The window and the multiplier are chosen together to hold budget-spend constant, not independently.
  • What breaks if you set the short window to half the long window instead of a twelfth?
    You lose most of the benefit. Reset time is governed by the short window, so a 30-minute short window on a one-hour alert still leaves the page firing for half an hour after mitigation. You also weaken the recency signal, because a burn that stopped twenty minutes ago still satisfies the short condition. A twelfth is a compromise that keeps reset fast without letting ordinary jitter flap the pair.

saying these in an interview costs you the question

  • Treating 14.4 as a magic constant rather than derived
  • Believing a one-hour window means one-hour detection
  • Adding the short window to speed up detection
  • Reusing the 30-day multipliers on a 7-day SLO window
  • Configuring one burn tier and calling it multi-burn-rate

context