skip to content

Burn Rates & SLO Alerting

Alerting on how fast the error budget burns rather than on raw thresholds. A senior-level SRE interview favorite: multi-window multi-burn-rate alerting is the canonical answer to 'how do you page on SLOs without noise?'.

on this pageshow

questions

5

A service has a 99.9% availability SLO measured over a rolling 30-day window. Explain what error-budget burn rate means, what a burn rate of exactly 1 signifies, and how you compute it from an observed error ratio.

level: middleimportance: must knowfreq 65%

answer

  1. consumption expressed as a speed
  2. rate one lasts exactly one window
  3. divide observed ratio by the budget
  4. ten times on 99.9% means three days

basics

~20 s

Burn rate expresses error-budget consumption as a multiple of the sustainable pace: rate 1 spends the whole budget exactly at the window's end. Compute it as the observed bad-event ratio divided by (1 minus the SLO target).

solid answer

~50 s

The error budget is the complement of the target — for a 99.9% SLO that is 0.1% of requests, which over 30 days is also about 43.2 minutes of full unavailability. Burn rate turns that quantity into a speed: observed bad-event ratio divided by `1 - target`. A burn rate of 1 means you are consuming budget at precisely the pace that exhausts it at the end of the window, no sooner and no later. A sustained 1% error ratio against a 99.9% target is therefore a burn rate of 10, and at that pace a full 30-day budget is gone in three days. Two derived numbers matter operationally: time to exhaustion is window divided by burn rate, and budget spent over an interval is burn rate times interval divided by window — which is why 14.4x sustained for one hour spends exactly 2% of a 30-day budget.

code

python · 12 lines
python
def burn_rate(bad_ratio, slo_target):
    return bad_ratio / (1.0 - slo_target)

def hours_to_exhaustion(rate, window_hours=720.0):
    return window_hours / rate if rate else float("inf")

def budget_fraction(rate, interval_hours, window_hours=720.0):
    return rate * interval_hours / window_hours

print(burn_rate(0.01, 0.999))                 # 10.0
print(hours_to_exhaustion(10.0) / 24)         # 3.0 days
print(budget_fraction(14.4, 1.0))             # 0.02 -> 2% of the budget

go deeper

for a junior

Be able to state that the error budget is one minus the target and that burn rate says how fast you are spending it. Saying 'rate 1 uses it all up exactly at the end of the window' is enough at this level.

for a middle

You are expected to do the arithmetic out loud: divide the observed bad-event ratio by one minus the target, convert a rate into days of runway, and explain why burn rate is stated per observation window rather than instantaneously.

for a senior

Show you use the number to make calls — reading remaining budget alongside the rate on a rolling window, distinguishing a violent short spike from a slow drain of equal budget cost, and knowing which events count in the denominator.

for a principal

Own the framing: burn rate is what makes reliability targets comparable across services with different SLO targets, and it is the unit that lets one alerting standard serve a whole estate. Be ready to defend that normalisation as a platform decision.

## The quantity and the speed An SLO sets a target for a service level indicator — say, 99.9% of valid requests succeed over a rolling 30-day window. The complement of the target is the error budget: `1 - 0.999 = 0.001`, or 0.1% of requests. Because 30 days is 43,200 minutes, the same budget is often quoted as time: 0.1% of 43,200 minutes is 43.2 minutes. That quantity alone cannot drive an alert. Knowing you have 43.2 minutes of budget says nothing about whether the current error rate is a rounding error or an emergency. Burn rate supplies the missing dimension: it is consumption expressed as a **multiple of the exactly sustainable pace**. ## Rate 1 is the anchor Burn rate 1 is defined so that the budget is spent precisely at the end of the SLO window — you finish the 30 days having used exactly 100% of the budget. Everything else scales from that anchor: - rate 2 → budget gone in 15 days - rate 10 → gone in 3 days - rate 0.5 → half the budget still unspent at window end - rate 0 → a perfect window This is why burn rate, not raw error ratio, is the natural alerting unit: the number already answers "how long do I have?" ## The arithmetic For a ratio-based SLI the formula is one division: ``` burn_rate = observed_bad_ratio / (1 - slo_target) ``` With a 99.9% target, `1 - target` is 0.001, so: - 0.1% errors → burn rate 1 - 1% errors → burn rate 10 - 100% errors (total outage) → burn rate 1000 Two corollaries fall out. Time to exhaust a full budget is `window / burn_rate`. And the fraction of budget consumed by burning at rate B for an interval T is: ``` budget_fraction = B * T / window ``` A 30-day window is 720 hours, so burning at 14.4x for one hour spends `14.4 * 1 / 720 = 2%` of the budget, and burning at 6x for six hours spends `6 * 6 / 720 = 5%`. Those two products are exactly where the canonical alert thresholds come from. ## Burn rate is always an average over a window A burn rate is not an instantaneous property; it is computed over some observation window. "Burn rate 14.4 over the last hour" means the average bad-event ratio across that hour was 14.4 times the budget rate. This matters twice. First, a short, violent spike inside a long window is diluted: ten minutes of total failure inside a one-hour window averages to roughly 16.7% errors, not 100%. Second, the same real-world event produces different burn-rate readings depending on which window you measure over — so a burn-rate threshold is meaningless unless the window is stated alongside it. ## Why it is the right unit for paging Burn rate is dimensionless with respect to the target. A 99.9% service and a 99.99% service can share the same burn-rate thresholds, because each is measured against its own budget; a raw error-percentage threshold cannot be shared that way, since 0.5% errors is comfortable for one and catastrophic for the other. And burn rate maps directly onto the only question that matters at 3 a.m.: at this pace, how much runway is left? ## Traps The most common mistake is quoting burn rate as if it were an error percentage — "we're burning at 2%" is meaningless. The second is assuming the budget starts full: on a rolling window the budget already reflects the trailing 30 days, so a burn rate of 10 may exhaust an already-half-spent budget in a day and a half, not three days. The third is forgetting that the denominator is bad *valid* events, so misclassifying load-test traffic or health checks as user requests silently moves the number. ## A worked example A 99.9%/30-day service normally serves 0.02% errors (burn rate 0.2 — healthy). A bad deploy pushes it to 3% for 40 minutes. Burn rate during the incident is `0.03 / 0.001 = 30`. Budget consumed is `30 * (40/60) / 720 = 2.8%`. That is a real page-worthy event, and yet it costs under 3% of the month's budget — which is precisely the calibration a fast-burn alert is trying to achieve.

  • The SLO window is rolling rather than calendar-based. How does that change what a given burn rate implies?
    On a rolling window the budget is never guaranteed to start full — it already reflects the trailing 30 days. So burn rate still tells you the pace, but runway is `remaining_budget / burn_rate`, not the full window divided by the rate. A service that spent 60% of its budget last week has under half the runway the raw rate suggests, which is why burn-rate dashboards show remaining budget beside the rate.
  • How would you express burn rate for a latency SLO rather than an availability one?
    Identically, because the SLI is still a ratio of good to valid events — you just redefine "good" as requests served under the latency threshold. If the SLO is 99% of requests under 300 ms, the budget rate is 0.01, and an observed 4% of requests over 300 ms is a burn rate of 4. The math never touches the percentile itself; it counts events on the wrong side of the threshold.
  • Your dashboard shows a burn rate of 0.8. Is that good or bad?
    It is technically within budget but leaves almost no margin: at 0.8 you finish the window having spent 80% of the budget on business as usual, so a single significant incident pushes you over. A healthy steady-state burn rate is well under 1 — often an order of magnitude under — precisely so that the budget is available to absorb the unplanned events it exists for.

saying these in an interview costs you the question

  • Quoting burn rate as a percentage of requests
  • Thinking burn rate 1 means the SLO is violated
  • Assuming the budget always starts full each window
  • Reporting a burn rate without saying over which window
  • Believing a higher SLO target needs different burn thresholds

context

open as a page

You are designing SLO-based paging for a service with a 99.9% availability SLO over 30 days, using multi-window multi-burn-rate alerts. Which burn-rate and window pairs would you choose, and what job does each of the two windows in a pair do?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Use two paging tiers: 14.4x averaged over 1 hour for fast burns and 6x over 6 hours for slower ones, each ANDed with a short window (5 minutes and 30 minutes) of the same threshold. The long window sets sensitivity; the short one proves the burn is still happening.

open as a page

Many teams page on a static rule such as 'error ratio above 2% for five minutes'. Against a 99.9% availability SLO over 30 days, explain the two opposite ways that static threshold fails, and what burn-rate alerting changes.

level: middleimportance: should knowfreq 45%

basics

~20 s

A static error-ratio threshold ignores both the SLO and how long the condition lasts. It pages too early on brief or low-volume blips that cost almost no budget, and stays silent through a sustained sub-threshold error rate that quietly drains the entire budget in days.

open as a page

A page configured as 'average error ratio over the last hour exceeds 14.4 times the budget rate' fires during a ten-minute total outage of a 99.9% service. The outage is mitigated, but the page keeps firing for most of the following hour. Explain why, and how the standard fix works.

level: seniorimportance: should knowfreq 32%

basics

~20 s

The outage stays inside the trailing one-hour average until it rolls out of the window, so the condition remains true long after the burn stops. The fix is to AND the long window with a short one, typically a twelfth of its length, so the alert clears roughly five minutes after recovery.

open as a page

You own the alerting standard for a platform of several hundred services, some with 30-day SLO windows and some with 7-day windows. Would you mandate a single burn-rate alert configuration for all of them? Explain what must vary per service and what you would hold fixed.

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Hold the budget-spend policy fixed, not the multipliers. Burn rate already normalises across different SLO targets, but the multipliers are derived from the window length, so a 7-day SLO needs 3.36x where a 30-day SLO needs 14.4x, and low-traffic services need an extra event-count guard.

open as a page