skip to content

A service has a 99.9% availability SLO measured over a rolling 30-day window. Explain what error-budget burn rate means, what a burn rate of exactly 1 signifies, and how you compute it from an observed error ratio.

level: middleimportance: must knowfreq 65%

answer

  1. consumption expressed as a speed
  2. rate one lasts exactly one window
  3. divide observed ratio by the budget
  4. ten times on 99.9% means three days

basics

~20 s

Burn rate expresses error-budget consumption as a multiple of the sustainable pace: rate 1 spends the whole budget exactly at the window's end. Compute it as the observed bad-event ratio divided by (1 minus the SLO target).

solid answer

~50 s

The error budget is the complement of the target — for a 99.9% SLO that is 0.1% of requests, which over 30 days is also about 43.2 minutes of full unavailability. Burn rate turns that quantity into a speed: observed bad-event ratio divided by `1 - target`. A burn rate of 1 means you are consuming budget at precisely the pace that exhausts it at the end of the window, no sooner and no later. A sustained 1% error ratio against a 99.9% target is therefore a burn rate of 10, and at that pace a full 30-day budget is gone in three days. Two derived numbers matter operationally: time to exhaustion is window divided by burn rate, and budget spent over an interval is burn rate times interval divided by window — which is why 14.4x sustained for one hour spends exactly 2% of a 30-day budget.

code

python · 12 lines
python
def burn_rate(bad_ratio, slo_target):
    return bad_ratio / (1.0 - slo_target)

def hours_to_exhaustion(rate, window_hours=720.0):
    return window_hours / rate if rate else float("inf")

def budget_fraction(rate, interval_hours, window_hours=720.0):
    return rate * interval_hours / window_hours

print(burn_rate(0.01, 0.999))                 # 10.0
print(hours_to_exhaustion(10.0) / 24)         # 3.0 days
print(budget_fraction(14.4, 1.0))             # 0.02 -> 2% of the budget

go deeper

for a junior

Be able to state that the error budget is one minus the target and that burn rate says how fast you are spending it. Saying 'rate 1 uses it all up exactly at the end of the window' is enough at this level.

for a middle

You are expected to do the arithmetic out loud: divide the observed bad-event ratio by one minus the target, convert a rate into days of runway, and explain why burn rate is stated per observation window rather than instantaneously.

for a senior

Show you use the number to make calls — reading remaining budget alongside the rate on a rolling window, distinguishing a violent short spike from a slow drain of equal budget cost, and knowing which events count in the denominator.

for a principal

Own the framing: burn rate is what makes reliability targets comparable across services with different SLO targets, and it is the unit that lets one alerting standard serve a whole estate. Be ready to defend that normalisation as a platform decision.

## The quantity and the speed An SLO sets a target for a service level indicator — say, 99.9% of valid requests succeed over a rolling 30-day window. The complement of the target is the error budget: `1 - 0.999 = 0.001`, or 0.1% of requests. Because 30 days is 43,200 minutes, the same budget is often quoted as time: 0.1% of 43,200 minutes is 43.2 minutes. That quantity alone cannot drive an alert. Knowing you have 43.2 minutes of budget says nothing about whether the current error rate is a rounding error or an emergency. Burn rate supplies the missing dimension: it is consumption expressed as a **multiple of the exactly sustainable pace**. ## Rate 1 is the anchor Burn rate 1 is defined so that the budget is spent precisely at the end of the SLO window — you finish the 30 days having used exactly 100% of the budget. Everything else scales from that anchor: - rate 2 → budget gone in 15 days - rate 10 → gone in 3 days - rate 0.5 → half the budget still unspent at window end - rate 0 → a perfect window This is why burn rate, not raw error ratio, is the natural alerting unit: the number already answers "how long do I have?" ## The arithmetic For a ratio-based SLI the formula is one division: ``` burn_rate = observed_bad_ratio / (1 - slo_target) ``` With a 99.9% target, `1 - target` is 0.001, so: - 0.1% errors → burn rate 1 - 1% errors → burn rate 10 - 100% errors (total outage) → burn rate 1000 Two corollaries fall out. Time to exhaust a full budget is `window / burn_rate`. And the fraction of budget consumed by burning at rate B for an interval T is: ``` budget_fraction = B * T / window ``` A 30-day window is 720 hours, so burning at 14.4x for one hour spends `14.4 * 1 / 720 = 2%` of the budget, and burning at 6x for six hours spends `6 * 6 / 720 = 5%`. Those two products are exactly where the canonical alert thresholds come from. ## Burn rate is always an average over a window A burn rate is not an instantaneous property; it is computed over some observation window. "Burn rate 14.4 over the last hour" means the average bad-event ratio across that hour was 14.4 times the budget rate. This matters twice. First, a short, violent spike inside a long window is diluted: ten minutes of total failure inside a one-hour window averages to roughly 16.7% errors, not 100%. Second, the same real-world event produces different burn-rate readings depending on which window you measure over — so a burn-rate threshold is meaningless unless the window is stated alongside it. ## Why it is the right unit for paging Burn rate is dimensionless with respect to the target. A 99.9% service and a 99.99% service can share the same burn-rate thresholds, because each is measured against its own budget; a raw error-percentage threshold cannot be shared that way, since 0.5% errors is comfortable for one and catastrophic for the other. And burn rate maps directly onto the only question that matters at 3 a.m.: at this pace, how much runway is left? ## Traps The most common mistake is quoting burn rate as if it were an error percentage — "we're burning at 2%" is meaningless. The second is assuming the budget starts full: on a rolling window the budget already reflects the trailing 30 days, so a burn rate of 10 may exhaust an already-half-spent budget in a day and a half, not three days. The third is forgetting that the denominator is bad *valid* events, so misclassifying load-test traffic or health checks as user requests silently moves the number. ## A worked example A 99.9%/30-day service normally serves 0.02% errors (burn rate 0.2 — healthy). A bad deploy pushes it to 3% for 40 minutes. Burn rate during the incident is `0.03 / 0.001 = 30`. Budget consumed is `30 * (40/60) / 720 = 2.8%`. That is a real page-worthy event, and yet it costs under 3% of the month's budget — which is precisely the calibration a fast-burn alert is trying to achieve.

  • The SLO window is rolling rather than calendar-based. How does that change what a given burn rate implies?
    On a rolling window the budget is never guaranteed to start full — it already reflects the trailing 30 days. So burn rate still tells you the pace, but runway is `remaining_budget / burn_rate`, not the full window divided by the rate. A service that spent 60% of its budget last week has under half the runway the raw rate suggests, which is why burn-rate dashboards show remaining budget beside the rate.
  • How would you express burn rate for a latency SLO rather than an availability one?
    Identically, because the SLI is still a ratio of good to valid events — you just redefine "good" as requests served under the latency threshold. If the SLO is 99% of requests under 300 ms, the budget rate is 0.01, and an observed 4% of requests over 300 ms is a burn rate of 4. The math never touches the percentile itself; it counts events on the wrong side of the threshold.
  • Your dashboard shows a burn rate of 0.8. Is that good or bad?
    It is technically within budget but leaves almost no margin: at 0.8 you finish the window having spent 80% of the budget on business as usual, so a single significant incident pushes you over. A healthy steady-state burn rate is well under 1 — often an order of magnitude under — precisely so that the budget is available to absorb the unplanned events it exists for.

saying these in an interview costs you the question

  • Quoting burn rate as a percentage of requests
  • Thinking burn rate 1 means the SLO is violated
  • Assuming the budget always starts full each window
  • Reporting a burn rate without saying over which window
  • Believing a higher SLO target needs different burn thresholds

context