skip to content

Many teams page on a static rule such as 'error ratio above 2% for five minutes'. Against a 99.9% availability SLO over 30 days, explain the two opposite ways that static threshold fails, and what burn-rate alerting changes.

level: middleimportance: should knowfreq 45%

answer

  1. it knows nothing about the budget
  2. cost is intensity times duration
  3. one failed request out of ten
  4. fifteen times burn stays under two percent

basics

~20 s

A static error-ratio threshold ignores both the SLO and how long the condition lasts. It pages too early on brief or low-volume blips that cost almost no budget, and stays silent through a sustained sub-threshold error rate that quietly drains the entire budget in days.

solid answer

~50 s

The rule has no notion of budget, so it misjudges in both directions. Too early: a five-minute spike to 3% on a 99.9% service costs about 0.35% of the month's budget, and on a low-traffic endpoint serving ten requests in that window a single failure reads as 10% — a page for one failed request. Too late: a steady 1.5% error ratio never crosses 2%, so nothing fires, yet that is a burn rate of 15 and the full 30-day budget is gone in two days. Burn-rate alerting fixes both by normalising against `1 - target` and by requiring the burn to persist long enough to cost a stated share of the budget. The threshold then means something operational — 'we are on track to lose 2% of the month's budget in this hour' — rather than an arbitrary percentage someone picked once.

go deeper

for a junior

Be able to say that a fixed error-percentage alert ignores how long the problem lasts and how big the service is, and that SLO-based alerting measures budget consumption instead.

for a middle

Work the two failure directions with numbers: compute the budget cost of a short spike, and show that a sustained rate below the threshold can still be a burn rate in the double digits. Name what burn rate normalises.

for a senior

Show the operational consequence — a noisy static threshold is why nobody trusts the pager, and the blind band beneath it is why breaches arrive as a surprise. Be ready to say where a static threshold is still correct.

for a principal

Frame it as a standards question: hand-tuned thresholds per service cannot be reviewed, compared or audited across an estate, whereas a burn-rate threshold derived from a budget-spend policy makes the organisation's interrupt tolerance explicit and uniform.

## What a static threshold actually asserts A rule like *error ratio above 2% for five minutes* encodes exactly two facts: a percentage and a duration, both chosen by hand. It knows nothing about the reliability target the service has promised, nothing about how much of the month's allowance the condition consumes, and nothing about how many requests the ratio was computed from. Every one of those omissions produces a specific failure. ## Failure one: pages too early The threshold fires on conditions that cost almost nothing. Take a five-minute excursion to 3% errors. Against a 99.9% target the burn rate is 30, but it lasts only five minutes, so the budget cost is `30 * (5/60) / 720 = 0.35%`. Three tenths of one percent of the month, in exchange for waking a human. Do that a few times a week and the pager becomes background noise while the budget stays untouched. The low-traffic case is worse because the ratio itself becomes unstable. An internal endpoint serving 10 requests in a five-minute window produces error ratios quantised to 0%, 10%, 20% — a single timeout, possibly a client that gave up, reads as 10% and clears the 2% bar by five times over. The alert is not measuring reliability at that volume; it is measuring the arrival of one bad request. ## Failure two: fires too late, or never The same threshold is blind to everything beneath it, however long it lasts. A service sitting at 1.5% errors is comfortably under 2%, so the rule never fires — and 1.5% against a 0.1% budget rate is a burn rate of 15. The entire 30-day budget is gone in `30 / 15 = 2 days`. Nothing pages, nothing tickets, and the SLO is breached before the end of the week by a condition the alert was literally configured to watch. The general problem is that a static threshold treats intensity as the only variable and ignores duration entirely, whereas budget cost is the product of the two. A brief severe event and a long mild one can cost identical budget while landing on opposite sides of any fixed line. ## What burn rate changes Burn-rate alerting replaces both hand-chosen numbers with derived ones. The percentage is normalised. Instead of an absolute ratio, the condition is `observed_ratio / (1 - target)`, so the threshold is stated relative to what the service has promised. The same 14.4x condition means 1.44% errors on a 99.9% service and 0.144% on a 99.99% one — automatically, with no retuning. The duration is derived from a budget-spend policy. Instead of "five minutes feels right", you choose how much budget you are prepared to lose before interrupting someone, and the window and multiplier follow: 2% of budget in one hour gives 14.4x/1h. The number now has a defensible meaning, and when someone asks "why 14.4?" the answer is a policy decision, not a shrug. And because several tiers cover several burn shapes, the slow drain that a static threshold sleeps through is caught by a lower multiplier over a longer window. ## Where static thresholds are still the right tool This is not an argument that every threshold must be a burn rate. Some conditions have no meaningful ratio — a certificate expiring in 48 hours, a queue that must never exceed a hard capacity, a replica count of zero. Those are binary or absolute facts, and a static rule states them precisely. The argument is narrower: for a service-level indicator that already has an SLO, expressing the alert in any unit other than budget consumption throws away information you already have. ## The residual problem burn rate does not solve Low traffic remains hard. Burn rate is still a ratio, so ten requests still quantise. The usual mitigations are to require a minimum absolute count of bad events alongside the ratio, to widen the window until enough events accumulate, to aggregate several small services into one SLO, or to accept a looser target for that service. What you should not do is pretend the ratio is meaningful and page on it.

  • Would raising the static threshold from 2% to 5% fix the noise problem?
    It trades one failure for the other. Fewer brief blips reach 5%, so the pager quietens — but the blind band beneath it widens, and now a sustained 4% error rate (burn rate 40, budget gone in eighteen hours) is invisible. Any single fixed line has this property: moving it shifts which mistake you make, it does not eliminate the mistake.
  • A team says their static threshold has worked fine for years. What would you check before agreeing?
    Whether it is working, or whether nobody is looking. I would check how many pages it produced last quarter and how many corresponded to real budget spend, then check the budget history for periods of significant drain that produced no page at all. Both silences are informative: a threshold that never fires during a month that lost 40% of its budget is not calibrated, it is decorative.
  • How would you alert on an SLI for a service that genuinely serves only a few hundred requests a day?
    Not on a five-minute ratio. Either widen the window until enough events accumulate to make the ratio meaningful, require a minimum absolute bad-event count alongside the burn-rate condition, roll the service into a shared SLO with related services, or drive the SLI from synthetic probes so the request volume is known and constant. Choosing a looser target is also honest if the service genuinely warrants one.

saying these in an interview costs you the question

  • Tuning the threshold percentage until the noise stops
  • Assuming a low error ratio means low budget impact
  • Treating a five-minute window as inherently safe
  • Ignoring request volume when reading an error ratio
  • Believing burn-rate alerting solves low-traffic noise

context