skip to content

You own the alerting standard for a platform of several hundred services, some with 30-day SLO windows and some with 7-day windows. Would you mandate a single burn-rate alert configuration for all of them? Explain what must vary per service and what you would hold fixed.

level: principalimportance: nice to knowfreq 26%

answer

  1. standardise the policy, generate the numbers
  2. burn rate already normalises the target
  3. the multiplier depends on the window length
  4. seven-day window means three point three six
  5. low traffic needs an event-count guard

basics

~20 s

Hold the budget-spend policy fixed, not the multipliers. Burn rate already normalises across different SLO targets, but the multipliers are derived from the window length, so a 7-day SLO needs 3.36x where a 30-day SLO needs 14.4x, and low-traffic services need an extra event-count guard.

solid answer

~50 s

I would standardise the policy and generate the numbers, not standardise the numbers. What transfers cleanly is the budget-spend decision — 'we page when an hour's burn would cost 2% of the budget' — and because burn rate divides by `1 - target`, a 99.9% and a 99.99% service can share that policy untouched. What does not transfer is the window: 2% of a 7-day budget in one hour is a multiplier of 3.36, not 14.4, so copying the canonical table onto a weekly SLO silently makes the page nearly five times less sensitive, letting about 8.6% of the budget go before anyone is woken. Low-traffic services need a minimum bad-event count or a wider window before any ratio is trustworthy, and non-request SLIs like freshness need a different denominator entirely. So: fixed policy and a generator, per-service inputs of target, window and traffic floor, and a documented exception path — because a mandate teams cannot adjust gets routed to a channel nobody reads.

code

python · 10 lines
python
def tiers(slo_window_days):
    w = slo_window_days * 24.0
    policy = [(0.02, 1.0), (0.05, 6.0), (0.10, 72.0)]  # budget fraction, long window (h)
    for frac, long_h in policy:
        yield frac * w / long_h, long_h, long_h / 12.0

for days in (30, 7):
    print(days, [f"{m:.2f}x/{lh:g}h+{sh:.2f}h" for m, lh, sh in tiers(days)])
# 30 ['14.40x/1h+0.08h', '6.00x/6h+0.50h', '1.00x/72h+6.00h']
# 7  ['3.36x/1h+0.08h', '1.40x/6h+0.50h', '0.23x/72h+6.00h']

go deeper

for a junior

Know that burn-rate thresholds are derived from the SLO window and cannot simply be copied between services with different windows, and that very low-traffic services need special handling.

for a middle

Be able to recompute a multiplier for a different SLO window and show what breaks when the canonical 30-day numbers are pasted onto a 7-day objective.

for a senior

Demonstrate that you separate the portable part from the per-service part — the spend policy travels, the window and traffic floor do not — and that you know the low-traffic mitigations rather than just naming the problem.

for a principal

Own the tradeoff explicitly: a rigid mandate gets routed around and leaves phantom coverage, full autonomy leaves an unauditable estate. Argue for generated thresholds from declared inputs plus an exception log that doubles as the feedback signal on the standard itself.

## The real question: what is actually portable? Uniformity in alerting is valuable — it makes pages reviewable, comparable and teachable across an estate, and it stops every team inventing thresholds by feel. But mandating the wrong layer produces alerts that are uniformly wrong. The design task is to identify which part of a burn-rate configuration is genuinely service-independent. ## Portable: the budget-spend policy The organisational decision underneath the canonical table is not "14.4". It is *how much of the error budget the company is willing to lose before interrupting a human*: 2% in an hour for the fast page, 5% in six hours for the slow one, 10% over three days for a ticket. That statement is about interrupt tolerance and it applies across an estate. It also survives differences in target, because burn rate is normalised. A 14.4x condition means 1.44% bad events on a 99.9% service and 0.144% on a 99.99% one, computed automatically. This is exactly why burn rate makes a fleet-wide standard possible at all — a fleet-wide *error-percentage* threshold could never be right for both. ## Not portable: the SLO window Here is where a naive mandate breaks. Budget spent is `rate * interval / window`, so the multiplier that corresponds to a given budget fraction depends on the window length: ``` multiplier = budget_fraction * window_hours / long_window_hours 30-day window: 0.02 * 720 / 1 = 14.40x 7-day window: 0.02 * 168 / 1 = 3.36x ``` A team that copies 14.4x onto a 7-day SLO gets an alert that requires `14.4 * 1 / 168 = 8.6%` of the budget to be spent within an hour before it fires — more than four times the intended tolerance. The alert still looks canonical in review, and it is meaningfully deaf. The inverse mistake, copying a short-window multiplier onto a longer SLO window, produces a page that fires on trivially small budget loss. ## Not portable: traffic volume Burn rate is a ratio, and ratios need events. A service serving a few requests a minute produces a short-window ratio quantised so coarsely that one failure trips any multiplier. The standard therefore needs a traffic-dependent guard: a minimum absolute count of bad events ANDed with the burn condition, a longer long window for low-volume tiers, aggregation of small services into a shared SLO, or synthetic probes to guarantee a known request volume. Deciding which of those applies is a per-service input, and pretending otherwise gives low-traffic teams an alert that only ever cries wolf. ## Not portable: the shape of the SLI Availability and latency SLIs are good/valid event ratios and slot straight into the arithmetic. Freshness, correctness and durability objectives often are not per-request, and a pipeline whose SLI is "data no older than 15 minutes" has a denominator that is time or batches, not requests. Those need their own derivation rather than a table lookup. ## The design I would ship A generator, not a document. Teams declare the inputs — SLI definition, target, window length, expected request volume — and the platform emits the alert tiers from the fixed budget-spend policy, including the paired short windows and any minimum-count guard. Three properties follow: every service's thresholds are derivable and therefore reviewable; changing the organisation's interrupt tolerance is one change in one place rather than several hundred; and a team that needs something different has to file an exception that names why, which is a far better artefact than a silently hand-edited threshold. ## The cost of getting the balance wrong Both extremes have a real price. A rigid mandate with no exception path is defeated in the way all rigid mandates are defeated: the team routes the page somewhere nobody watches, and now the platform believes it has coverage it does not have. Full autonomy produces several hundred bespoke configurations that no one can audit, no one can compare, and no one can fix centrally when the underlying policy changes. The generator plus exception path is the middle, and the exception log is itself the signal that tells you where the standard needs to change.

  • A team argues their service should page at 6x over one hour instead of 14.4x because it is business-critical. How do you respond?
    By translating it back into the policy: they are asking to be woken when 0.83% of the monthly budget is at risk rather than 2%. That is a legitimate request if the service's budget is small in absolute terms, but it is a decision about interrupt tolerance, so it belongs in the exception log with a named owner and a review date. Criticality is usually better expressed by tightening the SLO target than by tightening the alert against a loose target.
  • How would you detect that some team has quietly copied the 30-day multipliers onto a 7-day SLO?
    Make it structurally impossible rather than detectable — if thresholds are generated from declared inputs, the mismatch cannot be expressed. Where hand-written rules still exist, audit by recomputing the implied budget fraction for each alert (`rate * long_window / slo_window`) and flagging anything far from the policy values. An alert implying 8.6% budget spend before paging stands out immediately in that view.
  • What is the strongest argument against generating alerts centrally at all?
    That the team owning the service understands its failure modes better than the platform does, and that generated alerts arrive without the context that makes them actionable. The honest answer is that generation should cover the SLO-based paging tier only — the layer where the math is objectively derivable — and leave cause-level and service-specific alerting entirely to the owning team.

saying these in an interview costs you the question

  • Copying the 14.4x table onto every SLO window
  • Assuming a stricter SLO target needs a higher multiplier
  • Mandating thresholds with no exception path
  • Applying request-ratio burn rates to freshness objectives
  • Treating low-traffic services as a rounding error

context