Two teams both report a 99.9% availability SLO, but one tracks its error budget in bad minutes and the other in failed requests. How does that choice change what the budget shows during a partial outage, and how would you pick?
answer
- same target, two ways to count
- bad minutes versus bad requests
- a threshold turns partial into all-or-nothing
- request-based weights an incident by traffic
basics
~20 sTime-based budgets count bad minutes and weight every minute equally; request-based budgets count failed requests and weight an incident by the traffic it hit. Partial outages and off-peak incidents cost very differently under the two.
solid answer
~50 sA time-based budget marks each minute good or bad against a threshold and spends whole minutes. A request-based budget divides bad events by valid events and spends in proportion to the requests actually harmed. During a partial outage where 5% of requests fail for an hour, the request-based budget records 5% of that hour's traffic, while a bad-minute count records either the entire hour or nothing at all, depending on where the threshold sits. Time-based counting also makes a 3 a.m. incident cost exactly as much as one at peak, which does not match user harm. I default to request-based for services with steady, substantial traffic, because it tracks blast radius. I use time-based for low or spiky traffic, where a handful of requests swings the ratio into noise, and for anything that has to line up with a contractual availability figure quoted in minutes.
go deeper
Know that an error budget can be counted either as failed requests or as bad minutes, and that the two are not interchangeable. Be able to say which one a given dashboard is showing.
Work a partial outage through both bases out loud, and explain how the per-minute threshold turns a continuous failure rate into an all-or-nothing charge.
Demonstrate the operational consequence: which basis the budget policy is wired to, why traffic weighting can hide a regional failure, and how you would slice the SLI to fix it.
Own consistency across the fleet. Be ready to argue for one accounting standard so services are comparable, and to explain to a commercial stakeholder why the engineering number and the contractual minutes will never match exactly.
## Two ways to spend the same budget The target — 99.9% — says nothing about how consumption is counted. Two accounting bases are in common use, and they produce different numbers from the same incident. **Event- or request-based.** The SLI is `good_events / valid_events`. The budget is a count of bad events: 0.1% of everything served in the window. Every failed request draws down the account by exactly one unit. **Time-based (bad minutes).** Each fixed slice of time — typically a minute — is classified good or bad by evaluating a condition over that slice, for example "fewer than 99% of requests in this minute succeeded". The budget is a count of bad slices: 43 bad minutes out of 43,200. ## What a partial outage does to each This is where they diverge, and it is the heart of the question. Suppose a bad deploy makes 5% of requests fail for a full hour. Request-based: you spent 5% of an hour's traffic. If the service handles a tenth of its monthly requests per day, this is a modest, precisely-proportional bite out of the budget. Time-based: it depends entirely on the per-minute threshold. If a minute is "bad" when success drops below 99%, all 60 minutes are bad — you just spent 60 of your 43 budgeted minutes and blew through the window from a failure that touched one request in twenty. If the threshold sits at 95%, the same hour costs nothing at all, and a real user-visible degradation is invisible in the budget. That cliff is the defining weakness of bad-minute accounting: it converts a continuous quantity (fraction of requests harmed) into a binary one, and the threshold you pick decides everything. ## The traffic-weighting difference Request-based counting weights incidents by traffic automatically. A ten-minute total outage at 3 a.m. on a consumer service might harm a few thousand people; the same ten minutes at peak might harm a million. The request-based budget charges you roughly 200x more for the second, which is what actually happened to users. A bad-minute budget charges the same ten minutes for both. That cuts both ways. Traffic weighting means a quiet region's total outage barely registers if the global ratio is what you track, which can hide a complete failure for one customer segment behind healthy traffic elsewhere. The standard mitigations are to slice the SLI (per region, per tier, per large customer) or to track a "worst window" variant — but note that choosing what to slice by is an SLI design decision, not a budget-arithmetic one. ## Where each one fits Prefer **request-based** when: - traffic is high and reasonably continuous, so ratios are statistically stable - the user harm really is proportional to requests affected - you want partial degradations to cost partially rather than all-or-nothing Prefer **time-based** when: - traffic is low or bursty — at 200 requests an hour, two failures put you at 1% and the budget becomes noise - the thing being measured is not request-shaped at all: a batch pipeline's freshness, a queue's drain, a scheduled job that either ran or did not - you must reconcile with a contractual availability number, since commercial availability commitments are almost always expressed as minutes of downtime per month A very common practical answer is to run both: request-based for the engineering budget that gates releases, time-based for the number quoted to customers, and to accept that they will not agree during partial failures. If you do that, be explicit about which one the budget policy is wired to — otherwise a freeze can be triggered by whichever number happens to look worse. ## Second-order effects worth naming **Gaming.** Because a bad-minute threshold is a cliff, a team can tune the threshold rather than the service. Any change to the threshold should be as visible as a change to the target itself. **Comparability.** Two services quoting 99.9% on different bases are not comparable. If you run a fleet-wide reliability review, normalise the basis first or you will draw the wrong conclusions. **Aggregation.** Request-based ratios compose reasonably across shards and regions by summing numerators and denominators. Bad-minute counts do not compose cleanly — a minute bad in one region and good in three is a judgement call you have to define in advance. The interview answer that lands is not "request-based is better". It is: name what each does to a partial outage, name the traffic-weighting consequence, and pick based on the traffic shape and on what the number has to be reconciled with.
- Your service is request-based and a whole region went dark for an hour, but the global budget barely moved. What went wrong?Nothing went wrong with the arithmetic — a global ratio hides a local failure when the failing slice is a small share of traffic. The fix is to stop reporting one global number: evaluate the same SLI per region or per customer tier so a total failure of any slice consumes that slice's own budget. A global-only budget systematically under-reports concentrated harm.
- How would you set the per-minute threshold for a bad-minute budget?Derive it from user harm, not from convenience. Pick the success rate at which a typical user notices — often a few percent of failures on an interactive path — and be aware that you are choosing a cliff. Then check it against history: replay past incidents and see whether the threshold classified them the way the incident reviews did. Changing the threshold later is a change to the SLO and should be reviewed as one.
- Can the same incident exhaust one basis and leave the other healthy?Yes, routinely. A long, shallow degradation burns almost no request budget but can burn every bad minute; a short, total outage at peak burns enormous request budget while consuming only a few minutes. That is why the budget policy has to name exactly one basis as the one that triggers consequences.
saying these in an interview costs you the question
- Assuming both bases give the same number for the same incident
- Setting the bad-minute threshold arbitrarily and never revisiting it
- Believing a global success ratio catches a single region's total outage
- Using a request ratio on a service with a handful of requests per hour
- Comparing two services' nines without checking they count the same way