A service has a 99.9% availability SLO measured over a 30-day rolling window. How much error budget does that give you, and what does it mean for a team to "spend" it?
answer
- what the target leaves over
- one minus the target, over a window
- 0.1 percent of 43,200 minutes
- about 43 minutes, or 0.1% of requests
basics
~20 sAn error budget is one minus the SLO target over the window. A 99.9% target across 30 days allows about 43 minutes of unavailability, or 0.1% of requests, before the objective is missed. Spending it means shipping change against that allowance.
solid answer
~50 sThe budget is the complement of the target: `1 - 0.999 = 0.001`. Over a 30-day window that is 0.1% of 43,200 minutes, so roughly **43 minutes** of unavailability; measured on requests instead, it is 0.1% of the requests served — 100,000 bad requests out of 100 million. Spending it means doing things that consume that allowance: shipping releases, migrating a database, running a load test in production, or simply having an incident. The point is that the budget is an allowance, not a limit you are supposed to hoard. If the budget is intact, the team has earned permission to ship faster; if it is gone, reliability work takes priority. On a rolling window the budget refills gradually as old failures age out, so there is no clean-slate date — an outage on the 1st still counts on the 29th.
code
python · 10 linesTARGET = 0.999
WINDOW_MINUTES = 30 * 24 * 60 # 43,200 minutes in a 30-day window
budget_minutes = (1 - TARGET) * WINDOW_MINUTES
print(f"budget: {budget_minutes:.1f} minutes per 30 days")
total_requests = 100_000_000
failed_requests = 62_000
budget_requests = (1 - TARGET) * total_requests
print(f"spent: {failed_requests / budget_requests:.0%} of the request budget")go deeper
Be able to do the sum out loud: one minus the target, times the window. Know that 99.9% over 30 days is roughly 43 minutes and that the budget is an allowance the team is allowed to use.
Explain both bases — bad minutes and bad requests — and what a rolling window does to refill. Be ready to say why each extra nine divides the budget by ten and what that costs in architecture.
Show how you actually use the number: gating a risky migration on remaining budget, translating a proposed target into minutes before agreeing to it, and refusing to reset the budget after a fix.
Own where the target comes from. Be ready to argue what reliability the business is willing to pay for, why the internal SLO sits tighter than the contractual promise, and how you would defend that gap to a customer-facing executive.
## What an error budget actually is An SLO states a reliability target — say, 99.9% of valid requests succeed over a 30-day window. The **error budget** is the complement of that target: the amount of unreliability the organisation has explicitly decided is acceptable. It is not a forecast of how much you will break, and it is not a tolerance for sloppiness. It is a stated, negotiated allowance, and the whole SRE mechanism rests on it: because the allowance is a number, arguments about "are we shipping too fast?" become arithmetic instead of opinion. ## The arithmetic Time-based form. A 30-day window contains 30 x 24 x 60 = 43,200 minutes. At a 99.9% target the budget is 0.1% of that: ``` 0.001 * 43200 = 43.2 minutes ``` The same target over different targets and windows: - 99% over 30 days -> 432 minutes (7h 12m) - 99.9% over 30 days -> 43.2 minutes - 99.95% over 30 days -> 21.6 minutes - 99.99% over 30 days -> 4.32 minutes - 99.9% over 365 days -> 525.6 minutes (about 8h 46m) Notice the shape: **every additional nine divides the budget by ten**. That is why "just add a nine" is never a free request — going from 99.9% to 99.99% takes you from three quarters of an hour of slack a month to four minutes, which usually means redundancy, automated failover and a different on-call posture. Event-based form. If the SLI is a ratio of good events to valid events, the budget is a count of bad events. Serving 100 million requests in the window at a 99.9% target gives 100,000 requests you may fail. Consumption is then `bad_events / allowed_bad_events`, which is the "42% of budget remaining" number on a dashboard. ## Tracking the spend A budget dashboard normally shows two things: how much is left, and how fast it is going. "Left" is the cumulative figure above. "How fast" is the rate at which it is being consumed relative to the window, which is what drives budget-based alerting rather than raw thresholds. For this question, the important part is the cumulative view: at any moment the team can say "we have consumed 61% of this window's budget" and act on it. On a **rolling** window, the budget refills continuously — an outage 31 days ago no longer counts, and one from yesterday will keep counting for another 29 days. There is no reset date, which is deliberate: it removes the temptation to wait out the calendar. ## What "spending" means in practice Spending is not only incidents. Every deliberate act that risks user-visible failure draws on the same account: - shipping a release, especially a large or infrequent one - a schema migration or a datacentre/zone migration - a load test or chaos experiment run against production - a planned failover drill - a dependency's outage that reaches your users That framing is the payoff. Product wants features; the people carrying the pager want stability. With a budget, the answer to "can we ship the risky thing this week?" is "we have 70% of the budget left, so yes" or "we are at 4%, so not until we recover" — a decision with a shared, checkable input. ## Common mistakes **Treating an unspent budget as a win.** A budget that is never touched means the team is either buying reliability the users did not ask for, or has a target looser than reality; either way it is paying in velocity for nothing. **Resetting the budget after an incident.** "We fixed the bug, so those minutes shouldn't count" destroys the entire mechanism. The users experienced the failure; the budget records experience, not blame. **Quoting a budget without stating its window and basis.** "We have 43 minutes" is meaningless without "per 30 days, time-based". A 43-minute annual budget and a 43-minute monthly budget are two very different services. **Confusing budget with SLA penalty.** The error budget is an internal engineering control. A contractual SLA usually sits at a looser target than the internal SLO precisely so the budget runs out — and triggers a response — well before money is owed. A final sanity check to keep in your head: 99.9% is roughly "about three quarters of an hour a month", 99.99% is "a few minutes a month". If somebody proposes a target, translate it into minutes out loud before agreeing to it.
- Your SLO is 99.9% but the contract with customers promises 99.5%. Why would anyone set the internal target higher than the contractual one?So the internal control fires long before money or reputation is at stake. The error budget is meant to run out while the consequence is still "stop shipping features and fix reliability", not "pay credits". The gap between the two targets is your reaction time: by the time the 99.9% budget is exhausted you still have room before the 99.5% commitment is breached.
- Does planned maintenance consume the error budget?It depends on a decision you have to make explicitly and write down. If users see errors during the window, the honest default is that it counts — they were affected. Many organisations carve out announced maintenance windows instead, which is defensible because users were warned, but every carve-out widens the gap between the number and lived experience, so keep them few, documented and counted.
- What changes if you move the same 99.9% target from a 30-day rolling window to a calendar quarter?The budget grows to roughly 130 minutes but arrives as one pot with a hard reset date. That creates end-of-period effects: a bad January can freeze work for weeks, and teams learn to wait for the reset. A rolling window refills gradually and keeps recent history in view, which is why it is the more common choice for engineering decisions.
saying these in an interview costs you the question
- Thinking the goal is to keep the error budget at 100%
- Quoting a budget with no window or measurement basis
- Believing an incident stops counting once the bug is fixed
- Assuming each extra nine costs about the same as the last
- Treating the error budget and the contractual SLA as the same number