A service ends every quarter with more than 90% of its error budget unspent. What does that tell you, and what would you actually change?
answer
- a green budget is not automatically good
- unspent budget is unbought velocity
- does user pain track the budget at all?
- tighten the target, or spend the slack
basics
~20 sA permanently unspent error budget is not a win. It usually means the target is looser than what users notice, or the team is buying reliability nobody asked for. Tighten the SLO, or deliberately spend the budget on velocity.
solid answer
~50 sAn untouched budget means the mechanism is not doing its job, because it never influences a decision. There are three explanations and they call for different actions. The target may be looser than users' real expectations — the tell is complaints and support tickets while the budget reads green, and the fix is to tighten the SLO. The team may be over-buying reliability: extra redundancy, slow cautious releases, manual gates on every change. There the budget is permission to go faster — larger rollout steps, fewer manual approvals, retire a costly standby tier. Or the SLI may simply not measure what breaks, in which case a green budget during real incidents is measurement failure, not health. I'd distinguish them by checking whether user-reported pain correlates with budget spend at all. And I'd size the change carefully: moving 99.9% to 99.99% cuts the monthly budget from about 43 minutes to about 4.
go deeper
Understand that the error budget is meant to be used, and that a budget which is never touched is a sign something should change rather than proof the service is excellent.
Give the competing explanations — target too loose, reliability over-bought, SLI measuring the wrong thing — and show you know each nine costs ten times the previous one.
Show how you would settle it with evidence: correlating user-reported pain with budget spend, then naming the specific velocity or cost levers you would relax with the slack.
Own the strategic angle. Be ready to argue what reliability is worth to the business at the margin, and to explain how a chronically over-reliable service teaches its callers to stop handling its failures.
## Why an unspent budget is a finding, not a trophy The budget exists to make a decision tractable: how much risk may we take this window? A budget that is never meaningfully consumed never changes anyone's behaviour, so the organisation is carrying the cost of the whole apparatus and getting no decision out of it. Worse, the cost is usually paid in velocity — cautious releases, extra approvals, redundancy that nobody sized against a requirement. "We were 99.997% against a 99.9% target, every quarter, for a year" is a prompt to investigate, not to celebrate. ## The three explanations, and how to tell them apart **1. The target is looser than the users' expectations.** This is the common case and the important one. Your budget says everything is fine while the support queue fills with reports of slow checkouts. The diagnostic is correlation: pull the last four quarters of user-visible complaints, escalations from account managers, and abandoned-session or retry metrics, and see whether any of them line up with budget spend. If pain happens while the budget is green, the number is not tracking user experience and the target (or the SLI behind it) is wrong. Action: tighten the target — but price it first. Each additional nine divides the budget by ten. Going from 99.9% to 99.99% on a 30-day window takes the allowance from about 43 minutes to about 4.3. That is usually not a config change; it is multi-zone redundancy, automated failover, faster rollback, and a stricter on-call posture. Tighten by the smallest step that reflects reality (99.9% to 99.95%, say) rather than jumping two nines because it sounds better. **2. The team is over-buying reliability.** The service genuinely is far more reliable than it needs to be, and the money is going somewhere: a hot standby region nobody has ever needed, a release process with three manual gates, a rollout that takes four days to reach 100%. Here the budget is a mandate: it is telling you that you have slack and are refusing to spend it. Action: convert the slack into velocity or cost savings deliberately. Increase rollout step sizes and shorten bake times, remove a manual approval, ship the migration you have been deferring, run the chaos experiment you have been avoiding, or drop a redundancy tier and take the savings. Each of these is a decision with a cost, made with evidence, which is exactly what the budget is for. **3. The SLI does not measure what actually breaks.** The budget is green during real incidents because the measurement misses them — it watches a health-check endpoint, or aggregates so broadly that a failing journey disappears into healthy traffic, or counts a fast error response as a success. Action here is to fix the measurement; the target is not the problem. (Which signals to measure and where to measure them is its own design exercise — the point for this question is that a budget that stayed green through a known outage is evidence of a measurement defect, not of reliability.) ## The over-reliability trap for dependencies There is a second-order effect worth naming: when a service is chronically far more reliable than it promises, its callers stop treating it as fallible. They drop the retry, remove the fallback path, and build assumptions on the observed behaviour rather than the published target. Then the first time the service uses its budget legitimately, the dependents break harder than they should. Google's SRE book describes deliberately taking the Chubby lock service down when it overperformed its SLO, precisely so that dependent teams would keep designing for the promised level rather than the observed one. You do not have to adopt planned outages, but you should recognise the failure mode and at minimum test dependents against the target you actually promise, for instance during a game day. ## What not to do **Do not carry the budget over.** Unspent budget does not roll into the next window; the whole point is that the window bounds the decision. A team arguing for accumulated credit has stopped treating the budget as a control and started treating it as currency. **Do not manufacture failures to burn it.** Injecting errors for the sake of consuming budget is theatre and harms real users. Spending is a by-product of taking useful risk — shipping, migrating, testing — not a goal. **Do not tighten reflexively.** Ratcheting the target every time it is comfortably met eventually produces a number the service cannot hold, and then the budget is permanently empty and equally useless in the other direction. Change the target when evidence about user expectations says to, and change it at a review, with the new number stated as clearly as the old one. The answer that lands in an interview: name that a green budget is a signal to investigate, give the three explanations, say how you would distinguish them with evidence, and state what you would spend the slack on.
- How would you decide how much to tighten the target by?Work backwards from user tolerance and price each step. Find the reliability level at which users start complaining or abandoning, then pick the smallest target above it that you can actually hold. Remember each nine divides the budget by ten, so 99.9% to 99.99% moves you from roughly 43 minutes a month to roughly four — usually a redundancy and failover project, not a tuning exercise.
- Users complain constantly but the error budget has never been touched. Where do you look first?At the SLI, not the target. A budget that stays green through pain is measuring the wrong thing or measuring it in the wrong place — probing a health endpoint rather than a real journey, aggregating a failing segment into healthy traffic, or counting fast failures as successes. Fix what is measured before arguing about what number it should hit.
- Can unspent budget be carried into the next window?No. The budget is bounded by its window by design: it answers "how much risk this window", and carry-over turns it into a savings account that would let a team bank quiet months and then spend them on a reckless launch. If the budget is routinely unspent, the correct response is to change the target or spend the slack now, not to accumulate it.
saying these in an interview costs you the question
- Treating a permanently unspent budget as a success to protect
- Jumping straight from 99.9% to 99.99% without pricing it
- Assuming a green budget proves users are happy
- Wanting to roll unspent budget into the next window
- Injecting failures purely to consume the remaining budget