skip to content

A cloud zone failure took your service down for an afternoon and burned 80% of its quarterly error budget, while your own code behaved perfectly. Should that spend be exempted from the budget, and what does your answer commit you to?

level: principalimportance: nice to knowfreq 33%

answer

  1. users were down either way
  2. measurement honest, consequence tailored
  3. a budget is not a blame ledger
  4. the dependency chain caps the target

basics

~10 s

Count it. An error budget measures user experience, not team fault, and exempting dependency failures hides the case for removing the dependency. Adjust the policy response to the cause instead of adjusting the measurement.

solid answer

~50 s

My default is that it counts, because the users were down and the budget records what they experienced rather than who is to blame. The moment you subtract failures you did not cause, the number stops describing the service and starts describing the team, and you lose the strongest evidence you will ever have for funding multi-zone work. But I'd separate measurement from consequence. The measurement stays honest; the policy response should fit the cause — freezing your own feature work does nothing about a provider outage, so the right consequence is that the reliability investment becomes removing the single-zone dependency, funded now. I'd also check whether the target was ever achievable: if you promise 99.99% on a single-zone deployment, the budget just reported an impossible commitment. If an exemption is granted anyway, it should be counted, logged and approved outside the team that benefits.

go deeper

for a junior

Know that the error budget records what users experienced, so an outage caused by a cloud provider still counts against it even when your own code was faultless.

for a middle

Explain why exemptions are corrosive: they turn the budget into a blame ledger, break comparability between services, and erase the evidence that would justify redundancy work.

for a senior

Separate measurement from consequence out loud. Keep the spend, then argue that the policy response for an external cause is funding removal of the dependency rather than freezing your own feature work.

for a principal

Own the ceiling argument and the governance. Be ready to show that serial dependencies multiply into a hard limit on any promise, and to define who may grant an exemption, how it is counted, and how both raw and adjusted numbers get reported.

## The question behind the question An interviewer asking this is testing whether you understand what the error budget is *for*. If you believe it is a scorecard for the team, exempting an external failure is obviously right. If you understand it as a measurement of what users got, exempting it is obviously wrong. Both instincts are defensible in isolation; the senior answer is that the two purposes have been collapsed together and need separating. ## The case for counting it **The users were down.** They did not experience a well-run service unlucky in its provider; they experienced an outage. An SLI is a statement about the service as delivered, and the dependency is part of the service as delivered. You chose it, you deployed into it, and you did not build around it. **Exemptions destroy the signal you most need.** The 80% spend is the single most persuasive artefact you will ever hold for funding a multi-zone architecture. Subtract it and, on paper, nothing happened — and next quarter you will be arguing for redundancy with an anecdote instead of a number. **A budget with exemptions stops being comparable.** Once every team subtracts the incidents it considers unfair, the fleet's reliability numbers stop meaning the same thing, and the reviews built on them become theatre. **Blame creeps in.** "Was it our fault?" is a postmortem-adjacent question and a corrosive one to bolt onto an accounting mechanism. The budget should be one of the few numbers in the organisation that nobody argues about the attribution of. ## The honest counter-argument The objection is real and you should voice it: if the standard consequence is a feature freeze, then counting the spend punishes the team with a remedy that cannot possibly help. Freezing your feature releases does not make the provider's zone come back, and the team correctly perceives it as arbitrary. That perception is how budget policies lose legitimacy. But notice the defect is in the *response*, not the *accounting*. The fix is to make the policy cause-aware: - Budget exhausted mainly by your own changes -> freeze feature work, fix release safety. - Budget exhausted by a dependency you could have survived -> the reliability work is removing that dependency, funded and scheduled now, and feature work bends around it. - Budget exhausted by an event nothing available could have survived -> escalate to a target conversation, because the promise exceeds the architecture. That third branch is the one people miss. **A service cannot promise more reliability than its dependency chain delivers.** Serially dependent components multiply: three dependencies at 99.9% each put a ceiling near 99.7% before your own code has failed once. If you committed to 99.99% on a single-zone deployment, the budget did not misfire — it correctly reported an unachievable target, and the honest outcome is either to buy the redundancy or to lower the number publicly. ## If you do grant an exemption Some organisations do carve out categories, and there is a defensible version. Make it meet all of these conditions: - **Defined in advance.** The categories are written into the budget policy before any incident, not invented afterwards. - **Decided outside the beneficiary.** The team whose budget is restored does not get to approve the restoration. - **Counted.** Exemptions are a countable, reported quantity. Three exemptions in a quarter is itself a finding. - **Narrow.** "Announced maintenance windows" is a narrow, user-informed carve-out. "Anything upstream of us" is not — on a modern stack that is most of the failure surface. - **Reported alongside the raw number.** Publish both: the budget as measured, and the budget after exemptions. If they diverge much, the exemptions have become the story. Commercial SLAs routinely exclude force majeure and announced maintenance, and people cite that as precedent. It is a poor one: an SLA is a liability instrument written by lawyers to bound payouts, while an internal SLO is an engineering control written to drive decisions. They optimise for different things, and copying the exclusions across imports the wrong incentives. ## How to answer in the room Say it counts, and say why in one line: the budget measures experience, not fault. Then immediately show the second move, because that is what distinguishes the answer — the consequence should follow the cause, and the reliability work this incident has just funded is removing the single-zone dependency. Close with the ceiling argument: check whether the target was ever compatible with the architecture, because if it was not, the exemption debate is a distraction from a promise that needs rewriting.

  • How would you set a realistic target for a service sitting on a dependency that promises 99.9%?
    Start from the chain, not from ambition. Serial dependencies multiply, so three components at 99.9% cap you near 99.7% before your own failures. Either promise below that ceiling, or change the architecture — redundant instances of the dependency, a degraded-mode fallback, or caching that survives its absence. Promising above the ceiling guarantees a budget that is always empty and a policy nobody respects.
  • Should announced maintenance windows be excluded from the SLI?
    It is the one carve-out that is genuinely defensible, because users were told in advance and could plan. Keep it narrow: announced, bounded, and published. Then count how much of your reliability story depends on it — if a meaningful share of would-be budget spend sits inside maintenance windows, your users' experienced reliability is materially worse than your number claims.
  • Your provider publishes credits for the outage. Does that change the budget treatment?
    No. Service credits settle a commercial liability; they do not restore the requests your users lost. Treat them as an unrelated finance matter and leave the SLI alone. If anything, credits are further evidence for the reliability investment, because they quantify how much the dependency's failure was worth to somebody.

saying these in an interview costs you the question

  • Exempting any incident whose root cause was outside the team
  • Treating the error budget as a scorecard of team fault
  • Promising more nines than the dependency chain can deliver
  • Copying an SLA's force-majeure exclusions into an internal SLO
  • Granting exemptions informally, uncounted, by the team that benefits

context