Your service exhausts its error budget in week two of the quarter. What should a written error budget policy already commit the team to at that moment, and what stops that commitment from being quietly ignored?
answer
- agreed before, never during
- the freeze is a means, not the point
- who signs it decides whether it holds
- overrides allowed, counted, and visible
basics
~20 sA budget policy, agreed in advance by engineering, SRE and product, states what happens at zero: feature releases stop while reliability and security fixes continue, reliability work is funded, and a single named owner may override — with every override logged and counted.
solid answer
~50 sThe policy has to exist before the budget runs out, because at the moment of exhaustion every incentive pushes towards renegotiating it. A usable one names four things. First, the consequence: typically a feature freeze where only reliability fixes, security patches and rollbacks ship. Second, what the freed capacity is spent on — a committed share of the next sprint on the reliability work the incidents actually pointed at, not a vague promise. Third, an exit criterion, so the freeze ends on evidence rather than on mood: budget recovered on the rolling window, or the named remediation landed. Fourth, an escalation path — a specific role who can override the freeze, with overrides counted and visible. What keeps it honest is that it was signed by product leadership too, and that the override is allowed but never invisible. A freeze nobody can override gets bypassed; one nobody can see gets abused.
go deeper
Know that exhausting the error budget is supposed to change what the team works on, and that the rule for what changes is written down beforehand rather than argued at the time.
Explain the mechanics: what a freeze covers, what still ships, where the freed capacity goes, and what condition ends it. Be able to state why blocking rollbacks would make things worse.
Show the judgement. Name the costs of the freeze — growing batch size, delayed commitments — argue when renegotiating the SLO beats freezing, and describe the exit criterion you would have written in advance.
Own the negotiation. Be ready to explain how you get product leadership to sign consequences that will one day cost them a launch, and how a counted, visible override keeps the policy usable instead of ceremonial.
## Why the policy has to pre-exist The error budget only does work if running out has a consequence. If exhaustion produces a conversation rather than an action, the budget is a dashboard, not a control. And the conversation held at the moment of exhaustion is the worst possible one: the quarter's commitments are already made, a launch date is already public, and the person arguing to keep shipping outranks the person carrying the pager. So the commitment is made in advance, in writing, and signed by the people who will later want to break it — engineering leadership, the SRE or platform lead, and the product owner. This structure is described in Google's SRE Workbook, which is where most organisations' policies are derived from. ## What the policy should actually say **The trigger.** Not just "budget exhausted" — say which SLO, on which window, measured on which basis, and whether the trigger is exhaustion or a stated remaining threshold (some policies act at 25% remaining, which gives the team room to respond before the cliff). **The consequence.** The canonical one is a feature freeze: changes that add user-visible functionality stop; changes that improve reliability, patch security, or roll back a bad release continue. Note what the freeze is not — it is not "stop deploying". Blocking all deploys makes the service *less* reliable, because you lose your fastest remediation path and you accumulate an ever-larger batch of untested change. **Where the capacity goes.** A freeze that produces idle engineers produces nothing. The policy should commit a share of team capacity to reliability work drawn from the incidents that spent the budget: the action items nobody funded, the missing automation, the retry storm nobody fixed. "Shift work to reliability" is the actual purpose; the freeze is only the mechanism that frees the hands. **The exit criterion.** Freezes without an exit are how the practice dies. Legitimate exits: budget recovers above a stated threshold on the rolling window; a named set of remediation items ships; a fixed review date arrives and the group formally re-decides. Write down which one applies. **The escalation and override.** Someone must be able to say "we ship anyway" — a regulatory deadline or a contractual launch is a real thing. Make that person specific and senior, and make each override a counted, logged event. A useful pattern is a small fixed number of overrides per period (sometimes called silver bullets): the count creates a natural budget on exceptions, and running out of them is itself a signal. ## The costs you must be able to name An interviewer is checking whether you understand that a freeze is expensive, not just virtuous. - **Batch size grows.** Two weeks of frozen change ship together when the freeze lifts. A larger, less-tested release is a riskier release — you may spend the recovered budget on the unfreeze. - **Business cost is real.** Delayed revenue features and missed commitments are not free, which is exactly why product leadership must sign the policy up front rather than be told about it later. - **The freeze may not be the right remedy.** If the budget was destroyed by a single provider outage, stopping your own feature work does nothing to prevent a repeat. The policy should distinguish causes and permit a different response. ## The other legitimate outcomes at zero A freeze is not the only answer, and a strong candidate names the alternatives: **Renegotiate the SLO.** If the target was set aspirationally and the service has never met it, the budget is reporting an unrealistic promise rather than a regression. Lowering the target is legitimate — but it must be a deliberate, reviewed decision with the affected users considered, made at a review and recorded, never a quiet edit during the crisis. The test: would you tell your customers the new number? **Change what the SLI measures.** Sometimes exhaustion reveals that the measurement counts things users do not care about — health-check traffic, a deprecated endpoint, retried requests counted twice. Fixing the measurement is valid; doing it because you dislike the answer is not, and the way you tell the difference is whether the change was proposed before the number went red. **Escalate the reliability investment.** Exhaustion in week two of a quarter is evidence the service cannot hold its target with its current architecture or staffing. That is a resourcing conversation, and the budget is the artefact that makes it fundable. ## What keeps it honest Three properties, in order of importance. It was agreed before it was needed. Its consequences were signed by the people who lose from them. And its exceptions are visible — allowed, but counted. A policy with an unusable override is bypassed within a quarter; a policy with an invisible one is bypassed permanently and nobody notices.
- During a budget freeze, which changes should still ship?Anything that reduces risk: reliability fixes, security patches, rollbacks and the remediation items from the incidents that spent the budget. Blocking all deploys is counterproductive — it removes your fastest mitigation path and grows the batch that eventually ships. The freeze targets new user-visible functionality, not the deploy pipeline itself.
- How does the freeze end?On a criterion written before it started: budget recovered above a stated threshold on the rolling window, a named set of remediation items landed, or a scheduled review at which the group formally re-decides. Ending it on mood or on pressure teaches everyone the policy is theatre, and the next exhaustion will be argued rather than acted on.
- A product director insists on shipping a contractually committed launch during the freeze. What do you do?Use the override path rather than pretending the policy does not apply. Record who authorised it, that the budget was exhausted at the time, and what mitigation accompanies the launch — a smaller rollout, a tested kill switch, extra staffing. An override that is granted and logged keeps the mechanism intact; one that is granted silently ends it.
- Isn't lowering the SLO just cheating?Only if it is done quietly and reactively. A target that has never been met is describing an aspiration, not a commitment, and a budget that is always empty stops driving any decision. Lowering it deliberately, at a review, with the user impact examined and the new number stated publicly, is honest. Editing it mid-incident to end a freeze is not.
saying these in an interview costs you the question
- Deciding the consequences after the budget has already run out
- Freezing all deploys, including rollbacks and reliability fixes
- Freezing with no exit criterion and no owner
- Quietly lowering the SLO target to end a freeze
- Treating a freeze as free — ignoring batch size and delayed commitments