skip to content

How much pre-provisioned headroom should a platform team mandate for control-plane outages, and how do you justify paying for it?

level: principalimportance: should knowfreq 36%

answer

  1. insurance with a monthly premium
  2. size from window times demand
  3. per tier, never fleet-wide
  4. the buffer a workload already has
  5. mandate create-nothing, not a percentage

basics

~20 s

Size headroom to the demand a workload must absorb while it can create nothing, not to average utilisation, and set it per tier rather than fleet-wide. Cheaper than any percentage: mandate that every failover path completes without creating a resource.

solid answer

~50 s

Headroom is insurance with a monthly premium, so the honest question is what it buys and for whom. Size it from two numbers a team can actually defend: how long a create-nothing window plausibly lasts, and how much demand must be served through it. That yields different answers per tier - a synchronous, user-facing path with no buffer may justify standing warm capacity, while a batch fleet with a durable queue can often justify none, because its backlog is the buffer. The cheapest mandate is not a percentage at all: **require that every failover and mitigation path complete without creating a resource.** That is a design property, it costs nothing continuously, and it removes the dependency that headroom was compensating for. Justify the remaining spend against the cost of the delay it prevents, in that workload's own terms.

go deeper

for a junior

Recall that headroom is capacity bought and paid for in advance, because during a management-API incident no new capacity can be created at any price.

for a middle

Explain how a headroom figure is derived: an assumed window in which nothing can be created, multiplied by the demand that must be absorbed during it.

for a senior

Show that you would set it per workload shape, reserve and measure it rather than assume idle capacity counts, and walk a failover path step by step marking creates.

for a principal

Own the estate-wide trade: defend a recurring premium in cost review, prefer the create-nothing property as the enforceable standard, and state the planning window as an assumption the organisation revisits.

## Headroom is a standing premium, not a project Pre-provisioned headroom is capacity you pay for every hour in order to have it during the hours you cannot create any. Unlike most reliability work it does not end - it is a permanent line on the bill, it scales with the fleet, and it is the first thing challenged in a cost review. So a platform team that mandates it has to be able to defend it in the same sentence as what it prevents. It also has a competitor that is nearly free, and the honest answer usually mixes the two: **making the mitigation path stop depending on creates at all.** ## Sizing it from two defensible numbers A headroom figure that comes from a round percentage is indefensible in review. One that comes from these two inputs survives: 1. **The create-nothing window.** How long do you assume the platform may refuse new resources? This is a planning assumption you own, written down, not a figure read off a published commitment. Pick it, state it, and revisit it after every incident that tests it. 2. **The demand you must absorb in that window.** Not average load and not last month's peak, but what the workload must serve, or accumulate, between the incident starting and the point where degradation becomes unacceptable. Headroom is then the capacity that covers input two for the duration of input one, on top of whatever the fleet is running at the moment the window opens - which is why a conservative scale-in floor is part of the same decision. A fleet allowed to shrink to nothing overnight has no headroom at 03:00 no matter what the policy says. ## Not every workload gets the same number | Workload shape | Buffer it already has | Reasonable posture | |---|---|---| | Synchronous, user-facing path | Only the latency budget, then errors | Standing headroom, possibly a warm standby that needs no create to take over | | Queue-fed batch fleet | A durable backlog measured in hours | Often no standing headroom; conservative scale-in and deep retention instead | | Scheduled or periodic job | The schedule itself can slip | Usually none; document the acceptable slip | | Stateful tier with a standby | The standby, if promotion needs no provisioning | Headroom in the standby, plus a promotion path that creates nothing | The table is the actual deliverable of this decision. A single fleet-wide percentage is simultaneously too expensive for the batch fleets and too thin for the one path that has no buffer, and it teaches every team that the standard is arbitrary. ## Mandate the property, not the percentage If a platform team can enforce only one rule across many teams, it should be this one: **a failover or mitigation path must complete without creating a resource.** In practice that means the standby already exists, the alternate path is already registered, the configuration is already distributed, the credentials are already present, and the switch is a change to something live rather than a request for something new. This is a better mandate than a headroom number for three reasons: - It has **no recurring cost** beyond what the standby itself costs, so it survives cost review. - It is **checkable in a design document**, before anything is built, by reading the failover steps and marking each one create or change. - It **keeps working when the assumption is wrong.** A headroom figure sized for a short window is simply insufficient for a long one; a path that creates nothing is indifferent to the window's length. Headroom then covers only what the property cannot: genuine demand growth during the window. ## Justifying the spend Express the premium against the harm in that workload's own units, and let each team make the trade in public: - For the user-facing path: the cost of standing capacity per month against the revenue or obligation exposed by being unable to add capacity for the assumed window. - For the batch fleet: the cost of standing capacity against the cost of a backlog measured in hours, which is often genuinely small - and saying so is the point, because it stops the mandate from being cargo-culted everywhere. - Across the estate: the cost of the standbys required by the create-nothing rule, which is usually far smaller than uniform headroom and far more predictable. Two traps to name explicitly. First, headroom that exists on paper but is consumed by the fleet's own normal operation is not headroom; it must be reserved and measured, and its consumption alerted on. Second, a create-nothing failover path that nobody has ever exercised is an assumption, not a capability - the step everyone discovers under load is the one that turned out to provision something after all.

  • Why is a single fleet-wide headroom percentage a poor standard?
    Because workloads differ in the buffer they already hold. A queue-fed batch fleet converts excess demand into delay and often needs none; a synchronous path has only its latency budget and may need a great deal. One number overpays for the first, underprotects the second, and signals to every team that the standard is arbitrary.
  • How do you tell whether a failover path really creates nothing?
    Walk its steps and mark each one as a change to something that exists or a request for something new. Promotion, registration behind an entry point, address allocation, provisioning a replacement and reading boot configuration from a management endpoint all count as new. Anything marked new is a dependency on the half of the platform that fails first.
  • What stops reserved headroom from quietly disappearing?
    Measurement and an alert. Headroom that is merely un-utilised gets absorbed by normal growth, and the fleet reaches the incident with none. Reserve it explicitly, track the margin as its own signal, and treat a sustained fall in that margin as a capacity item rather than as healthy efficiency.

saying these in an interview costs you the question

  • Picks a round headroom percentage with no stated window
  • Applies one headroom number to every workload in the estate
  • Counts merely idle capacity as reserved headroom
  • Assumes a failover path creates nothing without walking it
  • Justifies the premium by citing a published availability figure