skip to content

Capacity plans usually size a fleet so peak load lands well below 100% of its capacity — often somewhere near 60%. Why leave that much unused, and what do you gain and lose by raising the target?

level: juniorimportance: should knowfreq 52%

answer

  1. headroom is insurance you pay for
  2. what has to fit in the gap?
  3. redundancy usually sets the ceiling
  4. capacity arrives in minutes, spikes in seconds
  5. 60% target means buying 1.67x peak

basics

~20 s

The gap absorbs what the plan cannot: a lost failure domain, forecast error, the minutes it takes to add capacity, and the extra load of a rollout or a retry surge. Raising the target cuts spend proportionally and removes that cushion.

solid answer

~50 s

Headroom is insurance with a price tag. It covers four things: surviving the loss of a failure domain, which usually sets the ceiling on its own; error in the demand forecast; the time it takes to add capacity, since even elastic capacity arrives in minutes while a spike arrives in seconds; and transient extra load from a rolling deploy taking instances out, a failover, or clients retrying. There is also the plain fact that response time degrades before a resource is fully saturated, so 100% is not a place you can operate anyway. The cost side is simple arithmetic: at a 60% target you buy about 1.67x your peak, and moving to 75% removes roughly a fifth of the fleet. The right target is the minimum of your redundancy ceiling and whatever you have measured as safe — and it should be tracked per resource, because a fleet at 55% CPU can still be out of connections.

code

python · 15 lines
python
import math

peak_qps = 6000
per_server_qps = 150

def fleet(target_util):
    return math.ceil(peak_qps / (per_server_qps * target_util))

for t in (0.50, 0.60, 0.67, 0.75, 0.90):
    print(f"target {t:.0%} -> {fleet(t)} servers")
# target 50% -> 80 servers
# target 60% -> 67 servers
# target 67% -> 60 servers   <- N+1 across 3 zones already caps you here
# target 75% -> 54 servers
# target 90% -> 45 servers

go deeper

for a junior

Be ready to say that the unused portion is deliberate insurance — for failures, forecast error and the delay in getting new capacity — and that raising the target saves money by removing that insurance.

for a middle

Explain the arithmetic in both directions: a 60% target means buying about 1.67x peak, and the redundancy ceiling from your failure-domain count usually decides the number before performance does.

for a senior

Show that you measure utilization per resource and per failure domain at the peak interval that matters, and that you set different targets for tiers with different failure costs.

for a principal

Own the spend-versus-risk position across the fleet: what company-wide default you set, where you allow exceptions, and whether you monetize headroom with preemptible work rather than shrinking it.

## Headroom is a purchase, not an accident Idle capacity looks like waste on a bill, and someone will eventually ask why the servers are only 60% busy. The answer is that the gap is doing four jobs, each of which has a failure attached if you remove it. **Redundancy.** If the service must keep serving peak after losing a zone or a cell, the arithmetic already caps utilization: with three equal zones you cannot exceed about two-thirds, because the survivors must absorb the third. For most multi-zone services this is the binding constraint and everything else is a smaller adjustment on top of it. **Forecast error.** The demand number is an estimate. Headroom is what turns a 20% miss into an ordinary week rather than an incident. **Time to acquire capacity.** Even fully elastic systems react rather than anticipate: a metric window, a provisioning call, instance boot, dependency registration, cache warm-up and JIT warm-up all cost minutes. A promotional spike or a push notification lands in seconds. Headroom is what serves traffic during the gap between demand arriving and capacity arriving. **Transient self-inflicted load.** A rolling update takes instances out of service while adding load to the rest. A failover moves a shard's traffic onto peers with cold caches. A blip causes clients to retry, briefly multiplying request volume. All of these consume headroom precisely when you least want to be near the limit. And separately from all four: response time rises well before a resource is fully saturated, so the practical operating ceiling is below 100% regardless of what the plan says. ## What raising the target buys The cost relationship is linear and easy to state. Fleet size is roughly `peak / (per_unit_capacity x target)`. So: ``` target 50% -> 2.00x peak purchased target 60% -> 1.67x peak purchased target 75% -> 1.33x peak purchased target 90% -> 1.11x peak purchased ``` Moving from 60% to 75% removes about 20% of the fleet — real money at scale. Moving from 60% to 90% removes a third. What you give up is proportionate: at 90% across three zones you cannot lose a zone, you cannot absorb a forecast miss, and a rolling deploy alone may push you into saturation. ## Choosing the number Start from the redundancy ceiling, because it is usually the strictest. Then subtract for forecast uncertainty (larger for a volatile or launch-heavy service, smaller for a steady one) and for how fast you can actually add capacity. Then check the number against measured behaviour — the point where response time starts climbing for this specific service — and take the minimum. Different tiers deserve different targets. A latency-critical user-facing tier and an asynchronous batch tier that can be late do not need the same cushion, and applying one company-wide number to both overspends on one and underprotects the other. ## Measure per resource and per domain "Utilization" is not one number. The binding resource is whichever one saturates first, and it is often not CPU: a connection pool, memory, file descriptors, a downstream quota, disk throughput, or a lock. A fleet reported at 55% CPU while its database connection pool is at 95% has no headroom at all, and the CPU graph will not say so. Granularity matters just as much. A 60% hourly average can contain minutes at 100%. Compute utilization against the peak interval you actually need to survive, not against a comfortable average. And measure per failure domain. Aggregate utilization across zones hides one zone at 90% when traffic is unevenly distributed, which is exactly the zone whose loss you were planning for. ## Making the headroom cheaper The honest way to reduce the cost of headroom without reducing the headroom is to fill it with work that can vanish instantly: batch jobs, reprocessing, model training, anything preemptible and sized so it can be evicted in seconds when foreground demand arrives. You are still holding the capacity; you are just not holding it idle. The discipline required is that the filler work must genuinely yield immediately — filler that takes ten minutes to drain is not filler, it is a second workload competing during your incident.

  • How do you make the headroom cheaper without giving it up?
    Fill it with preemptible work — batch, reprocessing, training — sized so it can be evicted within seconds when foreground demand arrives. You keep the capacity available for the failure you bought it for, but stop paying for it to sit idle. The requirement is that the filler yields immediately; work that takes ten minutes to drain is not filler.
  • Why can a fleet sitting at 55% CPU still be out of capacity?
    Because CPU is not necessarily the binding resource. Connection pools, memory, file descriptors, disk throughput, a lock, or a downstream service's quota can saturate first, and the CPU graph shows nothing. Track utilization per resource and per failure domain, and set the target against whichever one saturates first for that service.
  • Should every service in the company run the same utilization target?
    No. The target follows from each service's failure-domain count, its forecast volatility, how fast it can add capacity, and the cost of being short. A latency-critical user-facing tier and a batch tier that can finish late warrant very different cushions; a single company-wide number overspends on one and underprotects the other.

saying these in an interview costs you the question

  • Calling idle capacity pure waste and pushing utilization to 90%
  • Averaging CPU over an hour and calling it peak utilization
  • Assuming autoscaling replaces headroom with no delay
  • Using one utilization target for every service and every tier
  • Tracking only CPU while connections or memory saturate first

context