In a rolling update you can configure how many instances may run above the desired count (surge) and how many of the desired count may be missing at once (unavailable). What do these two budgets trade off, and how does each extreme setting fail?
answer
- floor and ceiling on the fleet
- surge costs resources, unavailable costs capacity
- zero surge means kill before create
- zero unavailable needs real headroom
- rollback runs at the same speed
basics
~20 sSurge buys capacity with spare resources: extra instances start before old ones stop. Unavailable buys speed with reduced capacity: old instances stop first. Surge zero degrades a busy service mid-rollout; unavailable zero stalls the rollout when there is no headroom to schedule the extras.
solid answer
~50 sThe two budgets set the floor and the ceiling of the fleet during a replacement. Unavailable is the floor — with a desired count of twenty and an unavailable budget of five, at least fifteen instances are always serving. Surge is the ceiling — a surge budget of five means at most twenty-five exist at once, so the extra five cost real CPU and memory while the rollout runs. Set surge to zero and the controller must kill before it creates, so the fleet runs below target for the whole rollout; on a service already near saturation that is a self-inflicted brownout. Set unavailable to zero and you have promised full capacity, which only works if the cluster can actually schedule the surge instances — if it cannot, the new ones sit pending and the rollout stalls indefinitely with the old version still live. Both budgets at their smallest is correct but slow, which stretches the mixed-version window and your time to roll back.
go deeper
Know that a rolling update replaces instances in batches rather than all at once, and that the two settings control how many extras may run and how many may be missing while it happens.
Do the arithmetic out loud: desired count minus the unavailable budget is the guaranteed floor, desired plus surge is the ceiling. Explain that surge spends resources to protect capacity and unavailable spends capacity to avoid needing resources.
Show the failure modes you have actually seen: a stalled rollout because surge instances could not be scheduled, a brownout because a busy service ran below its floor, and rollouts that dropped traffic because readiness meant 'process started'. Tie the budget to rollback time.
Own the platform default and the exceptions to it. Decide how much standing headroom the estate carries so surge always has somewhere to go, and set the expectation that rollout duration is a rollback SLO, not a deployment convenience.
## The two budgets define a corridor A rolling update replaces a fleet piece by piece. The two budgets tell the controller how far it may stray from the desired count in each direction while it does so: ``` desired = 20, surge = 5, unavailable = 5 floor: 20 - 5 = 15 instances always available ceiling: 20 + 5 = 25 instances may exist at once ``` The floor is a promise to your users: capacity never drops below it. The ceiling is a bill: those extra instances consume real CPU, memory, IP addresses, database connections and licence seats for the duration of the rollout. Every rolling update is a negotiation between the two, and the negotiation is settled with spare capacity — that is the resource the whole strategy runs on. ## Surge set to zero With no surge allowance the controller cannot create before it destroys. It must take instances out of service first, then replace them. The fleet therefore runs below the desired count for the entire rollout — best case by one instance, typically by whatever the unavailable budget allows. For a service running at 30% utilisation, nobody notices. For a service running at 75% utilisation with a 25% unavailable budget, you have just removed the headroom the service was using to absorb its own variance. Latency climbs, queues build, timeouts start firing upstream, and — the part that catches people — autoscaling may react to the load spike by trying to add instances in the middle of a rollout, which the controller is simultaneously trying to converge. The release looks like it caused a performance regression when what it actually caused was a capacity shortfall. ## Unavailable set to zero This is the safe-looking setting, and it is the one that stalls. Promising that no instance may be missing forces the controller to start a replacement and wait for it to become healthy before it removes anything. That requires the platform to find room for the surge instances right now. If the cluster is full, the node pool is at its scaling ceiling, the quota is exhausted, an IP range is exhausted, or a pod anti-affinity rule leaves nowhere to place them, the new instances never start. The rollout does not fail loudly — it sits there, half-converged or not started at all, with the old version still serving. Teams discover this hours later when they wonder why the fix is not live. The same setting also multiplies downstream pressure. Twenty-five instances instead of twenty means 25% more database connections, more open file handles, more licence consumption, and more concurrent calls to a rate-limited partner API — briefly, but exactly while you are also introducing new code. ## Both budgets small Surge of one and unavailable of zero is the most conservative correct configuration and it is often the wrong one. It replaces instances one at a time, so a fleet of sixty takes sixty sequential health-check cycles. Two consequences matter more than the wall-clock time: - **The mixed-version window is long.** Old and new code serve side by side for the entire rollout, which means every request must be safe under both, and any schema or contract change must be backward compatible for far longer than anyone pictured. - **Your rollback is slow too.** Rolling back is another rolling update, subject to the same budgets. If it takes forty minutes to roll forward, it takes about forty minutes to roll back — which is the number that actually matters during an incident. ## Both budgets large A surge equal to the desired count is effectively a blue-green switch performed by the rolling controller: a whole second fleet appears, then the old one drains. A large unavailable budget is the opposite — it replaces most of the fleet at once, which is fast and cheap and removes the blast-radius protection that made you choose rolling in the first place. If half the fleet can be new before anyone looks at a metric, a rolling update has stopped being a progressive rollout. ## The budgets are only as real as the readiness signal All of this arithmetic assumes the controller can tell when a new instance is genuinely serving. If the health signal is a process that has started or a TCP port that is open, instances are counted as available while they are still loading configuration, filling connection pools, or compiling hot paths. The floor you configured is then fictional: you promised fifteen serving instances and delivered fifteen instances of which several return errors. Getting the readiness signal to mean *ready to take a real request* is what makes the budget mean anything, and it is the single most common reason a correctly configured rolling update still drops traffic. ## Choosing Start from the service's utilisation and its spare capacity. If there is headroom, prefer surge over unavailable: pay resources, keep capacity. If there is no headroom, either accept a small unavailable budget with a plan for the load, or pre-scale the fleet before the rollout so the surge has somewhere to go. Then size both so that the total rollout time is a rollback time you are willing to live with.
- A rolling update has been stuck at half-converged for an hour with no errors reported. What do you check first?Whether the replacement instances can actually be placed and can actually become ready. Two usual causes: no room for the surge instances — cluster full, quota or IP exhaustion, a scaling ceiling, an unsatisfiable placement rule — so they never start; or they start but never pass the readiness check, so the controller correctly refuses to remove any more old instances. Both look identical from the outside: old version still serving, rollout not progressing.
- How do the surge and unavailable budgets affect how long a bad release is live?Directly, in both directions. A slow rollout limits exposure while it runs but delays full deployment; more importantly, the rollback is another rolling update governed by the same budgets, so a forty-minute roll-forward is roughly a forty-minute roll-back. Size the budgets against the rollback time you can tolerate during an incident, not just against the deploy time.
- Why can a rolling update drop requests even with a generous unavailable budget?Two reasons. If readiness means only that the process started, instances are counted as available before they can serve, so the promised floor is fictional. And if terminating instances are killed without draining — removed from the load balancer and given time to finish in-flight work — every request already in progress on them fails, regardless of how many other instances are up.
saying these in an interview costs you the question
- Setting unavailable to zero makes a rolling update safe
- Surge is free because the extra instances are temporary
- A rolling update never reduces capacity
- The rollout is done when the controller reports the new count
- Rollback is instant because the old version is still around