skip to content

Three zones each run an order service at 70% of capacity at peak; one zone is lost — what happens, and what sizing rule prevents it?

level: middleimportance: must knowfreq 61%

answer

  1. size for what survives, not what exists
  2. three zones at 70% is 210%
  3. two survivors hold only 200%
  4. cap peak at (N-1)/N per zone
  5. headroom is idle capacity you buy

basics

~20 s

Each survivor is asked for about 105% of its own capacity, so the service queues, slows and sheds load at the worst moment. Surviving one zone loss means capping each zone's peak utilisation at (N-1)/N — roughly 66% with three zones.

solid answer

~40 s

Peak load here is three zones times 70%, or 2.10 zone-capacities of traffic. Losing one zone leaves 2.00 zone-capacities of compute, so each survivor is asked for about 105% of what it can serve: queues build, latency climbs, and roughly 5% of peak demand is shed or times out. The rule that prevents it is to size for the surviving set rather than the deployed set — with `N` zones and one lost, no zone may exceed `(N-1)/N` of its capacity at peak, so about 66% with three zones and 50% with two. That headroom is real money sitting idle almost all the time, which is exactly why the number gets argued over, and it is the point of the question.

code

pseudocode · 16 lines
pseudocode
zones = 3
peakUtilisationPerZone = 0.70          // fraction of one zone's capacity

totalPeakLoad = zones * peakUtilisationPerZone        // 2.10 zone-capacities
survivingCapacity = zones - 1                         // 2.00 zone-capacities
demandPerSurvivor = totalPeakLoad / survivingCapacity  // 1.05

if demandPerSurvivor > 1.0:
    overshoot = demandPerSurvivor - 1.0                       // 0.05 per survivor
    unservedShare = (totalPeakLoad - survivingCapacity) / totalPeakLoad  // ~0.048
    safeUtilisation = (zones - 1) / zones                     // 0.667
    report "survivors over capacity by", overshoot
    report "peak demand shed", unservedShare
    report "cap per-zone peak utilisation at", safeUtilisation
else:
    report "one zone may be lost without shedding load"

go deeper

for a junior

Know that losing one of three zones moves that zone's traffic onto the other two, so each of them needs room to spare. Being able to say 'the survivors carry more load' is already most of the point at this level.

for a middle

Do the arithmetic aloud and state the (N-1)/N cap, then explain why scaling out during the incident arrives too late: boot time, warm-up and health checks all land after the load has already shifted.

for a senior

Talk about what actually breaks first in your system — connection pools, a cache that lost a third of its warm entries, a downstream per-caller limit — and about the retry feedback that turns a small shortfall into a visible outage.

for a principal

Own the money question: continuous idle headroom against a rare event, and whether deliberate degradation of low-value traffic is the cheaper answer. Say who decides that and where the decision is written down.

## The arithmetic, done explicitly Assume the load is spread evenly and each zone can serve one unit of work at 100%. - Peak load = 3 zones x 0.70 = **2.10 units** of demand. - After one zone is lost, capacity = **2.00 units**. - Demand per survivor = 2.10 / 2 = **1.05**, or 105% of what one zone can serve. - Unserved share = (2.10 - 2.00) / 2.10 ≈ **4.8% of peak demand**. So the failure is not dramatic and it is not clean either. The survivors do not fall over; they saturate. Queues grow, latency climbs past the point where callers start timing out and retrying, and retries add load, which is how a 5% shortfall turns into a much larger visible failure. This is why 'we had three zones' is not on its own an answer. ## The rule: size for the surviving set The general form is straightforward. To survive the loss of one of `N` zones without shedding load, each zone's peak utilisation must satisfy `utilisation <= (N - 1) / N`. | Zones | Max peak utilisation per zone | Capacity you are paying for above peak need | |---|---|---| | 2 | 50% | 100% extra | | 3 | ~66% | 50% extra | | 4 | 75% | ~33% extra | | 6 | ~83% | 20% extra | The table shows the real trade. **Headroom is cheaper the more zones you spread across**, because the lost share is smaller. It also shows why two zones is an uncomfortable posture: to survive losing one, half your fleet is idle at peak. ## Why 'we will scale up during the incident' is a weak plan It is a plan that arrives late. The load shifts the instant the entry point stops selecting the dead zone; capacity takes minutes to request, boot, warm caches or connection pools, and pass its own health check. Three things compound during those minutes: 1. **Time.** The overload is immediate and the mitigation is not; the gap is exactly when your peak traffic is present. 2. **Contention.** Every other tenant in that region is asking for capacity of the same shape at the same moment, and capacity in a region is finite. 3. **Feedback.** Slow responses cause client retries, which raise demand while you are trying to add supply. Scaling out is a good second move. It is a poor first one, and the honest version of the answer says so. ## What the simple rule does not cover - **The data tier does not scale like the web tier.** Adding stateless instances is minutes; adding a database copy is a data copy. - **Caches lose a share of their warm entries with the zone**, so what sits behind them sees a miss burst just as the surviving instances get busier. - **Downstream services were sized against your normal footprint**, and their per-caller limits do not move because your topology changed. - **Connection pools are per instance**, so concentrating the same traffic on fewer instances can exhaust a pool even when CPU looks fine. - **Uneven placement breaks the arithmetic.** If a scale-in or a backfill left 45% of the fleet in one zone, losing that zone is worse than the even-spread maths predicts. ## Making the headroom cheaper The number is negotiable in a few honest ways: spread across more zones so each one's share is smaller; hold the headroom as capacity you are willing to have reclaimed, accepting that it may not be there; degrade deliberately under pressure by shedding the cheapest traffic first, so the shortfall lands on a feature you chose rather than on checkout; or accept a defined amount of slow-but-working during a rare event and write that down as a decision rather than discovering it during one. ## How to answer this out loud Do the arithmetic in front of the interviewer — three times 70% against two zones is the whole answer — then state the `(N-1)/N` rule, then name the cost: headroom is idle capacity you pay for continuously to cover an event that may not happen this year. Finish by saying what you would degrade if the headroom turned out to be insufficient, because that is the part most candidates never reach.

  • Does the rule change if the load is not spread evenly across the zones?
    Yes, and it gets stricter. The constraint is the worst case, so what matters is the largest share any single zone carries, not the average. If one zone holds 45% of the fleet because a scale-in or a backfill drifted, losing that zone removes 45% of capacity and the even-spread arithmetic understates the damage. Enforce and monitor the spread, not just the total.
  • Is holding idle headroom in each zone the only way to survive the loss?
    No. You can spread across more zones so each one's share is smaller, hold part of the headroom as capacity you accept may be reclaimed, or plan to degrade — shedding low-value traffic first so the shortfall lands where you chose. Each swaps money for either risk or reduced function, and the decision should be recorded rather than discovered during an incident.
  • Why does a 5% shortfall often look much worse than 5% of requests failing?
    Because saturation is not linear. Queues build, latency crosses client timeouts, and timed-out callers retry, which adds load to an already-overloaded tier. The visible failure rate can far exceed the raw capacity gap unless the service sheds load deliberately and callers back off.

A three-lane road running at 70% of capacity looks fine until one lane closes: the same cars now need two lanes, which is more than they fit into, and the queue forms in minutes even though nothing about the traffic changed.

saying these in an interview costs you the question

  • Saying three zones make the service safe without doing the arithmetic
  • Assuming replacement capacity arrives in time to cover the peak
  • Sizing each zone to its own peak rather than the surviving set's peak
  • Forgetting that retries raise load while you are already short
  • Assuming the deployed spread is still even after months of scaling