Your payments ledger fails over into its standby zone and half its replacement machines will not launch — why is capacity scarcest exactly then?
answer
- the plan drains the pool the plan needs
- everyone fails over in the same hour
- demand correlates with the failure itself
- unlaunched capacity is a forecast
- warm, held, or merely hoped for
basics
~20 sFailover is the moment every tenant wants the same thing at once. The loss of a zone pushes many organisations onto the same surviving zones and the same popular shapes within the same minutes, so the pool the plan intended to launch into is drained by every other plan.
solid answer
~50 sThe plan assumed a pool that the plan itself helps to drain. When a zone becomes unusable, everyone whose standby sits in the same surviving zone launches replacements in the same few minutes, so demand on those pools is correlated with the very event the design was meant to survive. Capacity that has not been launched is a forecast about somebody else's inventory, not a resource. What survives correlation is capacity that already exists: machines running warm in the standby zone, or an allocation held in advance so the launch is drawn from stock set aside for your account. Everything else — a preference list of shapes, more than one fallback zone, a smaller size — improves the odds and is worth having, but it is a lottery played against every other tenant holding the same ticket.
go deeper
Recall that a standby zone only helps if machines can actually be launched there, and that many other customers are launching into that same zone at the same moment.
Explain why demand correlates with the failure itself, and name the three grades of standby capacity: already running, held in advance, or merely planned for.
Show the failover plan you would actually sign: what is held, what is launched, what degrades, and how you rehearsed it knowing the rehearsal is optimistic by construction.
Own the spend decision — how much idle or held capacity the organisation buys, against how much degradation it is willing to accept in its worst hour.
## The assumption hiding in most failover plans A standby zone is usually described as `we run in two zones and fail over to the other one`. Read closely, that sentence contains a purchase nobody made: it assumes that at the moment of the failure, machines of the right shape will be free in the surviving zone and will launch on request. Capacity that has not been launched is not a resource. It is a forecast about another company's inventory, made for the one hour in which the forecast is least likely to hold. ## Why demand correlates with the failure The event that removed a zone is a **shared cause**, and the reactions to it are alike for structural reasons: - **The same trigger.** One physical or platform-wide event starts everybody's recovery inside the same few minutes. - **The same destination.** Two-zone and three-zone designs are the norm, so the surviving zones of that region are where everyone goes. There is nowhere else to go. - **The same shapes.** Fleets are built from a small set of popular families and sizes, so the demand lands on the same pools. - **Automation rather than people.** Health checks and orchestrators react in seconds, compressing everyone's demand into one window instead of spreading it over an hour. - **Amplification.** Failed launches are retried, often aggressively, so the offered demand is larger than the real demand and arrives faster. The consequence is blunt: the pool your plan depends on is drained by the plans of every other tenant, and regional capacity in aggregate is no comfort, because you cannot launch in a zone that is not the one you are failing into. ## Three grades of standby capacity | What sits in the standby zone | What it costs while nothing is wrong | What it is worth at failover | |---|---|---| | A plan to launch | nothing | whatever happens to be free that minute | | An allocation held for your account | close to the price of running those machines | a launch drawn from stock no other tenant may take | | Machines already running | the price of machines doing little | capacity that needs no launch at all | Only the lower two rows are insensitive to correlated demand. The top row is what most organisations actually have, and it works most of the time — which is exactly why it is rarely questioned before the hour in which it does not. ## Sizing the part you pre-pay for Nobody holds a duplicate fleet. The tractable question is how much of the fleet must exist immediately for the service to be serving at all, and only that much is held. Take a service running ninety machines spread evenly across three zones, thirty per zone, with peak demand needing all ninety. Lose a zone and sixty survive. If the service is acceptable at two thirds of peak throughput, sixty is already enough and nothing needs holding. If it is only acceptable above roughly eighty-five percent of peak, it needs about seventy-seven machines, so the gap is seventeen — held as, say, twenty split between the two surviving zones rather than a standby thirty. Those figures are illustrative and assume each machine contributes equally and that traffic can actually be steered onto the survivors. The method is the point: hold the **minimum serving set** and launch the remainder best-effort. ## What to do with everything above the held line 1. **Degrade deliberately.** Decide in advance which features, tenants or request classes are dropped first, so a partly-capacitated service is a designed state rather than a random one. 2. **Widen the fallback list for the remainder**, with entries genuinely unlike each other: another generation, a differently weighted family, the less fashionable surviving zone. 3. **Retry spaced, not tight.** Backoff with jitter places more machines than a loop, because the loop competes with every other tenant's loop and can be throttled. 4. **Prefer smaller units for the overflow**, since they fit the fragments that whole large shapes cannot use. ## Rehearsals are optimistic by construction A failover drill run on a quiet afternoon proves the automation works, the data is where it should be and traffic steers correctly. It cannot prove the pool will be there, because the variable that matters — how many other tenants want the same shapes in the same zone in the same minutes — is absent from the drill and supplied only by the real event. Treat drill timings as a lower bound on the good path, and rehearse the degraded path explicitly: force the launches to fail and confirm the service still serves on what it holds. How the data stays consistent across zones, and how traffic is steered after the loss, are their own decisions. This one is narrower and comes first: whether the machines you intend to run will exist when you ask for them.
- Your failover rehearsal launched every replacement in under two minutes. Why is that not evidence?Because the rehearsal ran while no other tenant was failing over. The variable that decides the real outcome — how many others want the same shapes in the same zone in the same minutes — was absent, and only the real event supplies it. A drill proves your automation works; it cannot prove the pool will be there.
- Does running the standby fleet warm remove the problem entirely?It removes the launch from the critical path for the capacity you already hold, which is the main risk. It does not cover scaling beyond that fleet, replacing a machine that fails during the event, or a second loss, and it costs the full price of machines doing little. Most teams run a minimum serving set warm and accept degradation above it.
- Why does retrying harder make correlated scarcity worse?Every tenant's automation reaches the same conclusion at the same time, so the retry traffic becomes a second correlated wave against the provisioning interface, and aggressive retries can be throttled. Spaced attempts with jitter, plus an early decision to accept a smaller shape, place more machines than a tight loop does.
A hotel beside a conference centre has rooms most nights. The night the conference hotel floods, everyone walks over within the same hour, and only the guests who already held a booking get a room.
saying these in an interview costs you the question
- Assumes the standby zone will have machines because it usually does
- Treats a failover tested on a quiet afternoon as proof of capacity
- Believes a fallback shape list is enough for a must-not-fail launch
- Thinks regional capacity in aggregate guarantees one zone can absorb failover
- Says paying on-demand rates confers priority over other tenants