You cannot afford held capacity in the standby zone for every service — how do you decide which workloads get a guaranteed allocation?
answer
- certainty is bought, not configured
- rank by the cost of the worst hour
- hold the minimum serving set
- most workloads get flexibility, not allocation
- publish what the untiered ones will do
basics
~20 sRank workloads by what a refused launch costs in the worst hour, not by how important the owning team feels. A small set gets held allocation or warm machines; the rest get shape flexibility, a written degradation plan, and an acknowledged possibility of waiting.
solid answer
~50 sGuaranteed capacity is bought, not designed, and a held allocation is priced close to a running machine, so this is a monthly spend decision with a number attached. Rank by consequence in the worst hour: what breaks if this cannot start for an hour, whether it can serve degraded, whether it can be placed flexibly, and whether other teams' recovery depends on it. A small top tier gets warm machines or a held allocation **sized to the minimum serving set, never the peak fleet**; a middle tier gets an allocation only for what must start immediately; everything else gets a genuinely diverse preference list and a published degradation. Then govern it — every held allocation needs an owner, an expiry, and a periodic check that its shape and zone still match what the current failover plan would request.
go deeper
Recall that guaranteed capacity in a standby zone is purchased rather than configured, and that an organisation can normally afford it for only a few workloads.
Explain how to size what is held — the minimum set that serves acceptably, not the peak fleet — and why the rest of the plan rests on shape flexibility instead.
Show the criteria you would apply per workload, and how you verify that each held allocation still matches the shape and zone the current failover would request.
Own the standard: who qualifies, who reviews it, what the organisation publishes about the services chosen to wait, and what that certainty costs every month.
## The decision is a purchase, not a design Guaranteed capacity in a standby zone is not something you configure; it is something you buy, every month, for a machine that is usually doing nothing. A held allocation is priced close to a running machine, so `critical services get held capacity` is a policy with a monthly number attached. The only honest way to run it is to decide deliberately who gets it, and to publish what everybody else gets instead. ## What actually qualifies a workload Rank by consequence in the worst hour, using tests a reviewer can apply without knowing the team: - **What breaks if this cannot start for an hour?** Money stops, safety is affected, a regulator is involved, or another team's recovery is blocked — against a report being late. - **Can it serve degraded?** A service that is useful at a third of capacity needs a far smaller held set than one that is all-or-nothing. - **Can it be placed flexibly?** A workload that shrinks onto smaller shapes and accepts several families has decent odds without held capacity. A single-writer component needing one large shape has almost none. - **Is it on somebody else's critical path?** A shared dependency deserves the tier of the most critical thing behind it, not its own team's opinion of itself. - **How long is its honest tolerance?** A workload that may legitimately wait until the surge subsides belongs in the flexible tier however important it feels. Notice what is not on that list: how senior the owning team is, how visible its dashboard is to leadership, and whether the team can fund the allocation from its own budget. The last is the most dangerous, because it quietly replaces risk with budget size as the allocation rule. ## A three-tier standard | Tier | What it has in the standby zone | Monthly cost | What it promises | |---|---|---|---| | Serves through the hour | machines already running, sized to the minimum serving set | roughly the price of running that set | serving immediately; degraded above the held set | | Recovers within the hour | a held allocation for the minimum serving set, launched on the event | roughly the price of that set, idle | a launch that does not depend on the contended pool | | Recovers when capacity allows | no held capacity; a broad and genuinely diverse preference list | engineering effort only | best effort, with a published degradation or wait | Three tiers is usually enough. More invite negotiation; fewer collapse into `important` and `not important`, which every team answers identically. ## Sizing what is held Hold the **minimum serving set**, never the peak fleet, and derive it from the degradation contract rather than from the current machine count: 1. State the throughput or latency the service must still meet in its degraded state. 2. Compute how many machines deliver that, given what survives the failure. 3. Subtract what survives; the remainder is the held set. 4. Split it across the surviving zones rather than concentrating it in one, so a second correlated event does not remove all of it. 5. Re-derive it whenever per-machine capacity changes, which happens at most significant releases. The arithmetic is usually kinder than people expect, because the degraded target sits well below peak and much of the fleet survives the event. ## Governance, or the policy rots Held capacity is the easiest spend in an organisation to acquire and the hardest to notice afterwards. Three controls keep it honest. - **An owner and an expiry on every held allocation.** Anything lacking both is a candidate for release. - **A periodic match check.** Do the held shape and zone still correspond to what the current failover plan would request? Architectures change faster than purchases, and an allocation for a retired shape is pure loss. - **A revocation path.** When a workload leaves the tier, its allocation goes with it. If nothing ever leaves a tier, it is not a standard, it is a ratchet. ## The part that must be published The uncomfortable half of this decision is the tier that gets nothing, and it is not acceptable for that to be discovered during the event. Write down, per service, what happens in the worst hour — serves degraded, waits for capacity, or is unavailable — and have the business acknowledge it. That document converts an engineering trade-off into an owned business decision, and it is also the fastest way to discover that a service everybody assumed was top tier has never been funded as one. A leader's job here is not to eliminate the risk, because the spend required for that is unbounded. It is to make the risk explicit, pay to remove the part worth removing, and be able to say exactly which services were chosen to wait, and why.
- A team argues its internal reporting service needs held capacity because leadership watches it. How do you answer?Ask what actually breaks in the hour after the event, and what a refused launch costs in money or safety. Visibility is not the test. A reporting service that can start an hour late belongs in the flexible tier, and the held budget it would consume buys more protection somewhere that must serve through the hour.
- What review keeps a held-allocation policy from rotting?A periodic check that each held allocation still matches a shape and zone the current failover plan would actually request, and that the workload still qualifies for its tier. Architectures change faster than purchases; the usual findings are an allocation for a retired shape and one whose owner nobody can name.
saying these in an interview costs you the question
- Gives held capacity to every team that asks for it
- Ranks workloads by team seniority rather than by failure cost
- Sizes held capacity to the peak fleet rather than a serving minimum
- Leaves held capacity in place after the architecture has changed
- Never tells the business which services will simply wait