skip to content

When Capacity Runs Out

Quota says you may; capacity says whether the provider can. A shape can simply be unavailable in a zone, and it is unavailable exactly when everyone wants it.

on this pageshow

questions

5

A launch request for a new machine is refused in one zone — how do you tell an empty capacity pool from an administrative ceiling?

level: middleimportance: must knowfreq 62%

answer

  1. who is short, you or the provider
  2. administrative permission against physical stock
  3. one of the two is a published number
  4. try the same request in another zone
  5. a smaller shape proves less than you think

basics

~20 s

A refused launch means one of two shortages: the provider has no machine of that shape free in that zone, or your account is not permitted more. A ceiling is a number you can look up; pool depth is never published.

solid answer

~50 s

Separate the two before doing anything else. An **administrative ceiling** is a count attached to your account within a region — the provider has the machine, you are not permitted it — and you can read your current usage against that number before you ever launch. An **empty capacity pool** means the provider has nothing of that shape free in that zone right now; permission is irrelevant, and the depth of a pool is never published to anyone. The refusal usually names which one it is. The cheapest confirmation is to submit the same request into another zone of the same region: a ceiling is counted against the account and follows you, while pool scarcity is local. Trying a smaller shape is suggestive but not conclusive, because some ceilings are counted in aggregate units such as cores.

code

pseudocode · 20 lines
pseudocode
result = requestMachine(shape, zone)

if result.accepted:
    return result.machine

if result.refusedBecause == 'not permitted':
    // the account is at a counted ceiling; the machines exist
    recordCeilingShortfall(shape, region)
    return raiseWithPlatformOwner(shape, region)

if result.refusedBecause == 'no capacity':
    // the pool for this shape in this zone is empty right now
    for each candidate in fallbackShapesAndZones(shape, zone):
        result = requestMachine(candidate.shape, candidate.zone)
        if result.accepted:
            return result.machine
    return retryLater(backoffWithJitter)

// anything else is not a capacity problem at all
return reportUnexpectedRefusal(result)

go deeper

for a junior

Recall that a launch can be refused for two different reasons, and that one of them is about what your account is permitted while the other is about whether the provider has a free machine of that shape in that zone.

for a middle

Explain what each refusal means mechanically and name a check that separates them: reading current usage against the published ceiling, or submitting the same request into a second zone of the same region.

for a senior

Show that you act differently on each. An administrative shortfall is handled with the provider ahead of time; an empty pool is answered by changing shape or zone, or by having held an allocation before the event.

for a principal

Argue where the organisation should spend so the distinction never matters for the workloads that cannot tolerate a refused launch, and what it consciously accepts for all the others.

## Two shortages that look identical from the outside A request to create a machine can be refused for reasons that feel the same in a terminal and are nothing alike underneath. - **An administrative ceiling.** The platform counts something your account holds — machines of a family, processor cores, addresses — and compares it with a number attached to the account in that region. Over the number, the request is refused. The hardware exists; you are not permitted it. A **soft quota** is a ceiling the provider is willing to move, and hitting it costs you **time**; a **hard limit** is one it will not move, and hitting it costs you an architecture change. - **An empty capacity pool.** The provider has no machine of that shape free in that zone at that moment. Your permission is irrelevant. Nothing you own, owe or agree to changes the physical stock in that building in the next second. Other refusals exist — a policy set above the account denying the action, a suspended account, a malformed request, a machine image not present in that zone — but they announce themselves differently and are not the pair that gets confused. ## Why the distinction is the whole question Because the two have opposite remedies, and each remedy is useless against the other. Against a ceiling the remedy is administrative and costs time: the number is changed with the provider, and nothing you do to the workload helps. Against an empty pool the remedy is substitution or prior allocation: a different family, a different size, a different zone, or capacity already set aside for you. Asking for a larger ceiling adds permission to a shortage of hardware and changes nothing at all. The expensive failure is a team that spends an outage raising a ceiling that was never the binding constraint. ## Reading the refusal | What you observe | Points to a ceiling | Points to an empty pool | |---|---|---| | The refusal text | names a count, quota or limit on a resource | names capacity or availability for a shape in a location | | Your usage against the published ceiling | at or very near the number | far below it | | The same request in another zone | refused the same way | frequently succeeds | | The same request a few minutes later | refused the same way | may succeed | | A different family at the same size | refused if the count covers it | frequently succeeds | The most reliable check is the second row, and it is reliable because of an asymmetry worth memorising: **your ceiling is a number the platform publishes to you, and the depth of a capacity pool is never published to anyone.** You can always answer `am I at my ceiling?` in one read. You can never answer `how much is free?` except by asking for a machine. The other checks are useful with a caveat each. A second zone is a strong discriminator because a ceiling is normally counted against the account within a region, so it follows the request across zones while pool scarcity is local to one. A smaller shape succeeding is **suggestive but not conclusive**: where the ceiling is counted in aggregate units such as cores, a smaller request can also fit under a ceiling that refused a larger one, so the same observation is consistent with both stories. ## What to do with each answer 1. **Confirm which one it is before acting** — one read of usage against the ceiling, and one identical request into a second zone. 2. **If it is a ceiling**, treat it as lead time rather than an incident: the number is moved with the provider, and that takes as long as it takes. 3. **If it is the pool**, walk a preference list written in advance — acceptable families, sizes and zones — rather than inventing substitutions while the launch is failing. 4. **Retry, but bounded and spaced.** Pools refill continuously as other tenants release machines, so a few attempts with backoff and jitter genuinely help. A tight loop does not: it loads the provisioning interface, it can be throttled, and when demand is correlated every other tenant is looping alongside you. 5. **If the launch must not fail**, neither remedy is a plan. Capacity already running, or an allocation held in advance, is the only arrangement that does not depend on the pool at the moment of need. ## The habit that prevents the diagnosis Check headroom against your ceiling *before* the event that will consume it — a launch, a migration, a scale-out — because that side is knowable in advance and slow to move. Then assume the other side is not knowable, and design the workload so that several shapes and more than one zone are acceptable. A workload pinned to exactly one shape in exactly one zone has the worst possible odds and no diagnosis left to make: every refusal is final.

  • The same request is refused in every zone of the region. What does that shift your suspicion towards?
    Towards the account side. A ceiling is normally counted against the account within a region, so it follows the request everywhere, while pool scarcity is a property of one zone and one shape. Region-wide scarcity of a single family does happen, usually for accelerator-equipped or very large shapes, so confirm by reading usage against the ceiling rather than inferring it from the pattern.
  • Why is a bounded retry with backoff still worth running against a refusal that names capacity?
    Because pools refill continuously as other tenants release machines, so the refusal is a statement about this second, not about the day. A few spaced attempts with jitter often succeed and cost almost nothing. What they must not be is the plan: a tight loop adds load to the provisioning interface and can be throttled, and when demand is correlated every other tenant is looping with you.

saying these in an interview costs you the question

  • Assumes every refused launch means a quota must be raised
  • Believes a bigger budget makes the provider find capacity
  • Retries the identical request in a tight loop until it works
  • Treats headroom under a ceiling as guaranteed machines
  • Thinks a smaller shape launching proves the pool was empty
open as a page

Your payments ledger fails over into its standby zone and half its replacement machines will not launch — why is capacity scarcest exactly then?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Failover is the moment every tenant wants the same thing at once. The loss of a zone pushes many organisations onto the same surviving zones and the same popular shapes within the same minutes, so the pool the plan intended to launch into is drained by every other plan.

open as a page

Your preferred machine family is unavailable in a zone while a smaller profile launches there — what does that reveal about capacity pools?

level: middleimportance: should knowfreq 46%

basics

~20 s

Capacity is not one fleet-wide reservoir but many narrow pools, roughly one per zone per machine family and size. Scarcity therefore has an address, and a smaller or differently weighted shape often launches where the preferred one will not.

open as a page

A team buys a multi-year spend commitment expecting machines to be waiting in its standby zone — what has it actually bought?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A term commitment is a billing instrument: it lowers the price of machines you manage to obtain and sets none aside. Guaranteeing an allocation in a named zone takes a capacity reservation, which is a separate purchase billed from the moment it exists.

open as a page

You cannot afford held capacity in the standby zone for every service — how do you decide which workloads get a guaranteed allocation?

level: principalimportance: should knowfreq 33%

basics

~20 s

Rank workloads by what a refused launch costs in the worst hour, not by how important the owning team feels. A small set gets held allocation or warm machines; the rest get shape flexibility, a written degradation plan, and an acknowledged possibility of waiting.

open as a page