skip to content

A failover into the second region is refused at a quota, although the primary region ran the same fleet for months — why?

level: middleimportance: must knowfreq 58%

answer

  1. look at the scope, not the number
  2. ceilings are counted somewhere specific
  3. per account, and per region
  4. raises do not follow the workload
  5. the standby kept its starting defaults

basics

~20 s

Quotas are counted per account and separately per region, so the standby region still holds the defaults the account started with. Years of small raises in the primary never followed the workload across, and nobody had exercised the second region at full size.

solid answer

~50 s

A customer does not have one ceiling; they have one ceiling per account, per region, and sometimes per narrower scope again. The primary region's numbers drifted upward over time through a series of small increase requests that nobody recorded as architecture, while the standby region kept whatever the account was given on day one. Because a raise applies only to the scope it was granted for, none of that history travelled. The automation that builds the second region succeeds at smoke-test size, so the ceiling is met for the first time at exactly the moment the full fleet is created — during an incident, when the only remaining lever is an increase request measured in days. The fix is to inventory what the failover path consumes at full size, raise those ceilings in the standby region ahead of time, and prove it once at production size.

code

json · 8 lines
json
{
  "quotaName": "concurrent-compute-instances",
  "raisable": true,
  "scopes": [
    { "account": "production", "region": "primary", "granted": 400, "used": 310 },
    { "account": "production", "region": "standby", "granted": 40,  "used": 0 }
  ]
}

go deeper

for a junior

Remember the scope: a quota is counted per account and per region, so the same resource has a different ceiling in each region. Say that a raise applies where it was granted.

for a middle

Explain why the primary's ceiling is high — years of small increase requests — and why none of that travels. Note that the refusal is on the creation path, so running workloads keep serving.

for a senior

Describe how you would have caught it: required-at-full-size compared against granted, per region, including the paths you only use in an emergency, plus one deliberate full-size exercise of the standby.

for a principal

Treat every number your recovery plan depends on as something that must be proven per scope. Decide who owns that inventory across accounts, and what evidence a recovery plan must carry before it is accepted.

## The number is not the interesting part — the scope is When people talk about a quota they usually quote a number. The number is the least portable thing about it. What actually determines whether a quota bites is **what it is counted against**, and for most platform resources that is a pair: one account, one region. Some ceilings are counted narrower still — per private network, per grouping of resources, per family of machine — but the account-and-region pair is the default mental model and it is the one this failure turns on. So a customer does not have *a* ceiling on running instances. They have one in each region they operate in, and each of those is a separate number with its own history. ## Why the primary region's history does not travel A raise is granted **for the scope it was requested for**. The primary region's ceilings did not start high; they climbed over several years through a sequence of small increase requests, each filed by whoever was blocked that week. That sequence is invisible in the running system — the ceiling simply stopped being a problem — so nobody thinks of it as a design artefact that a second region would need too. What does and does not travel when you stand up a second region: - **Travels:** the infrastructure definitions, the build artefacts, the pipelines, the runbooks — everything you keep in a repository. - **Travels:** the automation that creates the fleet, which is exactly why the second region looks correct. - **Does not travel:** the granted ceilings in the other region. - **Does not travel:** the accumulated history of raises, or the usage evidence the provider reviewed when granting them. ## Why the standby region is where this is found A standby region is created by the same automation and validated the way a standby is usually validated: bring up one instance of each tier, check it responds, tear it down. That succeeds comfortably under a starting ceiling. The first time the region is asked for the *full* fleet is the first time the ceiling is approached at all — and that day is, by construction, an incident. The shape of the failure matters. The refusal is on the **creation path**: the management API declines to build the new instances. Whatever is already running in that region continues to serve, and the primary region is not made worse by it. What you lose is the ability to *grow* the standby to the size the traffic needs, at the one moment you needed that. Two other refusals look superficially similar and are not this: - Being pushed back for issuing too many management calls too quickly is a ceiling on **request rate**, not on how much you may have, and it clears on its own with backoff. - The provider genuinely not having the resource to give you is a **capacity** failure; its error speaks about availability rather than about a ceiling, and no increase request fixes it. Read the refusal before choosing the remedy. ## What to do about it 1. Inventory what the failover path consumes **at full size** — instance counts by family, addresses, network constructs, storage volumes — and record the number per region, not once. 2. Raise the standby region's ceilings to that size ahead of time, as ordinary backlog work with an owner and a date. 3. Exercise the standby at production size at least once, deliberately, in daylight. A ceiling that is never approached is never proven. 4. Re-check after growth in the primary. The target moves, so a raise granted two years ago against a smaller fleet is no longer enough. ## The generalisation worth carrying Anything counted per scope has this failure mode, not just quotas: the scope you exercise daily accumulates fixes, and the scope you keep for emergencies keeps its defaults. The discipline is to ask, for every ceiling your recovery plan depends on, *in which account and which region is this number granted, and when was it last proven at the size I will actually need?*

  • The standby region was built by the same infrastructure code as the primary. Why did that not carry the ceilings across?
    Because ceilings are not resources your code creates; they are policy the provider holds about your account in that region. Infrastructure code describes what to build and the platform decides whether you may. The two live on opposite sides of the boundary, so identical definitions can succeed in one region and be refused in the other.
  • How would you have found this before the incident, without running the full fleet in the standby every day?
    Compare granted against required, not against used. Compute what the failover path would create at full size, list it per region, and check it against the granted ceiling in the standby scope. That is a paper exercise you can run weekly; the expensive live test then only needs to happen occasionally.

A card issued to the same company can carry a different spending limit in each country it was activated in. Raising the one you use every week does nothing for the one in the drawer, and you find that out at the till.

saying these in an interview costs you the question

  • Assumes a quota is one estate-wide number per customer
  • Believes a raise in one region applies in all of them
  • Treats two regions as identical because the definitions match
  • Checks headroom only where the workload runs today
  • Calls a standby proven after a single-instance smoke test