skip to content

At failover the second region refuses the fleet size the primary runs daily, though nothing was deleted — which defaults did you inherit?

level: middleimportance: should knowfreq 48%

answer

  1. the account is global, the numbers are not
  2. counted per account per region
  3. increases apply where they were granted
  4. an unused region is at opening values
  5. partial success hides the diagnosis

basics

~20 s

The second region's day-one defaults. Ceilings on how much of a resource an account may hold are scoped per region, so every increase ever granted applies where it was granted. A barely used region still sits at the numbers it was opened with.

solid answer

~50 s

Most platform ceilings — how many machines of a family, how many addresses, how many instances of a managed service — are counted and capped **per account per region**, not per account. The primary's numbers are large because two years of growth was accompanied by a long series of increases, each one requested and granted in that region. None of that travels. A region you have used only for a handful of test resources is still at the values it was opened with, which can be a small fraction of what production needs. Worse, defaults are not uniform: a newer or smaller region can open lower than the primary originally did, so even the starting point is not the same. The fix is boring and has to happen in advance, because the increase is a request with a lead time, not a setting you flip during an incident.

go deeper

for a junior

Remember that how much you may create is counted per region, not per account. The same account can be allowed a large fleet in one region and a small one in another.

for a middle

Explain which account properties are global — identities, permissions, policies above the account — and which are regional, and why every increase ever granted stayed in the region where it was asked for.

for a senior

Show that you size the second region against the full production shape, arrange increases with their lead time in advance, and prove it by starting the real fleet once rather than a test-scale one.

for a principal

Treat regional capacity readiness as a standing obligation with an owner and a review cadence, since the production shape grows continuously and last year's arrangement silently stops matching it.

## Ceilings are scoped to a region, and so is their history When a platform caps how much of something an account may hold, the counter is almost always kept **per account, per region**. That is the natural implementation: the capacity being protected is physical and lives in one region, so the ceiling that protects it lives there too. The consequence people miss is that the *history* is scoped the same way. A production region's generous numbers are not a property of the account. They are the accumulated result of every increase that was requested and granted **in that region**, usually one at a time, each triggered by something nearly failing. Nobody wrote that history down as a design artefact, so nobody thinks to reproduce it. A second region that has only ever held a handful of experiments has none of that history. It sits at its opening values. ## What travels to a new region and what does not | travels with the account | stays behind in the region where it was set | |---|---| | identities, groups and their permissions | how many machines of a family you may run | | policies applied above the account | how many addresses you may hold | | billing relationship and any term commitments | how many instances of a managed service you may create | | naming and tagging conventions you enforce | how much of a specific accelerator generation you may hold | | the organisation's ownership of the account | every increase previously granted, anywhere else | Because the left column is what people think of as "the account", a second region feels like the same account with the same capabilities. Structurally it is the same account with the **same permissions and different ceilings**. ## Why the starting numbers are not even equal Two further effects push the second region lower than intuition suggests: - **Defaults vary by region.** A provider protects a smaller or newer region's finite capacity with smaller opening numbers, so the second region can start below where the primary started. - **Defaults are not frozen.** Opening values change over time in both directions, so the number your primary was opened with years ago may not be the number any region is opened with now. - **Providers differ.** Some platforms reset every ceiling per region; some hold a few of them at the account level; some vary by resource. Where the design depends on it, check rather than assume. ## The failure this produces, and its timing The shape is always the same and always arrives at the worst moment: 1. The primary runs a fleet of some size every day, so nobody thinks of that size as exceptional. 2. The second region is stood up, and the pieces are created at test scale, which fits comfortably inside the opening values. 3. Everything is declared ready, because everything that was tried worked. 4. The day the real shape is requested there, the platform refuses part of it — not all of it, which is what makes the diagnosis slow. You get some of the fleet, and the rest is rejected with a ceiling error. Partial success is the cruel detail. A flat rejection is diagnosed in a minute; a fleet that comes up at forty percent of its size looks like a capacity problem, a scheduling problem or a configuration drift problem long before anyone thinks of a ceiling. ## What to do instead 1. **Write down the production shape as numbers.** How many machines of which family, how many addresses, how many instances of each managed service, at full size — not at steady state, at the size you would need if everything moved there. 2. **Read the second region's current values against that list.** Every number, in the target region, not the primary. 3. **Arrange the increases ahead of time.** An increase is a request that a human or a system has to approve, and it has a lead time measured in hours or days. That lead time is only affordable before the incident. 4. **Re-check on a schedule.** The production shape grows, and the increase you obtained last year was sized for last year's shape. 5. **Actually create the full shape once.** Numbers on a page are a claim; a fleet that started is evidence. Start it, confirm the count, tear it down. ## The point the question is really testing The candidate who has lived through this says "ceilings are per region and so was every increase we ever got" immediately. The weaker answer treats the second region as the same account and therefore the same allowances, which is true of permissions and false of capacity. Knowing which properties of an account are global and which are regional is the whole of the answer.

  • Which properties of an account are genuinely global, then?
    Broadly the identity and governance side: the principals and their permissions, the policies applied above the account, the billing relationship, and the organisation's ownership of it. Capacity ceilings and the resources themselves sit inside a region. So a second region gives you the same permissions with different allowances.
  • Why is a fleet that comes up at partial size harder to diagnose than one that fails outright?
    Because partial success looks like the failures engineers meet most often — scheduling pressure, a shortage of capacity, drifted configuration. A flat rejection points straight at the ceiling. The habit worth building is to check the regional ceiling first whenever a fleet stops short of its requested count in a region you rarely use.

saying these in an interview costs you the question

  • Assumes ceilings follow the account, so a second region inherits the primary's numbers.
  • Thinks the two regions must open at identical default values.
  • Reads a test-scale deployment as proof the region can hold production scale.
  • Expects an increase to take effect immediately during an incident.
  • Diagnoses a partially-started fleet as a capacity shortage without checking the ceiling.