skip to content

Regional Service Parity

Not every region offers every service: capabilities arrive on a rollout schedule, and features, hardware generations and quotas differ. Asked because recovery plans assume the second region matches.

on this pageshow

questions

5

Your recovery region lists the managed service as available, so why can it still refuse the feature your architecture depends on?

level: middleimportance: must knowfreq 60%

answer

  1. the list is coarser than the dependency
  2. services are bundles of features
  3. features ride hardware and other services
  4. the error names the field
  5. provision the real shape to find out

basics

~20 s

Availability is published per service; parity is decided per feature. A region can run an earlier generation of the same managed service, missing an option, a tier, an integration or a size that the architecture was built on, while still answering to the same name.

solid answer

~50 s

A service name is a coarse unit. What a provider rolls out region by region is not one indivisible product but a stack of features, tiers, hardware-backed sizes and integrations with other services, each of which arrives on its own schedule. So a region can legitimately list the service as available while the specific option your design uses is rejected at provisioning time — because that option depends on hardware not installed there, or on a second service that has not arrived, or simply because it sits later in the rollout queue. The failure surfaces as an error on a create call, not as a missing entry in a catalogue, which is why a design review that checked the service list feels safe and is not. The only reliable check is to issue the real request, with the real options, in the target region.

code

http · 20 lines
http
POST /v1/managed-store/instances HTTP/1.1
Host: management.region-b.provider.example
Content-Type: application/json

{
  "engineTier": "standard",
  "instanceFamily": "general-4",
  "crossZoneReplication": true
}

--- response ---

HTTP/1.1 400 Bad Request
Content-Type: application/json

{
  "error": "featureNotAvailableInRegion",
  "field": "crossZoneReplication",
  "message": "The service is offered in this region; this option is not yet enabled here."
}

go deeper

for a junior

Remember that a service being offered in a region does not mean every option of it is. If a create call is rejected on one field, read the field name before concluding the service is absent.

for a middle

Explain the resolution mismatch: availability is published per service while a design depends on features, tiers, sizes and integrations that roll out separately, each on its own hardware and service dependencies.

for a senior

Demonstrate the verification discipline — issue the real request with the real options, diff the created object against what you asked for, and record each accepted gap as a decision rather than a note.

for a principal

Own the standard: decide whether the organisation designs to the intersection of its regions, which caps capability, or to the primary and funds a per-region exception register, which caps nothing and costs attention forever.

## "Available" is published at the wrong resolution A provider's service-availability information is organised by **service**, because that is the unit customers shop for. What actually rolls out, though, is much finer: each service is a bundle of **features**, **tiers**, **hardware-backed sizes** and **integrations with other services**, and every one of those has its own arrival date in every region. The unit you can look up and the unit you depend on are not the same unit. So the common shape of a parity surprise is not "the service is missing". It is: the service is there, the create call is accepted in principle, and one field in the request is rejected. ## Where feature gaps come from - **The feature needs hardware the region does not have.** Anything backed by a specific accelerator, storage medium or network capability is limited to the racks that exist on that floor. - **The feature needs another service that has not arrived.** Managed services are built on other managed services; an option that reaches into a second service cannot ship before that second service does. - **The feature is later in the queue.** Providers ship a service into a new region in a reduced configuration to make it available sooner, and fill it in afterwards. - **The region's own shape limits it.** An option that depends on a minimum number of separate failure domains cannot be offered in a region built with fewer of them. - **The tier is not offered.** The same service can be sold in several tiers, and a region may carry only the mainstream one. ## What this looks like in practice | what you checked | what you actually needed to check | where it fails | |---|---|---| | "the managed store is available there" | that this engine tier is offered there | rejected create call | | "the compute service is available there" | that this instance family exists there | rejected, or silently different sizes | | "the key handling option exists" | that it is offered in **this** region for **this** service | rejected on one field | | "replication is supported" | which replication scope — within the region or across regions | accepted, but not doing what you meant | The last row is the dangerous one: a request that is accepted with a narrower meaning than you intended fails silently, and you find out when the property you assumed turns out never to have existed. ## How to find out before it matters 1. **Enumerate dependencies at option level.** Write down not "a managed store" but the tier, the size family, the replication scope, the retention behaviour and every integration the design leans on. That list is the parity checklist. 2. **Issue the real request in the target region.** Provision the actual configuration — not a simpler one — and read the error. A request that is rejected names the field, which is the most precise parity information you will ever get. 3. **Compare the resulting object to the request.** Some requests succeed with an option quietly dropped or downgraded. Read back what was created and diff it against what you asked for. 4. **Record the gap and the decision.** For each gap: change the design, change the region, or run differently there and accept it. An undocumented accepted gap becomes tomorrow's incident. ## Why this bites recovery regions hardest A primary region is usually one of the largest a provider runs, chosen when the product launched. A recovery region is chosen later, often for legal or distance reasons, and is frequently smaller and younger. So the gap is not random: it is systematically in the direction of the recovery region having less. And because nobody runs production traffic there, nothing exercises the gaps until the day something does. The architecture can also depend on a feature it never names. If the primary's managed store does replication across failure domains inside the region by default, and the recovery region offers a configuration where it does not, the design's durability assumption has quietly changed without a single line of the design being edited. ## The reasoning an interviewer wants back The strong answer has three moves in it. First, the resolution point: availability is per service, dependency is per feature. Second, the mechanism: features ride on hardware and on other services, so they arrive separately. Third, the check: the only trustworthy verification is a real provisioning attempt with the real options in the real region, because the error message operates at the resolution the design actually needs. A candidate who stops at "check the region list first" has described the step that produced the false confidence.

  • A create call in the second region succeeds. Have you proved parity for that component?
    Not yet. Read the created object back and compare it field by field with the request: an option can be dropped, downgraded or interpreted with a narrower scope while the call still returns success. Parity is proved by the resulting configuration matching, not by the absence of an error.
  • Why is a feature gap more likely in a recovery region than in the primary?
    Because the primary is usually one of a provider's largest and oldest regions, chosen when the product launched, while the recovery region is chosen later for distance or legal reasons and is often smaller and younger. The rollout order puts smaller and younger regions behind, so the gap runs systematically in one direction.
  • Which gap is hardest to detect?
    The one where the request is accepted with a narrower meaning than intended — for example a replication option that exists in both regions but covers a smaller scope in one. Nothing errors, so only reading the created configuration back, or measuring the behaviour, reveals it.

A chain restaurant's small-town branch carries the same sign and the same menu headings as the flagship, but not every dish — and the one you drove there for may be exactly the one it does not make.

saying these in an interview costs you the question

  • Treats a service-availability list as proof the architecture will provision there.
  • Assumes a service offered in two regions is the same service in both.
  • Thinks a rejected option means the whole service is missing from the region.
  • Believes a create call that succeeds must have honoured every option requested.
  • Says the gap is a bug in the provider rather than a stage of the rollout.
  • Verifies parity by provisioning a simplified test configuration instead of the real one.
open as a page

How do you prove a second region actually runs your architecture rather than trusting the provider's service-availability information?

level: seniorimportance: must knowfreq 52%

basics

~20 s

By deploying it there. Build the second region from the same description as the first and let every gap surface as a failed request, then record what was refused, what provisioned differently, and what you accepted. A table of service names proves nothing.

open as a page

Why does a provider's new capability reach some regions months before others, and what follows for a smaller region?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A region is a physical build-out, so capabilities ship on a staged rollout: racks, dependent services and trained operators arrive one region at a time. Smaller and newer regions get a capability last, and some never get it.

open as a page

At failover the second region refuses the fleet size the primary runs daily, though nothing was deleted — which defaults did you inherit?

level: middleimportance: should knowfreq 48%

basics

~20 s

The second region's day-one defaults. Ceilings on how much of a resource an account may hold are scoped per region, so every increase ever granted applies where it was granted. A barely used region still sits at the numbers it was opened with.

open as a page

The recovery region offers only an older accelerator generation, so how does that change the shape of the training job you planned to run there?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

It changes the job, not just the machines. An older generation typically offers less device memory and a slower interconnect, so the same work needs a smaller slice per device, more machines, more cross-machine traffic, and a longer wall clock at a different cost per unit of work.

open as a page