skip to content

A checkout API spans three zones but its primary datastore sits in one — which zone spread is a default and which must be declared?

level: middleimportance: must knowfreq 70%

answer

  1. spread is declared, not inherited
  2. read each resource's placement scope
  3. regional services hide the spread inside
  4. instances stay in their creation zone
  5. count distinct zones per tier

basics

~20 s

Very little spreads by itself. A service sold as regional often keeps its data redundant across zones inside the region as part of the service, while anything created as an instance is pinned to the zone it was created in until you declare a second one.

solid answer

~50 s

Two different things get called zone redundancy. The first comes with the service: when a capability is sold as regional, the provider places its copies across zones inside the region and you never see the placement. The second is yours to declare — subnets in several zones, capacity in each of them, a standby in another zone, a minimum instance count that keeps more than one zone occupied. A checkout API typically ends up in exactly this split: the web tier was given several zones because a load-balanced tier forces the question, the cache and the primary datastore were created once, in whichever zone the first person picked, and nothing ever prompted again. The check is mechanical rather than architectural — read the placement attribute of every resource, count the distinct zones each tier actually occupies, and treat any tier at one as a finding.

code

pseudocode · 16 lines
pseudocode
report = []

for each tier in application.tiers:
    zones = empty set

    for each resource in tier.runningResources:
        if resource.placementScope == "regional":
            zones.add("regional")        # the service places the copies
        else:
            zones.add(resource.zone)     # pinned where it was created

    if zones.size == 1 and not zones.contains("regional"):
        report.append({ tier: tier.name, finding: "single zone", zone: any(zones) })

for each row in report:
    print(row.tier, "has all running capacity in", row.zone)

go deeper

for a junior

Remember the default: a resource you create lives in the zone it was created in and stays there. Spreading across zones is something someone asked for, not something the region provides.

for a middle

Explain the split precisely — what the service places for you when it is sold as regional, against what you declare per tier: address ranges, running capacity, a standby, and a floor on instance count.

for a senior

Demonstrate the audit rather than the opinion: read every resource's placement scope, count distinct zones per tier from running capacity, and name the tier that sits in one zone and the zone it sits in.

for a principal

The tradeoff you own is which tiers are allowed to be single-zone on purpose, said out loud and written down, so that the exceptions are decisions with owners rather than leftovers from whoever created the resource first.

## What "the region has three zones" actually says It says three zones are **available to place things in**. It does not place anything. Every resource you create ends up with a placement: either it is **zone-scoped** — it lives in one named zone and stays there for its life — or it is **regional**, meaning the service that runs it decides where its copies sit inside the region and does not make that your problem. The confusion that produces the checkout API in the question is that both of these get described with the same phrase. One is a property of the service you bought. The other is a decision you either made or did not. ## What the platform usually spreads for you When a capability is sold as regional rather than as an instance, redundancy across zones is normally part of what you are paying for: - A **managed store sold as regional** keeps copies of an object in more than one zone inside the region as part of the service. - A **managed entry point** in front of a workload is typically a regional construct that can direct traffic at capacity in any zone you have registered. - A **managed queue or event bus** offered at the region level normally holds messages redundantly across zones. Providers genuinely differ here, and that difference matters for the answer: some replicate stored objects across zones in the region by default and also sell a cheaper single-zone option that does not, while others start single-zone and make cross-zone placement the thing you opt into. So "it is managed" is not evidence, and "it is regional" is — which is why the honest answer is to read the resource's own placement attribute. ## What you have to declare Everything that is created as an instance, and everything that holds state you own: - **Subnets in more than one zone**, because capacity can only be placed where an address range exists for it. - **Capacity in each zone**, not merely permission to use them — a tier allowed three zones but running one instance occupies one zone. - **A standby or second writable copy** for a stateful tier, with a declared placement in a different zone from the primary. - **A floor on instance count**, so an overnight scale-down cannot quietly collapse the tier back into a single zone. | Tier | What makes it genuinely multi-zone | What the single-zone symptom looks like | |---|---|---| | Stateless web tier | Capacity registered and running in several zones | Subnets exist in three zones, instances in one | | Cache | A second node placed in another zone, or a regional service | One node, created once, never revisited | | Primary datastore | A declared standby in another zone, or a regional service | A single instance with a zone attribute | | Object storage | Sold as regional, or the multi-zone option chosen | The cheaper single-zone option was picked | ## Reading the placement instead of the label The audit is boring and that is its value: 1. **List every resource** the application depends on, including the ones nobody deploys any more. 2. **Read each resource's placement scope** — regional, or a single zone. 3. **Group by tier and count distinct zones**, counting regional resources as already handled. 4. **Count running capacity, not configuration.** A declared zone list is an allowance; instances are placement. 5. **Report any tier whose count is one**, and say which zone, because that is the sentence the team will act on. What this finds, again and again, is the shape in the question: the web tier spread because a load-balanced tier forces you to pick subnets, and everything stateful stayed where it was first created because nothing ever asked. ## Two traps worth naming **A copy is not a copy in another zone.** A snapshot, an export or a backup held elsewhere is a way to rebuild; it is not a running second copy, and equating the two is how a tier gets described as redundant when it is only recoverable. **Declared spread is not running spread.** This is the single most common false positive in the audit. Configuration that permits three zones, a scaling floor of one, and a quiet night is a single-zone tier with a multi-zone description. ## Where this answer stops This is about **placement**: which tiers are where, what came with the service, and what had to be asked for. What you then build so the loss of a zone is survivable — the failover path, the duplicated capacity and what that duplication costs — belongs to the reliability material, and what the cross-zone bytes add to the bill belongs to the charge model. The finding here is simply that one tier sat in one zone while everything around it did not.

  • How do you tell whether a managed tier is genuinely spread across zones or merely managed?
    Read what the service says it places where. Some managed capabilities are sold as regional and keep copies in more than one zone inside the region as part of the offering; others run as a single instance in a zone you chose and offer a second zone only as an option you enable. The deciding evidence is the resource's placement attribute, never the word managed.
  • Capacity is declared in three zones but the tier scaled down to one instance overnight — is it zone-redundant?
    No. A declaration of where capacity may go is not capacity. At that moment the tier occupies exactly one zone, whatever its configuration allows, which is why the check must read running placement and why a scaling floor above one belongs in the configuration.
  • The cache is single-zone and the datastore is not. Is that automatically a defect?
    Not automatically, but it must be a decision rather than an accident. A cache that can be rebuilt from the datastore is a legitimate single-zone tier if losing it degrades rather than breaks the request path. The defect is finding out which it is during an incident instead of when it was placed.

saying these in an interview costs you the question

  • Assumes choosing a multi-zone region makes every resource multi-zone
  • Counts subnets declared in three zones as capacity in three zones
  • Thinks a tier is zone-redundant because it is managed
  • Reads one healthy tier's spread as proof for the whole stack
  • Equates a snapshot held elsewhere with a running copy in another zone