skip to content

You chose your first region for users and must now choose a second for recovery — why do different criteria decide it?

level: seniorimportance: should knowfreq 42%

answer

  1. the second answers a different question
  2. independence, not proximity to users
  3. far enough apart, near enough to replicate
  4. quota and services must exist first
  5. replication copies a bad delete

basics

~20 s

The first region optimises for the people using the product; the second optimises for independence from the first and for being able to run there at all. Distance stops being only a cost and becomes partly the point, while service availability, quota and the residency obligation become hard entry conditions.

solid answer

~50 s

The first choice answers "where are the users". The second answers "what must not fail with the first, and can we actually run there". That inverts one criterion and hardens two others. **Distance** flips: you want enough separation that a single event does not impair both, but not so much that replication lag or the round trip during failover becomes unusable. **Service availability and quota** become entry conditions rather than preferences — a recovery region that lacks a service, a hardware generation or the quota your architecture needs is a plan on paper, and capacity you have never launched is not reserved for you. **The residency obligation** follows the copy, so a second region outside the permitted jurisdiction is disqualified even if it never serves traffic. And the standing cost is different in kind: cross-region replication and idle standby capacity are charged continuously, whether or not you ever fail over.

go deeper

for a junior

Understand that a second region exists to be independent of the first, so it is not chosen the same way. Know that a copy of the data is still data, so the same legal constraint applies to it.

for a middle

Explain the inversion of the distance criterion and why availability and quota become entry conditions. Be able to say why a replica protects a failure domain but not a mistaken deletion.

for a senior

Show the whole decision: the failure being defended against, evidence of independence, proof the architecture runs there with quota granted in advance, and the continuous cost of holding it while nothing is wrong.

for a principal

Own the posture question — how much independence the business is buying, at what standing cost, and whether the recovery region is a maintained capability or a document nobody has exercised.

## The question changes, so the criteria change Choosing the first region is an optimisation for the people using the product: put the workload near them, inside whatever jurisdiction is required, in a place that offers what the design needs at a rate you accept. Choosing the second is a different question. It asks: *if the first becomes unusable, where do we run instead?* That reframes the same four inputs. ## Independence, the criterion that barely existed before The point of the second region is that whatever impairs the first does not reach it. Independence has more than one axis, and a candidate can be independent on some and correlated on others: - **Physical and infrastructural.** Distinct power, distinct network paths, distinct exposure to the same natural event. Two regions very close together may share more upstream than their names suggest. - **Jurisdictional.** If a legal order, a sanction or a regulator can act on both places at once, the pair is correlated in a way no engineering choice fixes. - **Operational.** Choosing a region your team has never touched means the recovery plan runs on unfamiliar ground on its worst day. But independence is not free: distance drives replication lag and the round trip you inherit while serving from the second region. The trade is to be **far enough apart that one event does not reach both, and near enough that the replication and the degraded experience are acceptable**. ## The entry conditions that harden | Criterion | In the first region | In the second region | |---|---|---| | Distance to users | Minimise it | Balance independence against lag and degraded latency | | Service and feature availability | Preference, workable around | Entry condition — the same architecture must run | | Quota | Raise it as you grow | Must exist before the incident, not during it | | Residency obligation | Filters the candidate list | Still filters it, because the copy is data too | | Cost | A rate you accept for work you do | A continuous charge for capacity you hope never to use | Two of these bite hardest in practice. **Quota** is the quiet one: an account's allowance in a region it has never used is typically small, some ceilings are raised on request and some are not, and finding out on the day of an incident costs you exactly the time you do not have. **Service availability** is the other: providers launch services, features and hardware generations region by region rather than everywhere at once, so the recovery region must be checked against the architecture it is supposed to run, not assumed to be a mirror. ## A replica is not a backup The most common misconception here deserves stating in the right direction. Cross-region **replication copies changes faithfully, including deletions and corrupt writes**. It protects you from losing a failure domain. It does not protect you from a mistake, because the mistake is replicated too, usually within seconds. Protection from a mistake comes from retained copies you can rewind to — a point-in-time restore or a backup with a retention period. A design that names its second region as its backup strategy has covered one risk and left the other open. ## What the standing cost actually is A second region is not paid for at the moment of failover. You pay continuously for: 1. **Replication traffic** crossing a regional boundary, charged for every byte, forever. 2. **Standby capacity**, at whatever fraction you keep warm — and a cold plan that keeps nothing warm trades that cost for recovery time. 3. **Duplication of everything around the workload**: configuration, identities, network layout, and the work of keeping them in step as the primary changes. ## How to defend the choice - Name the failure you are buying protection against, and show the candidate is independent of the first region on that axis. - Show the distance is a deliberate balance, not an accident of alphabetical order. - Prove the region runs the architecture: services, hardware generation, and quota granted in advance. - Confirm the obligation permits a copy there. - Say what it costs while nothing is wrong, since that is the state it will be in almost always. - State separately how you protect against a bad write, because replication does not.

  • Why is quota in the recovery region a placement criterion rather than an operational detail?
    Because an account's allowance in a region it has never used is typically small, and raising a ceiling is a request with a turnaround — sometimes minutes, sometimes a business process. Discovering that during an incident costs the hours you were trying to save, so the capacity has to be granted while things are calm.
  • Does the residency obligation apply to a recovery region that never serves traffic?
    Yes. The obligation attaches to the data, not to whether it is being read. A replica, a backup copy and a restored environment all hold the data, so a region outside the permitted jurisdiction is disqualified even if it sits idle for years.
  • How far apart is far enough?
    Far enough that no single event — power, network, natural, or legal — reaches both, and no further. Beyond that point extra distance only buys replication lag and a worse round trip while you are serving from it. It is a judgement about correlated failure, not a distance target.

saying these in an interview costs you the question

  • Picks the recovery region by alphabetical or list order
  • Assumes every region mirrors the primary's services and quota
  • Calls a continuously replicated second region a backup
  • Ignores the residency obligation because the copy is idle
  • Prices the second region only at failover, not continuously
  • Chooses the nearest neighbouring region for convenience