skip to content

Leadership asks whether checkout should survive losing an entire region — how would you frame the choice against a second zone, and what does each posture cost?

level: principalimportance: should knowfreq 38%

answer

  1. start from downtime cost, not architecture
  2. two numbers the business signs
  3. cold, warm, active-active differ in running capacity
  4. the second side rots unless exercised
  5. most outages are self-inflicted, not regional

basics

~20 s

Frame it as recovery time and acceptable data loss priced against the cost of each posture. Multi-zone is cheap and lossless; a second region costs duplicated capacity, duplicated data and continuous engineering to keep both true, and rises steeply from cold standby to active-active.

solid answer

~50 s

Turn the question into two numbers the business owns — how long checkout may be down, and how much confirmed work may be lost — and then price the postures that meet them. Multi-zone inside one region is the cheap baseline: synchronous replication, no data loss, automatic recovery, and only headroom to pay for. Crossing into a second region changes the shape: a cold standby holds data but little running capacity and recovers in hours; a warm standby runs a scaled-down copy and recovers in tens of minutes; active-active serves from both and recovers almost immediately while costing full duplicate capacity plus the hardest engineering. The cost people forget is not the second bill — it is keeping both sides genuinely equivalent as the system changes every week, and the fact that most outages are self-inflicted rather than regional.

go deeper

for a junior

Understand the ladder in outline: more of the second site running means faster recovery and a bigger bill, from data-only at one end to both sites serving traffic at the other.

for a middle

Be able to describe cold, warm and active-active by what is running and how long recovery takes, and to say why active-active is the one that changes how writes work rather than just how much you spend.

for a senior

Bring the costs that are missed in estimates — duplicated dependencies, keeping the second side equivalent, twice the operational surface — and say how you would stop the standby from quietly diverging.

for a principal

Drive the conversation from downtime cost and an agreed loss window, recommend per capability rather than per company, and be explicit that most outages are self-inflicted so the money may buy more availability elsewhere.

## Frame it as two numbers, then price the options Do not start from the postures. Start from what the business will sign: 1. **How long may checkout be unavailable** in a rare, large event? 2. **How much confirmed customer work may be lost** if the answer to the first is 'not long'? Those two numbers, plus an honest estimate of what an hour of lost checkout costs, turn an architectural argument into a comparison of prices. Without them, the conversation collapses into 'more resilience is better', which always loses to whoever is arguing about budget. ## The postures, side by side | Posture | What runs in the second place | Typical recovery | Data loss | What it costs | |---|---|---|---|---| | Multi-zone, one region | full capacity across zones | seconds to minutes, mostly automatic | none, with a synchronous standby | headroom in each zone | | Cold standby, second region | data copied; little or no compute | hours | the replication lag | data storage, plus a procedure that decays | | Warm standby, second region | a scaled-down but live copy | tens of minutes | the replication lag | part of a duplicate fleet, all of a duplicate data set | | Active-active, two regions | full capacity, serving traffic | near zero | depends on how writes are handled | duplicate everything, plus the hardest engineering | The jump that surprises people is not between cold and warm. It is between **anything with a single writable side** and **active-active**, because the moment both regions accept writes you own conflict handling, identifier allocation that does not collide, and a consistency story for every table. ## The costs that get left out of the estimate - **Keeping the second side true.** Every schema change, configuration change, credential rotation and new dependency has to land in both places. A standby that is not exercised diverges quietly, and the divergence is found during the event it exists for. - **Duplicated dependencies.** The posture is only as good as its least-portable component: a third-party service available in one place, a shared file store, a queue with one home, an identity integration that is pinned. Every one of them has to be solved or accepted. - **Operational surface.** Twice the estate to patch, monitor, secure and reason about, and alerting that has to say *which* side is unhealthy. - **Cognitive cost on every change.** 'Does this work in both regions?' becomes a standing question in design review, and it slows everything down slightly, forever. - **The procedure itself.** A failover path that is never exercised takes far longer than its published number the first time it is used in anger. ## The argument against jumping straight to a second region Two honest points belong in the conversation: - **Whole-region loss is rare; your own changes are not.** Most outages come from a deployment, a configuration change or a capacity mistake, and a second region does nothing about those — it is copied there too. Money spent on safer rollouts and faster rollback often buys more availability per unit than a second region does. - **Complexity has its own failure rate.** A cross-region posture adds machinery — replication, promotion logic, traffic steering — and that machinery can fail by itself, sometimes causing the outage it was bought to prevent. That is not an argument for never doing it. It is an argument for **doing it for a stated reason**: a regulatory requirement, a contractual commitment, a concentration of revenue that cannot tolerate hours, or a genuine history of regional events in the places you operate. ## A defensible middle Many organisations land on a split rather than a single posture. The **write path for money** gets warm standby with a small agreed loss window; reporting and batch get a cold copy and a longer recovery; the customer-facing read path is served from both places because reads are easy to duplicate and hard to get wrong. Picking per capability rather than per company is usually the answer that survives contact with the budget, and it makes each choice reviewable on its own. ## How to answer this out loud Show the frame first — recovery time and acceptable loss, priced against downtime cost — then walk the postures with their real costs, then say which one you would recommend and what would change your mind. Add the two honest caveats about self-inflicted outages and about the second side rotting, because they are what a lead is expected to say and what a candidate reciting a resilience checklist will not.

  • Why is the step to active-active so much larger than the step from cold to warm standby?
    Because cold and warm differ mainly in how much capacity is already running, which is a cost dial. Active-active changes the data model: both sides accept writes, so you own conflict resolution, non-colliding identifier allocation and a consistency story per data set. That is design work in the application, not a procurement decision.
  • How would you keep a second-region posture from decaying between events?
    Make the second side carry real traffic, even a small share, so that a break is visible during normal operation rather than during an incident. Where that is impossible, make every change apply to both by construction rather than by discipline, and treat a difference between the two as a defect with an owner.
  • What would make you recommend against a second region despite the resilience argument?
    A history showing that essentially all of your downtime came from your own changes, a set of dependencies that cannot be duplicated, or a downtime cost that is comfortably below the annual price of the posture. In those cases, safer rollouts, faster rollback and a rehearsed restore buy more availability for the same money.

A spare tyre in the boot against a second car in the driveway. One is cheap, sits unused and costs you half an hour by the roadside; the other is available instantly and costs you a second of everything, including the servicing.

saying these in an interview costs you the question

  • Recommending a second region without a stated downtime or data-loss target
  • Assuming a second region costs roughly twice the first
  • Treating active-active as a scaled-up warm standby rather than a data-model change
  • Forgetting that the standby must be kept equivalent as the system changes
  • Believing a second region helps against a bad deployment
  • Choosing one posture for the whole estate instead of per capability