skip to content

Zone & Region Outages

Surviving the loss of one zone against the loss of a whole region: duplicated capacity, replication lag, and who decides to fail over. Probed because the two postures cost very differently.

on this pageshow

questions

6

A checkout service runs in three availability zones and one zone goes dark at peak — what keeps serving, and what had to be true beforehand?

level: juniorimportance: must knowfreq 72%

answer

  1. who notices the zone is gone
  2. health check, then traffic removal
  3. stateless survives, stateful is promoted
  4. survivors carry the lost share
  5. headroom bought before, not during

basics

~20 s

Instances in the two healthy zones keep serving, but only if the traffic entry point health-checks the dead ones out, the writable data copy is not stranded in the lost zone, and the survivors already had spare capacity.

solid answer

~40 s

Spreading compute across zones is the easy half. When one zone goes dark, something has to notice: a health check has to fail the instances there and the traffic entry point in front of the service has to stop sending requests to them, which takes as long as the check's interval and threshold allow. The survivors then carry the lost zone's share, so they need headroom that is already running. The stateful pieces decide the rest — a service holding no state keeps serving untouched, while a database whose only writable copy sat in that zone is unavailable for writes until a copy elsewhere is promoted. So: stateless things with headroom recover by themselves; anything with exactly one home has to be failed over.

go deeper

for a junior

Recall that a zone is a failure domain inside a region and that identical copies in other zones keep answering once the failed ones are health-checked out of the traffic entry point. Say that the database is a separate question from the web tier.

for a middle

Explain the detection path in order: the health check fails, the entry point stops selecting those instances, in-flight requests fail, the survivors absorb the load. Name what needed configuring in advance rather than saying it is automatic.

for a senior

Show that you have watched one happen: the cache miss storm, the broken connection pools, the retry burst, and the capacity you could not get when everyone in the region wanted it at once. Say which of those you sized for.

for a principal

Frame it as which failures you have chosen to absorb and what that choice costs every month in idle capacity and standby data, and be explicit about where you deliberately accepted downtime instead of paying to remove it.

## What losing a zone actually looks like An availability zone is a **failure domain inside a region** — its own power, cooling and network path, sited so that one physical event should not take two of them, and close enough to its neighbours that a write can be confirmed in both without a human noticing the delay. Losing one is rarely a clean power-off. It usually looks like a slice of your fleet going quiet or slow while the rest is fine: instances stop answering, or answer slowly; a mounted volume stops acknowledging writes; connections hang instead of being refused. The first practical consequence is that your system has to **decide** the zone is gone, and nothing about that decision is instant. ## What keeps serving on its own, and what does not | Part of the stack | Keeps serving if… | Goes down if… | |---|---|---| | Stateless request handlers | copies already run in the other zones and the entry point stops routing to the dead one | the whole group happened to be placed in one zone | | The traffic entry point | it is itself spread across zones | it is a single instance living in the lost zone | | Read-only data copies | at least one copy is in a surviving zone and clients can reach it | every copy lived in the lost zone | | The writable copy of a database | a standby elsewhere is promoted into the writable role | there is exactly one writable copy and it was there | | Cache and session state | it is re-derivable, or held in more than one zone | it was the only home of state the request needs | | Queued and scheduled work | the broker and its leader span zones | the broker node or the job leader sat in the lost zone | The pattern is simple: **anything with many interchangeable copies survives automatically; anything with exactly one home needs a decision**. Interviewers ask this question because candidates routinely describe the first half and forget the second. ## The three things that had to be true in advance 1. **Something noticed.** A health check, run from outside the failing zone, has to mark those instances unhealthy, and the entry point has to stop selecting them. The time that takes is the check's interval multiplied by its unhealthy threshold, plus any draining, plus whatever a client caches about where the service lives. Until that completes, a share of requests is sent into the dark and fails. 2. **The survivors had capacity already running.** Load does not shrink because a zone did. The remaining zones absorb the lost share immediately, while replacement capacity takes minutes to boot, warm and pass its own health check — and everyone else in the region is asking for capacity at the same moment. 3. **Everything with one home had a defined way to get another.** For a database that means a standby ready to be promoted; for a queue, a broker that already spans zones; for a scheduled job, a leader election that can run without the lost member. ## Where the outage still reaches you - Requests already in flight to that zone fail, so callers need retries and the handlers need to be safe to retry. - Open connections to a database that gets promoted elsewhere break; pools reconnect, and requests in the gap fail. - Work that was accepted but recorded only in that zone stalls or is lost. - Latency shifts, because callers now land on instances that may be further from the data they read. - A cache that lost a third of its entries sends a burst of misses at whatever sits behind it, exactly when that thing is busiest. ## How to answer this out loud Split the system in two — the part with interchangeable copies and the part with exactly one home — and say what happens to each, in what order, and how long the change takes. Then name the thing you had to buy in advance: spare capacity that is idle most of the time, and a standby that costs money to keep current. Be careful not to slide from a lost zone into a lost region: they are different sizes of problem with very different answers, and conflating them is the fastest way to lose the thread.

  • How long does traffic actually take to stop reaching the dead zone?
    As long as the health check's interval multiplied by its unhealthy threshold, plus any connection draining, plus whatever the caller caches about where the service lives. That is usually seconds to a couple of minutes, and every request sent during that window fails. Shortening the interval detects faster but makes false removals more likely, so the setting is a trade-off rather than a maximum.
  • What keeps failing after traffic has been removed from the lost zone?
    Anything that was mid-flight or single-homed: requests already sent, open connections to a database that was promoted elsewhere, queued work recorded only in that zone, and background jobs whose leader lived there until a new one is elected. Cache misses also surge, because a share of the warm entries disappeared with the zone.
  • Does spreading across zones protect against a bad deployment?
    No. A zone spread protects against a failure domain being lost, not against a change you made yourself, which is rolled out to all zones and breaks all of them. Those are different risks with different controls — staged rollouts and fast rollback for the change, zone spread for the infrastructure event.

saying these in an interview costs you the question

  • Assuming multi-zone deployment makes every component fail over automatically
  • Believing traffic stops reaching dead instances the instant the zone fails
  • Assuming the survivors can simply be scaled up during the incident
  • Treating a replica in a second zone as if it were a backup
  • Assuming a whole region cannot be lost, only a zone
  • Forgetting that the database also has to move, not just the compute
open as a page

Three zones each run an order service at 70% of capacity at peak; one zone is lost — what happens, and what sizing rule prevents it?

level: middleimportance: must knowfreq 61%

basics

~20 s

Each survivor is asked for about 105% of its own capacity, so the service queues, slows and sheds load at the worst moment. Surviving one zone loss means capping each zone's peak utilisation at (N-1)/N — roughly 66% with three zones.

open as a page

Your order database has a synchronous standby in a second zone and an asynchronous copy in a second region — what differs when you promote each?

level: middleimportance: should knowfreq 57%

basics

~20 s

Promoting the in-zone synchronous standby loses no acknowledged writes, because every commit already waited for it. Promoting the cross-region asynchronous copy loses whatever the lag held — seconds or minutes of confirmed orders that exist nowhere else.

open as a page

Everything in your checkout path is deployed across three zones, yet losing one zone took checkout down — what class of dependency explains that?

level: seniorimportance: should knowfreq 49%

basics

~20 s

Something in the path had exactly one home in the lost zone — a single writable database, a broker node, a lock or leader, a session cache, an outbound path, or a fleet whose spread had quietly drifted into one zone. The path is only as multi-zone as its least-spread hop.

open as a page

For a cross-region failover of a checkout stack, would you automate the decision or require a human to declare it, and why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Automate inside a region, where a synchronous standby loses nothing and the move is cheap; require a declaration across regions, where promoting a lagging copy accepts real data loss and the failure signal is often ambiguous. Automation is only safe with independent observers and fencing.

open as a page

Leadership asks whether checkout should survive losing an entire region — how would you frame the choice against a second zone, and what does each posture cost?

level: principalimportance: should knowfreq 38%

basics

~20 s

Frame it as recovery time and acceptable data loss priced against the cost of each posture. Multi-zone is cheap and lossless; a second region costs duplicated capacity, duplicated data and continuous engineering to keep both true, and rises steeply from cold standby to active-active.

open as a page