skip to content

You are asked to give an Aurora-backed service a second AWS Region for disaster recovery and low-latency local reads. How does Aurora Global Database replicate between Regions, what recovery point and recovery time should you expect, and what does it force the application to handle?

level: principalimportance: should knowfreq 38%

answer

  1. replication runs in the storage layer
  2. one writable Region at a time
  3. switchover loses nothing, failover loses seconds
  4. the database's RTO is not the system's RTO
  5. stale reads cross a continent now

basics

~20 s

Aurora Global Database replicates at the storage layer to secondary Regions with typical lag under a second, so cross-Region reads are local and the writer pays almost no CPU for it. There is still one writer Region, and the application must handle stale reads and a Region-level promotion.

solid answer

~50 s

Aurora Global Database attaches read-only Aurora clusters in other Regions to a primary cluster, and the replication happens in the storage layer rather than in the database engine — dedicated infrastructure ships the redo stream, so the writer's CPU is barely affected and typical cross-Region lag is under a second. Each secondary Region serves local reads at local latency. Writes still go to one Region: either the application routes them there, or you enable write forwarding so a secondary accepts writes and forwards them to the primary, paying a cross-Region round trip and choosing a session consistency level. For recovery, distinguish a *planned switchover*, which is coordinated and loses no data, from an *unplanned failover*, which promotes a secondary and typically costs seconds of RPO with an RTO usually inside a minute. The application must handle stale reads, a global writer that can move, and idempotent retries across the cutover.

code

bash · 3 lines
bash
aws rds switchover-global-cluster \
  --global-cluster-identifier my-global \
  --target-db-cluster-identifier arn:aws:rds:eu-west-1:111122223333:cluster:eu-secondary

go deeper

for a junior

Know that Aurora Global Database adds read-only clusters in other Regions, that only one Region accepts writes, and that replication lag is typically under a second.

for a middle

Explain that replication happens at the storage layer rather than in the engine, why that keeps the writer's overhead low, and what write forwarding does and costs.

for a senior

Separate planned switchover (no data loss) from unplanned failover (seconds of RPO), and show that the real RTO includes DNS, application capacity, and consumers in the second Region — not just the promotion.

for a principal

Own the decision itself: what risk justifies a second Region at all, how warm the standby is kept against a stated RTO and budget, and how the promise is rehearsed and measured rather than asserted in a runbook.

## What is actually being replicated A plain Aurora cluster spans three Availability Zones inside one Region. Aurora Global Database extends that to Regions: one primary cluster that accepts writes, and up to five secondary Regions, each holding a read-only Aurora cluster over its own storage volume with its own readers. The important architectural point — and the one interviewers are probing — is **where the replication runs**. This is not a logical replication slot or a binlog feed consumed by a downstream engine. Aurora replicates through purpose-built infrastructure at the storage layer, streaming the redo record stream to the secondary Region's storage. The database instances in the primary Region are not doing the shipping, so the writer's CPU and its commit latency are essentially unaffected by adding a second Region. Typical replication lag is under a second across a Region pair, though that is a function of distance and write volume, not a guarantee — measure it, do not promise it. Because the secondary volume is a real Aurora volume, you can attach readers to it and they behave like any Aurora reader: low local lag, local latency for users in that geography. ## Writes still have one home Global Database is not multi-master. There is exactly one writable cluster at a time. Two ways to live with that: **Route writes explicitly.** The application knows the primary Region and sends writes there, accepting the cross-Region round trip for write paths while keeping reads local. Clean, explicit, and it makes the latency asymmetry visible in your code rather than hidden. **Write forwarding.** A secondary cluster accepts a write, forwards it to the primary, and returns when it is done. The wire latency is the same as routing it yourself; what you gain is a single connection story for the application. What you must then decide is the session's read consistency after such a write — the available modes range from eventual (fastest, may not see your own forwarded write locally) through session-level (your session sees its own writes) to a global mode that waits for the secondary to catch up to the primary's position. Each step tightens correctness and adds latency. ## RPO and RTO, precisely Collapse these two into "failover" in an interview and you have missed the question. - **Planned switchover.** Both Regions are healthy and you coordinate the move — for a Region migration, a compliance drill, or a maintenance event. The primary stops accepting writes, the chosen secondary catches up fully, and roles swap. **No data loss.** This is the operation you should be practising on a schedule. - **Unplanned failover.** The primary Region is unreachable, and you promote a secondary knowing it may not have received the last records in flight. Expect an RPO measured in seconds — bounded by the replication lag you were actually running at, which is why you should be alarming on that lag — and an RTO usually inside a minute for the promotion itself. ```bash aws rds switchover-global-cluster \ --global-cluster-identifier my-global \ --target-db-cluster-identifier arn:aws:rds:eu-west-1:111122223333:cluster:eu-secondary ``` The minute-scale RTO is the *database's* number. Your real RTO includes everything else: DNS or global traffic routing moving to the new Region, application capacity already warm there, secrets and IAM roles present, queue consumers repointed. A database that promotes in forty seconds into a Region with no running application servers has an RTO of however long it takes to scale those up. ## What the application has to own 1. **Stale reads are now cross-Region stale.** Sub-second is small, but a user who writes in one Region and reads in another can see the past. Route read-your-own-write traffic to the primary, or use the consistency controls write forwarding offers. 2. **The writer can move.** Endpoints are per-cluster, so a promotion changes which Region's cluster endpoint is writable. Configuration must be able to follow that without a code deploy — service discovery, parameter store lookup, or a global endpoint abstraction. 3. **Idempotent retries.** Everything in flight at the cutover fails. Retries must not double-charge anyone. 4. **Schema and version drift.** Both clusters run the same engine version; upgrades are coordinated across the global cluster rather than done Region by Region on a whim. ## The costs to name Cross-Region replication is billed — replicated write I/O and inter-Region data transfer — and each secondary Region runs real instances you pay for whether or not they serve traffic. A cheap DR posture keeps a minimal secondary and accepts a slower ramp; an expensive one keeps the secondary warm at production size. That choice, made explicitly against a stated RTO, is what separates a considered multi-Region design from a checkbox one. If the requirement is really "survive an AZ failure", a single-Region Aurora cluster already does that, and Global Database is the wrong tool for the stated risk.

  • Why does adding a secondary Region cost the Aurora writer so little, compared with adding a cross-Region logical replica?
    Logical replication is engine work: the primary decodes changes and feeds a subscriber, consuming CPU and a replication slot that can retain logs when the subscriber lags. Aurora Global Database ships the redo stream from the storage layer through dedicated replication infrastructure, so the writer instance is largely uninvolved.
  • Your requirement is to survive the loss of a data centre. Is Global Database the right answer?
    Usually no. A single Aurora cluster already keeps six copies across three Availability Zones and fails over inside the Region, which covers a data-centre loss. Global Database addresses a Region-level event or a latency requirement in a distant geography — a materially more expensive problem to solve, so make sure that is the one you have.
  • How would you keep the promised cross-Region RPO honest rather than aspirational?
    Alarm on the replication lag metric continuously, because your unplanned-loss window equals the lag at the moment of failure. Then rehearse: run a planned switchover on a schedule, time the full application cutover, and treat the measured number — not the database's promotion time — as the RTO you publish.

saying these in an interview costs you the question

  • Calls Aurora Global Database multi-master or active-active
  • Quotes the database promotion time as the whole system's RTO
  • Treats planned switchover and unplanned failover as the same operation
  • Assumes write forwarding removes cross-Region write latency
  • Reaches for a second Region to survive an Availability Zone failure

context