Your production Amazon RDS PostgreSQL database runs Multi-AZ in one Region. An interviewer asks how you would survive the loss of that entire Region. What AWS mechanisms do you reach for, and what does each cost you in recovery point and recovery time?
answer
- zones are not Regions
- one option is already running
- promotion is a decision, not an event
- and it cannot be undone
- the network and keys must exist first
basics
~20 sMulti-AZ never leaves its Region, so regional survival needs a cross-Region mechanism: a cross-Region read replica you promote manually, cross-Region automated backup replication, or copied snapshots. Replicas recover fastest with the smallest data loss; snapshot copies are cheapest and slowest.
solid answer
~50 sFirst, name the gap: a Multi-AZ deployment spans Availability Zones inside one Region, so it does nothing for a Region-level event. RDS offers three cross-Region tools, and they trade cost against recovery. A **cross-Region read replica** is a live instance in the second Region fed asynchronously; recovery means calling `promote-read-replica`, which makes it standalone and writable, and you lose whatever had not yet replicated. **Cross-Region automated backup replication** copies the backup set so that point-in-time restore is available in the second Region — cheaper, but recovery means provisioning and restoring a new instance, so it is measured in tens of minutes to hours. **Copied snapshots** are the cheapest and coarsest. Whichever you choose, the DR Region also needs the VPC, subnets, security groups, parameter groups and KMS keys in place beforehand, and RDS never promotes anything for you.
go deeper
Know that Multi-AZ stays inside one Region and that surviving a Region loss needs something cross-Region, such as a replica or copied snapshots.
Explain how a cross-Region read replica is created and promoted, how replicated automated backups differ, and why one recovers in minutes while the other recovers in tens of minutes or more.
Show the operating judgment: watch replication lag as the live measure of your recovery point, pre-provision network, parameter groups and destination KMS keys, and rehearse the promotion and repointing end to end.
Own the tiering decision across the portfolio — which systems justify a continuously running second-Region copy and its data-transfer cost, and what recovery objectives the business has actually signed up to fund.
## Start by naming what Multi-AZ does not do An Availability Zone failure and a Region failure are different events with different answers. Multi-AZ places a standby in a second AZ of the *same* Region and fails the writer endpoint over automatically. Every part of that — the standby, the endpoint, the automated backups — lives in one Region. If the requirement is to survive losing that Region, Multi-AZ is not an answer, and saying so first is what makes the rest of the answer credible. ## Mechanism one: cross-Region read replica You create a read replica in another Region from the source instance's ARN: ```bash aws rds create-db-instance-read-replica \ --db-instance-identifier prod-db-dr \ --source-db-instance-identifier arn:aws:rds:eu-west-1:123456789012:db:prod-db \ --region us-east-1 ``` The replica is a full, running DB instance receiving changes asynchronously over the AWS backbone. It is queryable, which means the DR copy can also serve read traffic local to that Region in normal times. Recovery is a deliberate act: ```bash aws rds promote-read-replica --db-instance-identifier prod-db-dr ``` Three properties matter in an interview. Promotion is **manual** — RDS never promotes a replica on its own, so a real plan needs a human decision or your own automation with a clear trigger. Promotion is **one-way**: the promoted instance becomes an independent, writable database and the replication link is gone for good; you cannot demote it back, and re-establishing the original topology later means building a fresh replica in the other direction. And promotion **involves a brief restart**, so it is fast but not instantaneous. The recovery point is bounded by how far behind the replica was at the moment the Region became unreachable — which is why the replica's replication metrics are something you watch continuously rather than during an incident. The recovery time is the shortest of the three options because the instance already exists and is already hydrated. ## Mechanism two: cross-Region automated backup replication RDS can replicate the automated backup set — the snapshots plus the transaction logs — to a second Region, giving you point-in-time restore there. You pay for storage rather than for a running instance, which for a large database is a substantial difference. The price is recovery time. Recovery means calling a restore in the second Region, which **provisions and hydrates a brand-new DB instance**. For a multi-terabyte database that is tens of minutes at best, and often longer, before anything can serve traffic. The recovery point is governed by how current the replicated logs are — good, but the total outage is dominated by the restore, not by the gap. ## Mechanism three: copied snapshots `copy-db-snapshot` into the DR Region, on whatever cadence you schedule. Cheapest and simplest, and the recovery point is as coarse as your copy interval — if you copy nightly, you can lose a day. It is the right answer for databases that can be rebuilt or whose data is reconstructible from an upstream system, and the wrong answer for anything transactional. ## The part candidates forget: everything around the database A promoted replica or a restored instance is useless in a Region where nothing else exists. A workable DR plan pre-provisions: - **Network**: a VPC, subnets across AZs, a DB subnet group, route tables and security groups that permit your application tier. - **Configuration**: parameter groups and, where the engine uses them, option groups — a restored instance defaults to the AWS defaults unless you name yours. - **KMS keys**: an encrypted database's snapshots and replicated backups must be encrypted with a key **in the destination Region**, so that key and its policy have to exist before the incident, not during it. - **Naming and routing**: something the application resolves — a Route 53 record or a config value — that can be repointed without a code deploy. - **Credentials**: the secret the application uses, replicated to the DR Region, since Secrets Manager secrets are regional. ## Choosing, and saying why The honest framing is a tiering exercise. For a payments-grade system the recovery-point budget is small enough that only a live cross-Region replica qualifies, and you accept the cost of a second running instance plus the data-transfer charges for continuous cross-Region replication. For an internal service that can be down an hour, replicated automated backups give point-in-time recovery at storage prices. For a reconstructible dataset, nightly snapshot copies are proportionate. What an interviewer is listening for is that you attach numbers to each option, that you name promotion as manual and irreversible, that you do not confuse Multi-AZ with regional resilience, and that you say the plan is only real if it has been rehearsed — a DR mechanism nobody has exercised is a hypothesis.
- Why is promoting a cross-Region read replica described as irreversible?Because promotion severs the replication link permanently: the instance becomes a standalone writable database with its own history, and RDS offers no way to demote it back into a replica of the original. Returning to the previous topology means creating a fresh replica in the opposite direction and cutting over again, which is a planned migration rather than a rollback.
- What continuous signal tells you whether your cross-Region DR promise is still true?The replica's replication metrics in CloudWatch, watched with an alert threshold. The recovery point you can honestly promise is bounded by how far behind the DR copy is right now, so a replica that has been drifting for hours quietly invalidates the plan. Pair that with periodic rehearsals that measure how long promotion and repointing actually take.
- Why can an encrypted RDS database complicate cross-Region disaster recovery?Because KMS keys are regional. A replicated backup or copied snapshot must be encrypted with a key that exists in the destination Region, and the operation needs permission to use it. If the key and its policy are not provisioned in advance, the copy or restore fails at exactly the moment you need it, so key setup belongs in the DR build, not the runbook.
saying these in an interview costs you the question
- Multi-AZ already protects against a Region outage
- RDS promotes the cross-Region replica automatically when the Region fails
- A promoted replica can be turned back into a replica afterwards
- Restoring from replicated backups is as fast as promoting a replica
- The DR Region needs nothing prepared in advance