In Amazon RDS, how does a Multi-AZ deployment differ from a read replica, and which of the two would you add to relieve a database whose CPU is saturated by reporting queries?
answer
- one is invisible, one is addressable
- synchronous standby versus asynchronous copy
- who gets an endpoint?
- DNS repoint versus explicit routing
- availability, not read capacity
basics
~20 sMulti-AZ keeps a standby copy in another Availability Zone purely for failover, and that standby serves no traffic. A read replica is an asynchronous, separately addressable copy you can query. Reporting load needs a replica; Multi-AZ buys availability, not read capacity.
solid answer
~50 sThey solve different problems. A Multi-AZ **DB instance** deployment provisions a standby in a second Availability Zone that RDS keeps in sync at the storage layer; you cannot connect to it, and it exists so that RDS can fail the single DB endpoint over to it when the primary, its AZ, or a maintenance action takes the writer down. Failover is automatic and the endpoint DNS name is repointed, so applications reconnect rather than reconfigure. A **read replica** is a separate DB instance fed asynchronously by the engine's own replication, with its own endpoint that you must explicitly send queries to; it can live in another AZ or another Region, and RDS never promotes it for you. So for reporting queries starving the writer of CPU, you add a read replica and point the reporting traffic at its endpoint. The two are complementary, and a replica can itself be Multi-AZ.
go deeper
Be ready to say plainly that Multi-AZ is about staying up and read replicas are about serving more reads, and that only the replica gives you an endpoint you can query.
Explain the mechanics: a synchronous, non-addressable standby with DNS-based endpoint failover versus an asynchronous separate instance with its own endpoint that your application must target on purpose.
Show the operational judgment — sizing a replica independently, isolating analytical traffic from the writer, and making clients survive a failover through connection validation and DNS TTL handling rather than manual intervention.
Own the tradeoff between the write-latency and cost of synchronous standbys and the staleness that asynchronous replicas introduce, and set a fleet-wide rule for which tiers of service get which combination.
## The question behind the question Interviewers ask this because the two features look similar on the console page — both create a second copy of your database in another Availability Zone — and candidates who have only read the marketing routinely answer "Multi-AZ, for high availability and to spread the read load". The second half of that sentence is wrong, and it is wrong in a way that produces a real incident: you pay double for an instance you believe is absorbing traffic, and it absorbs none. ## What a Multi-AZ DB instance deployment actually is An Availability Zone (AZ) is one or more physically separate data centres inside an AWS Region, with independent power and networking. A Multi-AZ *DB instance* deployment gives you one primary and one standby in a different AZ of the same Region. RDS keeps the standby's storage in step with the primary synchronously — a commit is not acknowledged to your client until the write is durable on both sides — which is why Multi-AZ adds a little write latency and why it protects you against losing committed transactions if the primary's AZ disappears. The crucial property is that **the standby is not addressable**. It has no endpoint you can connect to, it cannot serve a `SELECT`, and it does not reduce load on the primary by one query. Its whole purpose is to be there when RDS decides a failover is required: an AZ outage, primary host or storage failure, an instance-class change, an OS or engine patch, or a manual `reboot-db-instance --force-failover`. RDS then makes the standby the primary and repoints the DNS record behind your single writer endpoint at the new host. AWS documents this as typically completing in 60–120 seconds for a Multi-AZ DB instance deployment. Your application does not change its connection string; it does have to survive dropped connections and reconnect, which is why short DNS TTL handling and connection-pool validation matter on RDS. ## What a read replica is A read replica is a *separate DB instance* that RDS creates from a snapshot of the source and then keeps up to date using the engine's native asynchronous replication. Because the replication is asynchronous, the replica acknowledges nothing back to the writer, so it adds no write latency — and it can trail the source. It has its own endpoint, its own instance class (it need not match the source), and it can be placed in another AZ or another Region. Asynchronous also means **you must send reads there deliberately**. RDS does not route traffic for you: your application, your ORM's read/write splitting, or a proxy layer decides which queries go to the replica endpoint. And RDS will never promote a read replica automatically when the source dies; promotion is an explicit `promote-read-replica` API call that turns the replica into an independent, writable instance and permanently severs replication. ```bash aws rds create-db-instance-read-replica \ --db-instance-identifier prod-db-reports \ --source-db-instance-identifier prod-db \ --db-instance-class db.r6g.xlarge ``` ## Choosing for the reporting workload A writer pinned at high CPU by analytical scans is a read-capacity problem, so the answer is a read replica, sized independently of the writer and given its own endpoint for the BI tool. That isolates the scan-heavy workload's CPU and buffer-cache pressure from transactional traffic. Multi-AZ would change nothing about the CPU picture — it would only mean that when the overloaded primary falls over, a second, equally overloaded instance takes its place. ## They are not alternatives Production systems usually run both: Multi-AZ on the writer so an AZ failure is a blip instead of an outage, plus one or more read replicas for reporting or read-heavy endpoints. You can also enable Multi-AZ *on a replica*, so the replica itself has a standby and survives losing its own AZ. What Multi-AZ gives that replicas do not is a guarantee about committed data at failover time; what replicas give that Multi-AZ does not is queryable capacity and cross-Region reach. ```bash aws rds modify-db-instance \ --db-instance-identifier prod-db \ --multi-az \ --apply-immediately ``` ## The one-line summary to say out loud Multi-AZ is a *durability and availability* feature with an invisible standby and one endpoint; read replicas are a *capacity and locality* feature with visible endpoints and no automatic failover. Anything about read scaling is answered with replicas.
- If you already run Multi-AZ, does adding a read replica change how a failover behaves?Not by itself. Failover still moves the writer endpoint to the standby, and the replica keeps replicating from whatever instance now holds the source role. The replica is not promoted, is not consulted, and its own endpoint does not move. If you want the replica to become the writer you must call promote-read-replica explicitly, and that is a one-way operation.
- Can the standby in a Multi-AZ DB instance deployment be used to take backups or run maintenance?RDS uses it internally — backups on Multi-AZ deployments are taken without the I/O pause that a Single-AZ instance can see, and OS or engine patching is applied to the standby first and then failed over. But none of that is something you drive or query; from your side the standby remains an implementation detail with no endpoint.
- Your application connects fine but takes minutes to recover after an RDS failover. What is usually to blame?Client-side caching of the endpoint's DNS resolution and connection pools that hand out dead sockets. RDS moves the DNS record, so a JVM or resolver caching the old address keeps dialling the old host. Fix it by honouring short DNS TTLs, validating pooled connections before use, and keeping connection timeouts short enough that a stuck socket fails fast.
saying these in an interview costs you the question
- Multi-AZ standby serves read traffic to offload the primary
- RDS automatically promotes a read replica when the primary fails
- Read replicas are synchronous so they can never be behind
- Multi-AZ and read replicas are alternatives, never used together
- Applications must change the connection string after a failover