skip to content

Amazon RDS offers both a Multi-AZ DB instance deployment and a Multi-AZ DB cluster deployment. What is the difference between them, and what would make you pick the cluster?

level: seniorimportance: nice to knowfreq 30%

answer

  1. count the nodes: two versus three
  2. one of them has a reader endpoint
  3. semi-synchronous, quorum of one
  4. tens of seconds versus minutes
  5. engine support is not universal

basics

~20 s

A Multi-AZ DB instance has one hidden standby in a second Availability Zone. A Multi-AZ DB cluster has a writer plus two readable standbys across three AZs, with a reader endpoint and faster failover. Pick the cluster when you need shorter failover and read capacity from the HA copies.

solid answer

~50 s

The instance deployment is the classic two-node arrangement: one primary, one standby in another Availability Zone, no endpoint on the standby, and a failover AWS documents at roughly one to two minutes. The **Multi-AZ DB cluster** deployment provisions three instances across three AZs — one writer and two *readable* standbys — using the engine's own semi-synchronous replication, where a commit is acknowledged once at least one standby confirms it rather than waiting for both. You get a writer endpoint and a reader endpoint, so the HA copies also absorb read traffic, and failover is materially faster — AWS documents typically under 35 seconds. The cost is a third instance, a narrower support matrix (RDS for MySQL and PostgreSQL, on specific engine versions and storage types), and standbys that can still lag, so reads through the reader endpoint are not guaranteed current.

go deeper

for a junior

Know that RDS Multi-AZ comes in two shapes and that only the cluster version gives you standbys you can actually query.

for a middle

Explain the node counts, the writer and reader endpoints, and why a commit in a Multi-AZ DB cluster waits for one standby rather than both.

for a senior

Demonstrate the operating judgment: match the deployment to a recovery-time target, account for the third instance's cost against read replicas you would have provisioned anyway, and check the engine support matrix before promising it.

for a principal

Own the fleet-level policy — which service tiers justify a tens-of-seconds failover budget, and whether consolidating HA and read capacity into one construct is worth the coupling it creates.

## Two things AWS gave the same prefix The naming is genuinely confusing, and that is part of why the question gets asked. "Multi-AZ" in RDS refers to two distinct deployment options that you choose between when you create or modify a DB instance, and they differ in node count, in what those nodes can do, and in how quickly the writer role moves. ## Multi-AZ DB instance deployment Two nodes: a primary you connect to, and a standby in a different Availability Zone whose storage RDS keeps synchronously in step. The standby has no endpoint and answers no queries. On failure, maintenance, or a forced reboot, RDS moves the writer endpoint's DNS record to the standby. AWS's documented typical failover is 60–120 seconds. This option is available across the RDS engine range — MySQL, MariaDB, PostgreSQL, Oracle, SQL Server — and it is what most people mean when they say "we run Multi-AZ". ## Multi-AZ DB cluster deployment Three nodes across three Availability Zones: one writer and two standbys that are **readable**. Replication is the engine's own, and it is *semi-synchronous* — the writer acknowledges a commit once at least one of the two standbys has confirmed the change, rather than blocking on both. That quorum-of-one design is what makes the deployment tolerate one slow or unavailable standby without stalling commits, while still holding a confirmed second copy of every acknowledged write. Because the standbys are real, queryable instances, the cluster exposes more than one endpoint: - a **writer endpoint**, which always points at the current writer; - a **reader endpoint**, which distributes connections across the readable standbys. Failover is faster than the instance deployment — AWS documents typically under 35 seconds — because the promotion target is already a running, caught-up instance rather than a passive storage replica, and because the cluster's endpoints are managed as a unit. ```bash aws rds create-db-cluster \ --db-cluster-identifier prod-cluster \ --engine postgres \ --db-cluster-instance-class db.m6gd.large \ --storage-type io1 --allocated-storage 400 --iops 3000 ``` ## What you give up The cluster is not free of constraints, and naming them is what separates a real answer from a brochure answer: - **Three instances, not two.** You pay for a third node, so the HA premium is higher unless the readable standbys genuinely displace read replicas you would otherwise have run. - **A narrower support matrix.** Multi-AZ DB clusters are offered for RDS for MySQL and RDS for PostgreSQL, on particular engine versions, instance classes, and provisioned-IOPS-class storage. If you run Oracle or SQL Server on RDS, the option does not exist for you. - **Reads are still asynchronous in effect.** "Readable standby" does not mean "current": a session that writes through the writer endpoint and immediately reads through the reader endpoint can see stale data, exactly as with a read replica. Route reads there only where staleness is acceptable. - **It is not Aurora.** A Multi-AZ DB cluster keeps three independent copies of the data using engine-level replication; it does not have Aurora's shared distributed storage layer, and it is a different product with different scaling and failover characteristics. ## When to choose which Choose the **instance** deployment when your engine is not MySQL or PostgreSQL, when a one-to-two-minute failover window is acceptable, or when your read traffic is small enough that the writer handles it. It is the cheaper, broader, better-understood default. Choose the **cluster** when a shorter recovery time genuinely matters — a failover budget measured in tens of seconds rather than minutes — and when you have read traffic that can tolerate a little staleness, so the two standbys earn their cost by serving it instead of idling. The decision is essentially: are you paying for a third node anyway in the form of read replicas? If yes, the cluster consolidates HA and read capacity into one construct with managed endpoints. If no, the instance deployment plus a replica when you need one is usually the more economical shape. ## The interview trap The trap is answering as if there were only one Multi-AZ, and then being unable to explain how anyone gets a *reader* endpoint without Aurora. Knowing that the cluster deployment exists, that its standbys are queryable, and that its failover target is measured in tens of seconds is the differentiator here.

  • Why can a Multi-AZ DB cluster keep accepting writes when one standby becomes unreachable?
    Because its replication is semi-synchronous with a quorum of one: the writer acknowledges a commit as soon as either standby confirms it, not both. Losing one standby leaves the other able to satisfy that quorum, so commits continue. Losing both would remove the confirmed second copy, which is the condition the three-AZ layout is designed to make unlikely.
  • If a Multi-AZ DB cluster already gives you two readable instances, when would you still add a read replica?
    When you need read capacity beyond those two, a different instance size for a specific workload, or a copy in another Region for disaster recovery or read locality. The cluster's standbys exist primarily for availability, so treating them as elastic read capacity couples your HA posture to your reporting load.

saying these in an interview costs you the question

  • A Multi-AZ DB cluster is just another name for Aurora
  • The standbys in a Multi-AZ DB cluster are always fully current
  • Multi-AZ DB clusters are available for every RDS engine
  • A Multi-AZ DB instance also exposes a reader endpoint
  • Three nodes automatically means three times the write throughput

context