Amazon Aurora replicas typically show replica lag in the tens of milliseconds, while a standard RDS MySQL read replica can fall seconds or minutes behind. What is structurally different about how an Aurora cluster stores its data and feeds its replicas?
answer
- compute and storage are separate
- log records cross the wire, not pages
- six copies, three AZs
- readers attach, they never replay
- promotion copies nothing
basics
~20 sAurora separates compute from storage: every instance in a cluster attaches to one shared distributed volume that keeps six copies across three Availability Zones. Replicas replay nothing — they read the pages the writer already wrote, so lag stays in milliseconds.
solid answer
~50 sIn a normal RDS setup each replica is a separate server with its own full copy of the data, and it has to receive and replay a log or binlog stream to stay current — so apply lag is real and can grow under load. Aurora splits the engine from storage. The writer ships only **redo log records** to a purpose-built distributed storage service, which materializes pages itself and keeps six copies of the volume spread over three Availability Zones. Every reader in the cluster attaches to that same volume, so it never applies transactions to its own data; it just receives the redo stream to invalidate or refresh pages in its buffer cache and then reads from shared storage. That is why lag is milliseconds, why you can attach up to fifteen readers, and why promoting a reader on failover copies no data at all.
go deeper
Be able to say that an Aurora cluster is several compute instances sharing one storage volume, and that only one of them can write. Knowing that readers do not each hold their own copy is the point.
Explain the mechanics: the writer ships redo log records, the storage layer materializes pages, six copies live across three Availability Zones, and a four-of-six quorum acknowledges the write. Tie each to why replica lag collapses.
Show you still design for lag. State that milliseconds is not zero, name the read-after-write hazard through the reader endpoint, and connect the shared volume to fast failover and to the separately billed storage I/O that surprises teams on their first invoice.
Own the boundary of the model. Aurora scales reads and durability, not writes — argue when a single-writer cluster is the right ceiling for a workload and what the migration path looks like when it is not, weighing I/O-Optimized versus standard billing at fleet scale.
## The design Aurora replaced A conventional managed relational database is one server with its own disks. Durability comes from writing a write-ahead log plus data pages, and high availability comes from streaming that log to another complete server that replays it onto its own copy. Every replica therefore does the same work the writer did: parse the change stream, apply it to pages, flush them. When the writer is busy, the replica's single apply path becomes a bottleneck and lag grows. That apply lag is what you are measuring when RDS reports `Seconds_Behind_Source` or a Postgres replay delay. ## What an Aurora writer actually sends Aurora keeps the SQL engine (MySQL- or PostgreSQL-compatible) but replaces the storage layer with a distributed service. On a write, the instance does **not** push data pages down to disk. It ships redo log records — the compact description of the change — to the storage fleet. The storage nodes are not dumb block devices: they apply those log records themselves and materialize the data pages in the background. The often-quoted phrase from the Aurora paper is "the log is the database". The practical effect is far less data crossing the wire per commit than an engine that must eventually write full pages, plus a binlog, plus a doublewrite buffer. ## Six copies, three Availability Zones The cluster volume is sliced into 10 GiB segments. Each segment is replicated six ways: two copies in each of three Availability Zones. Writes use a quorum — the writer sends the log record to all six and acknowledges once four have accepted it. Reads use a smaller quorum of three. Those numbers are chosen so the volume tolerates losing two copies (a whole AZ) without losing write availability, and three copies without losing read availability. Because repair happens per 10 GiB segment rather than per volume, a lost copy is rebuilt from peers quickly instead of requiring a full-volume rebuild. The volume also grows automatically in 10 GiB increments as data arrives, up to a documented ceiling in the low hundreds of TiB on current engine versions, and you never pre-provision it the way you size an EBS volume for RDS. ## What the readers do Aurora Replicas are compute-only. They attach to the same cluster volume the writer is using, so there is no second copy of the data to keep in step. The writer's redo stream is also sent to each reader, but a reader uses it only to keep its **buffer cache** honest: pages it holds that the record touches are invalidated or updated, and anything else is read fresh from shared storage. This is why lag is measured in milliseconds rather than seconds. It is not, however, zero. There is still a short window in which a reader can serve a page slightly behind the writer, so read-after-write is **not** guaranteed through the reader endpoint. Design for that the same way you would with any asynchronous replica: route the read-your-own-write path at the writer, or wait on a replication position. ## The consequences you are being asked about - **Up to 15 Aurora Replicas** per cluster, all sharing one volume, so adding read capacity adds compute cost but not storage cost. - **Fast failover.** Promotion picks a healthy reader (promotion tiers decide which one) and re-points the cluster endpoint. Nothing is copied, so a failover typically completes in well under a minute rather than the minutes an instance rebuild would take. - **Backups are continuous** to Amazon S3 at the storage layer, so taking one does not tax the database instance. - **Reads scale, writes do not.** There is exactly one writer per Aurora cluster. If write throughput is the problem, more readers will not help; you are looking at sharding or a different data store. - **I/O is a billed dimension** in the standard configuration — each storage read and write operation costs money. A very chatty workload can end up paying more for I/O than for compute, which is what the Aurora I/O-Optimized cluster configuration exists to flatten in exchange for a higher instance and storage rate. ## What this does not change Commit latency still crosses Availability Zones — the four-of-six acknowledgement is a network round trip inside the Region, not free. And the engine on top is still MySQL or PostgreSQL, with the same transactions, locking, and query planning. Aurora changed where bytes live and who applies them, not how SQL behaves.
- If Aurora replicas lag by only milliseconds, can I safely send a read immediately after a write to the reader endpoint?No. Milliseconds is small but not zero, and the reader endpoint may hand you a page the writer has just moved past. Treat it as an asynchronous replica: send the read-your-own-write path to the writer, or wait until the reader has caught up to the commit position before reading.
- During an Aurora failover, why is there no risk of the promoted reader missing recent commits the way an asynchronous RDS replica can?Because durability lives in the shared cluster volume, not in the instance. A commit is acknowledged only after the storage quorum accepts the log records, so any committed transaction is already in the volume that the promoted reader attaches to. Promotion changes which instance may write, not which data exists.
- Aurora bills storage I/O as a separate line item. When does that dominate the bill and what is the lever?Write-heavy or scan-heavy clusters with a working set larger than the buffer cache generate huge read and write I/O counts, and on the standard configuration each is charged. Check the billed I/O against instance cost in Cost Explorer; if I/O is a large share, the Aurora I/O-Optimized cluster configuration trades a higher instance and storage rate for zero per-I/O charges.
saying these in an interview costs you the question
- Says Aurora replicas use binlog or WAL shipping like RDS
- Claims Aurora replicas are perfectly consistent with the writer
- Thinks each Aurora instance holds its own full copy of the data
- Assumes a commit waits for all six storage copies
- Believes adding Aurora readers increases write throughput