What is the difference between synchronous and asynchronous replication in a relational database, and what does each mean for data loss if the primary dies suddenly?
answer
- difference = when the client hears committed
- async: local durability only, non-zero RPO
- sync: extra round trip, RPO 0 for acknowledged writes
- sync standby loss can block commits
- cross-region latency is a physics floor
basics
~20 sAsynchronous: the primary confirms the commit as soon as it is durable locally and ships the change afterwards, so a sudden primary loss can lose recently committed transactions. Synchronous: the primary waits for a standby to acknowledge before confirming, so no acknowledged transaction is lost, at the cost of extra latency on every commit.
solid answer
~50 sBoth ship the same change stream; they differ in when the client is told committed. **Asynchronous**: the primary makes the change durable in its own log, answers the client immediately, and sends it to the standby afterwards. Commit latency is unaffected by the standby, and the standby's health never blocks writes. The exposure is that transactions acknowledged in the last moments before a primary failure may exist nowhere else, so a failover loses them: a non-zero recovery point objective. **Synchronous**: the primary sends the change and waits for a standby to acknowledge it before answering the client. Any transaction the client saw as committed is on at least two machines, so failover loses nothing acknowledged. The price is that every commit pays a network round trip plus the standby's disk write, and that losing the standby can stall commits unless you configure a fallback. The real question in an interview is which one per workload, and whether acknowledged means received, flushed, or applied.
go deeper
State the waiting point and the consequence: async can lose the most recent commits on a crash, sync does not but each commit is slower.
Add commit-latency numbers by distance, the acknowledgement-level nuance, and the fact that a dead synchronous standby can block writes.
Choose per workload, describe the mixed topology (sync in-region plus async cross-region), and how you would keep availability while promising zero loss.
Frame it as an RPO and latency budget tied to business cost of loss, and own the availability consequence of any zero-loss promise.
## The same pipeline, a different waiting point In both modes the primary does the same work: it executes the transaction, writes the change records to its write-ahead log, and makes that log durable on its own storage. It also streams those records to one or more standbys, which write and replay them. The only difference is whether the primary waits for the standby before telling the client the commit succeeded. ## Asynchronous replication The primary answers the client the moment its own log write is durable. Commit latency is therefore whatever the local storage costs, typically well under a millisecond on modern hardware. The standby applies changes shortly after, and the gap between them is replication lag. Consequences: - Write latency is independent of the network and of the standby, so a standby across a continent, or a temporarily slow one, costs nothing on the write path. - Standby failure is a monitoring event, not an outage: the primary keeps accepting writes. - On a sudden primary loss (host death, power cut, storage failure), everything committed but not yet transferred is gone. That is a non-zero RPO: you can measure it in transactions or in seconds of lag, but you cannot make it zero. This is the default in most deployments because it is cheap and robust, and because for many workloads losing the last second of writes in a rare hard failure is acceptable. ## Synchronous replication The primary sends the records and waits for an acknowledgement from a standby before returning success. Now every commit costs one network round trip plus whatever work the standby must do before acknowledging. Within a data centre or between availability zones in one region, that is typically a fraction of a millisecond to a couple of milliseconds; across regions it is tens of milliseconds and is dominated by the speed of light, which no tuning removes. Consequences: - Any transaction the client was told committed exists on at least two machines, so a failover to that standby preserves it: RPO of zero for acknowledged writes. - Commit throughput for a single session drops in proportion to the added latency, because each session's commits serialise on the round trip. Concurrency and group commit hide much of it in aggregate, but tail latency rises. - If the acknowledging standby disappears, commits block, because the primary can no longer satisfy the promise it makes. That availability cost has to be engineered away deliberately (a quorum of several standbys, an automatic degrade to asynchronous, or an accepted outage). ## The important nuance: what acknowledged means Synchronous is not one thing. The standby can acknowledge when it has received the bytes into memory, when it has written and flushed them to its own durable log, or when it has replayed them so queries on the standby can see them. Each level protects against a different failure and costs differently, and a candidate who says only synchronous means the standby has it has skipped the interesting part. ## An important thing synchronous replication does not give you It does not make the standby's queries current unless the acknowledgement is at the apply level. With acknowledgement at the flush level, the standby has the data durably but may not yet have made it visible, so reads there can still be stale. Synchronous replication is a durability mechanism first; read freshness is a separate decision. ## Choosing Ask what one lost second of committed writes costs. For a payments ledger, a trade capture system, or a system of record whose loss cannot be reconstructed from an upstream source, that is unacceptable and synchronous commit within a region is the answer. For event ingest, telemetry, session state, or anything replayable from an upstream log, asynchronous is right and the money is better spent elsewhere. Many systems mix the two: a synchronous standby nearby for zero-loss failover, plus an asynchronous replica in another region for disaster recovery, since paying cross-region latency on every commit is rarely justified. ## How to present it State the waiting point in one sentence, translate each mode into RPO and commit latency, then immediately raise the two follow-ups an interviewer wants: what the acknowledgement level means, and what happens when the synchronous standby is unavailable.
- Does synchronous replication guarantee that reads on the standby are up to date?Only if the acknowledgement is at the apply level. If the standby acknowledges after flushing the change to its log, the data is durable there but may not yet be replayed, so queries on the standby can still return the older value. Durability and read visibility are separate guarantees.
- Would you use synchronous replication across regions?Rarely. Cross-region round trips of tens of milliseconds are added to every commit, and no tuning removes the speed-of-light component, so interactive write paths suffer badly. The usual shape is a synchronous standby within the region for zero-loss failover plus an asynchronous replica in another region for disaster recovery, accepting a small RPO for the regional-disaster case.
Asynchronous is posting a letter and carrying on; synchronous is waiting on the phone until the other side says they have written it down.
saying these in an interview costs you the question
- Saying asynchronous means the standby might never get the data, rather than gets it slightly later
- Claiming synchronous replication eliminates all data loss, including unacknowledged in-flight transactions
- Ignoring that a lost synchronous standby can stall the primary's commits
- Assuming synchronous means the standby has applied the change and can serve it to readers
- Proposing synchronous replication across continents for a low-latency write path