skip to content

A team is deciding between synchronous and asynchronous replication for a database that must tolerate a primary node failure. What availability and durability guarantees does each give, and what specific failure mode does asynchronous replication risk that synchronous replication avoids?

level: middleimportance: must knowfreq 68%

answer

  1. synchronous = ack after replica persists, RPO=0, higher latency
  2. async = ack immediately, background ship, data-loss window on crash
  3. semi-sync = one replica sync, rest async, middle ground
  4. quorum sync avoids one slow replica blocking all writes
  5. correlated failure defeats sync guarantee if replica shares failure domain

basics

~20 s

Synchronous replication waits for a copy to confirm the write before telling the client it succeeded, so no acknowledged write is ever lost, but it's slower and can stall if a replica is unreachable. Asynchronous replication confirms instantly and copies data in the background, so it's fast, but a crash right after can lose the last few writes.

solid answer

~50 s

Synchronous replication requires at least one, or a quorum of, replicas to durably persist a write before the primary acknowledges it to the client, guaranteeing zero data loss on failover at the cost of added write latency and reduced availability if a required replica is unreachable. Asynchronous replication acknowledges the client immediately after the primary persists locally and ships the write to replicas afterward, giving low, predictable latency and full availability of the primary regardless of replica health, but risking data loss: if the primary crashes before a replica catches up, the acknowledged write is gone once a replica is promoted. This can also produce a silent revert, where a promoted stale replica no longer reflects a write the client was told succeeded. The choice is a direct expression of the durability/latency/availability trade-off, often resolved by hybrid designs like semi-synchronous replication that require only one replica to ack.

go deeper

for a junior

Can state, in plain terms, that synchronous replication is safer but slower and asynchronous is faster but riskier.

for a middle

Can explain precisely what guarantee is lost with async replication and why quorum sync is used instead of all-replica sync.

for a senior

Can reason about semi-synchronous replication as a deliberate middle ground, and can size an appropriate replication strategy against a workload's actual RPO/RTO requirements rather than defaulting to one mode.

for a principal

Can identify subtler failure modes like correlated failure domains defeating a synchronous guarantee, and can design replica placement and failover automation as part of the overall fault-tolerance architecture.

## What the acknowledgment actually promises **Replication** is the core redundancy mechanism distributed systems use to survive node failure: by keeping multiple copies of data on independent machines, the system can continue serving, and if designed carefully keep, data even after losing one or more of those machines. But how replication is performed - specifically, whether the primary waits for replicas before acknowledging a write - determines exactly what guarantee the write succeeded actually carries, and this is one of the sharpest trade-offs in fault-tolerant system design. ## Synchronous replication In **synchronous replication**, the primary node does not tell the client a write succeeded until at least one, or more robustly a quorum of, replicas has durably received and persisted that write. This means that at the instant the client is told success, the data is provably present on more than one machine, so if the primary immediately crashes, a replica can be promoted with zero loss of that write - this durability guarantee is what's usually meant by **RPO=0**, a recovery point objective of zero data loss. What it costs: - The cost is paid on the write path: every write's latency is bounded below by the round trip to the slowest replica required for acknowledgment, not just the primary's local disk. - Worse, if a required replica becomes unreachable, the primary faces an **availability dilemma**: it must either block or refuse new writes until the replica is reachable again, sacrificing availability to preserve the durability guarantee, or silently downgrade to asynchronous behavior, which breaks the guarantee it claims to offer. Systems that need true zero-loss synchronous replication typically use a quorum rather than all replicas specifically so that one slow or dead replica out of several doesn't halt the whole system. ## Asynchronous replication **Asynchronous replication** takes the opposite stance: the primary persists the write locally, immediately acknowledges the client, and ships the write to replicas in the background, independent of the client-facing response. This decouples client-perceived latency entirely from replica health and network conditions - writes are fast and the primary stays fully available even if every replica is temporarily unreachable, since replication is not on the acknowledgment critical path. This introduces two failure modes: - **A data-loss window.** The risk this introduces is precisely a data-loss window: if the primary crashes after acknowledging a write but before that write has propagated to any replica, the write is gone the moment a replica is promoted to primary, even though the client was already told it succeeded. This isn't a rare edge case in high-throughput systems - it's a routine, expected possibility whose size is governed by the replication lag at the moment of failure. - **A silent revert.** A related and subtler failure mode is a silent revert: a client is told write X succeeded, the primary crashes before shipping X, a lagging replica is promoted, and that replica now serves reads that don't reflect X at all - to the client, data that was confirmed durable has simply vanished, which can be worse than an outright failure because there's no error to react to. ## Matching the mode to the workload The choice between the two is a direct, concrete instance of the durability/latency/availability trade-off that underlies most of fault-tolerant systems design, and it's usually resolved by matching the replication mode to the actual **RPO** and **RTO** requirements of the workload rather than picking one universally. | Workload | Mode it usually picks | The reasoning | |---|---|---| | Financial ledgers, payment systems, and anything where a confirmed success must be an unbreakable promise | typically require synchronous, or synchronous-to-quorum, replication | despite the latency cost, because silent data loss is unacceptable | | High-throughput logging, analytics ingestion, or systems where a small, bounded loss window is tolerable | often use asynchronous replication | in exchange for much higher write throughput and lower tail latency | ## Semi-synchronous, the middle ground Many production systems land on a middle ground called **semi-synchronous replication**: the primary waits for acknowledgment from at least one replica, not all and not necessarily a full quorum, before confirming to the client, then ships to the remaining replicas asynchronously. This bounds the worst-case data loss to at most what wasn't yet on that one synchronous replica, while avoiding the full latency and availability cost of waiting on every replica or a full quorum. It's a pragmatic compromise, though it doesn't provide the same formal guarantee as full quorum-synchronous replication - if that one synchronous replica and the primary fail simultaneously, a **correlated failure** such as both being in the same rack or availability zone, the guarantee is void, which is why the physical placement of the synchronous replica matters just as much as the replication mode itself.

  • Why do teams typically use a quorum of replicas for synchronous acknowledgment rather than requiring all replicas to confirm?
    Requiring all replicas means a single slow or unreachable replica halts every write, turning one failure into a total availability outage. A quorum, such as a majority, still guarantees the durability property needed for safe failover while tolerating some number of replicas being down or lagging, trading a small amount of theoretical extra durability for much better availability.
  • How does placing the synchronous replica in the same availability zone as the primary undermine the durability guarantee?
    If both the primary and its synchronous replica share a failure domain such as a rack, power feed, or availability zone, a correlated event can take out both simultaneously, and the acknowledged-but-unreplicated-elsewhere writes are lost despite the synchronous configuration. The guarantee synchronous replication provides assumes independent failure of the primary and the replica it waited on, which physical placement must actually deliver.
  • What operational signal would tell you asynchronous replication lag has grown dangerously large?
    A widening gap between the primary's write position, for example a write-ahead log sequence number, and the replica's applied position, typically exposed as a replication-lag metric in seconds or bytes. A growing lag directly increases the size of the potential data-loss window on failover, so teams often alert on lag thresholds and may pause promotions or throttle writes if lag exceeds an acceptable RPO.

Synchronous replication is like not telling a customer their bank transfer went through until a second bank branch has also confirmed receiving the funds - slower, but the confirmation is trustworthy. Asynchronous replication is like confirming the transfer instantly and mailing the paper record to the backup branch later - fast, but if the main branch burns down before the mail arrives, that transfer is gone even though the customer was told it succeeded.

saying these in an interview costs you the question

  • Claims asynchronous replication never loses data as long as it eventually catches up
  • Thinks synchronous replication requires literally every replica to acknowledge, with no notion of quorum
  • Doesn't recognize that failover during replication lag is exactly when async data loss manifests
  • Assumes replication mode has no effect on write latency
  • Ignores failure-domain placement when discussing whether a synchronous replica actually protects against the primary's failure

context