In a single-leader replication setup, what's the practical difference between synchronous and asynchronous follower replication, and what does each cost you?
answer
- ack timing, not ordering
- sync = durability, cost latency+availability
- async = speed, risk lost acks on crash
- semi-sync = one sync follower + rest async
basics
~20 sSynchronous means the leader waits for a follower to confirm it got the write before telling the client 'done' - safer but slower. Asynchronous means the leader tells the client 'done' immediately and sends the copy in the background - faster but a crash can lose the last few writes.
solid answer
~40 sSynchronous replication makes the leader block the write's acknowledgment until at least one follower confirms it has durably applied the write, guaranteeing that write survives a leader crash at the cost of added write latency and reduced availability (if that follower is unreachable, writes stall). Asynchronous replication acknowledges the client as soon as the leader itself commits, then ships the write to followers in the background; this keeps write latency low and the leader available even if followers are slow or down, but a leader crash before the write replicates means it's lost even though the client was told it succeeded. Many systems use a middle ground - semi-synchronous - where one follower is synchronous and the rest async, balancing durability against latency/availability.
go deeper
Should state the core difference: sync waits for a follower before confirming, async confirms immediately. Doesn't need to know semi-sync.
Should articulate the latency vs durability trade-off in both directions and know that pure sync-to-all hurts availability.
Should know semi-synchronous as the practical default, reason about which failure scenario each mode protects against, and connect follower placement to real durability guarantees.
Should evaluate this choice per data class within one system (e.g., financial writes sync, activity-feed writes async), and account for failure-domain correlation and automated demotion/promotion policy design.
## What the choice actually decides Synchronous and asynchronous replication describe **when the leader is allowed to tell a client 'your write succeeded,'** and that timing decision is really a trade-off between durability, latency, and availability. Mechanically: - **Synchronous replication** - the leader appends the write to its own log, sends it to one or more designated synchronous followers, and only returns success to the client after receiving acknowledgment that those followers have durably persisted the write too. - **Asynchronous replication** - the leader appends the write locally, immediately returns success to the client, and then streams the write to followers on its own schedule, with no client-visible wait for their acknowledgment. The leader and followers still apply writes in the same order in both modes - the difference is purely about when the client is told 'done,' not about ordering. ## Why the distinction exists This distinction exists because a system has to decide how much it is willing to trust a single machine (the leader) with data before it's copied anywhere else. If a leader acknowledges a write and then immediately crashes before ever replicating it, that write is gone from the cluster's perspective - the client believes it succeeded, but no surviving node has it. Synchronous replication closes that window by refusing to acknowledge until the data exists on at least one other machine, so a leader crash right after acknowledgment cannot lose data (assuming the synchronous follower is promoted). Asynchronous replication accepts that risk in exchange for speed and availability. ## The trade-off, on both axes The trade-off is genuinely two-sided and shows up on both axes. | Axis | Which mode wins, and why | |---|---| | Latency and availability | Favor **asynchronous**: the client doesn't pay network round-trip cost to a follower on every write, and the leader keeps accepting writes even if every follower is down, slow, or network-partitioned - the leader just gets further ahead and followers catch up later. | | Durability | Favors **synchronous**: an acknowledged write is guaranteed to exist on more than one node, so failover after a leader crash never silently drops a 'successful' write. | But synchronous replication pays for that guarantee with tail latency (every write waits for the slowest required follower's round trip) and a nastier availability failure mode - if the synchronous follower becomes unreachable, the leader can't safely acknowledge any write at all, so the whole system can grind to a halt on a single follower outage unless it's designed to fail open (demote to async) or fail closed (block writes) under that condition. ## The middle ground In practice almost nobody runs pure synchronous replication to every follower, because losing any one of N followers would stall all writes. The common middle ground is **semi-synchronous replication**: one follower (or a rotating subset) is designated synchronous and must acknowledge before the client gets a response, while the remaining followers replicate asynchronously in the background. This bounds the durability guarantee ('the write survives loss of the leader alone') without paying the cost of waiting on every follower, and if the synchronous follower goes down, the system typically promotes a different follower to synchronous status so writes can keep flowing with only a brief gap in the stronger guarantee. ## Failure modes Failure modes differ sharply between the two. - **With pure async**, the dangerous scenario is: leader acknowledges a write, leader crashes before shipping it to any follower, a follower is promoted, and the write is permanently gone - the client was lied to. This is the classic cause of 'I saved my order and it disappeared after a failover' incidents. - **With sync (or semi-sync)**, the dangerous scenario is availability collapse: a synchronous follower experiences a slow disk, GC pause, or network blip, and every write in the system now waits on it, turning a single slow machine into a cluster-wide write stall; operators have to watch for this and either time out and demote the laggy follower or accept the stall as the cost of the durability guarantee. ## Where it shows up A concrete real-world example: MySQL semi-synchronous replication requires at least one replica to acknowledge receipt of a transaction before the primary commits, which is a deliberate middle ground between MySQL's default fully-async replication and a fully synchronous setup; PostgreSQL similarly supports naming synchronous standbys to require acknowledgment from one or more named standbys, with the well-known operational hazard that if the named synchronous standby is unreachable, write transactions on the primary hang until it recovers or an operator reconfigures the setting. This is precisely why teams pick asynchronous replication for latency-sensitive, high-write-volume systems where losing a few recent writes on a rare crash is an acceptable risk, and reserve synchronous or semi-synchronous replication for data where losing an acknowledged write (e.g., a financial transaction) is unacceptable even at the cost of higher write latency and a narrower availability envelope.
- What happens to write availability if the single synchronous follower in a semi-synchronous setup becomes unreachable?Depends on configuration: many systems block new write acknowledgments until the follower comes back or is failed out of the synchronous role, which can stall all writes; better-designed systems detect the outage and automatically demote to async (or promote a different follower to the synchronous role) after a timeout, trading a temporary durability weakening for continued availability.
- Does synchronous replication protect against losing data if the whole data center hosting leader and follower goes down together?No - if the synchronous follower is co-located with the leader (same rack/data center), a correlated failure like a power outage takes both out together, and the durability guarantee only holds against single-node failure. True multi-region durability requires the synchronous follower to be in a genuinely independent failure domain, at the cost of higher replication latency.
- Why might a team choose asynchronous replication even for financially sensitive data?If the business tolerance is a rare, bounded amount of lost recent writes versus the cost of materially higher write latency or availability risk on every transaction, async can be the right call - some systems compensate by writing critical events to a separate durable queue or by using async replication only for read replicas while the source of truth uses a different durability mechanism such as a distributed consensus log.
Synchronous replication is like a chef who won't tell the waiter 'order's out' until the sous-chef confirms they've also written it in the backup ticket book - slower per order but no order is ever forgotten if the chef trips and falls. Asynchronous is the chef shouting 'order's out!' immediately and scribbling the backup ticket whenever there's a spare second - fast, but if the chef collapses right after shouting, that ticket may never make it to the book.
saying these in an interview costs you the question
- Thinks synchronous replication changes write order, not just ack timing
- Believes async replication can never lose acknowledged data
- Doesn't know sync replication to every follower kills availability if any one is down
- Unaware of semi-synchronous as a middle ground
- Assumes sync replication protects against correlated/data-center-wide failures automatically