With three or more standbys, how does quorum-based synchronous commit (waiting for any k of n acknowledgements) differ from naming a single synchronous standby, and what does it change for durability and availability?
answer
- one named standby = second single point of failure
- ANY k of n: tolerate n-k losses
- latency = k-th fastest ack
- nobody designated current: compare positions at failover
- not consensus, still one writer; fence + rewind
basics
~20 sWith one named synchronous standby, that standby is a single point of failure for writes and the only node guaranteed to hold every acknowledged commit. Quorum commit waits for any k of n standbys, so any n minus k can be slow or dead without blocking writes, and commit latency follows the k-th fastest rather than one fixed node.
solid answer
~50 sNaming one synchronous standby couples writes to that specific machine: if it is slow, every commit is slow; if it dies, commits stall until you reconfigure. It also fixes which node is guaranteed current, which is convenient for failover but fragile. Quorum commit says wait for acknowledgements from any k of these n standbys. Two effects follow. Availability improves: with k = 1 and n = 3 you tolerate two failures without blocking writes. Latency improves in the tail: you wait for the k-th fastest acknowledgement, so one straggler no longer sets your commit time. The cost is that you no longer know in advance which standby holds the newest data, so failover must pick the most advanced one, comparing replication positions. That is why quorum commit belongs with a failover controller that compares positions, not with a hand-picked promotion. Also note k of n is about durability breadth, not about ordering across nodes.
code
text · 5 lines# wait for any one of three standbys (tolerates two failures)
synchronous_standby_names = 'ANY 1 (az_a, az_b, az_c)'
# wait for the highest-priority available standby only
synchronous_standby_names = 'FIRST 1 (az_a, az_b, az_c)'go deeper
Know that you can wait for one named standby or for any k of several, and that the second tolerates failures better.
Explain the durability, availability, and tail-latency effects of k of n and give a sensible k and n for a three-zone deployment.
Own the consequences: position-based promotion, fencing, repairing divergent standbys, and the cross-zone latency implications of raising k.
Set k and n from an explicit failure model and latency budget, and be clear that this is replication breadth, not consensus, so the failover controller carries the correctness burden.
## The problem with one named standby The simplest synchronous setup nominates a specific standby. The primary waits for that node before acknowledging any commit. Two properties follow, one good and one bad. Good: you always know which node has every acknowledged transaction, so failover is unambiguous. Bad: that node's latency becomes your commit latency, its garbage collection pause is your write hiccup, and its death stops writes altogether until an operator or controller changes the configuration. You have converted one availability risk (the primary) into two. ## Quorum commit Quorum commit generalises the wait: acknowledge the client once any k of the n listed standbys have acknowledged. PostgreSQL expresses this in `synchronous_standby_names` with `ANY k (list)`, and the older priority form `FIRST k (list)` waits for the k highest-priority available ones. Other systems use group- or set-based settings with the same idea. Durability: an acknowledged transaction exists on the primary plus at least k standbys. With k = 1 that is two copies, which is the usual zero-data-loss requirement. With k = 2 you also survive losing the primary and one standby simultaneously. Availability: writes continue as long as k standbys are healthy, so the system tolerates n minus k failures or slow nodes. Choosing n comfortably larger than k is what buys the availability that a single named standby lacks. Latency: you wait for the k-th fastest acknowledgement, not for one fixed node. This is the tail-latency argument for quorum commit, and it is significant in practice, since one standby's momentary I/O stall no longer shows up as a global write spike. ## The cost: who is current? With one named standby the answer is fixed. With `ANY 1 (a, b, c)` any of the three might hold the latest commit and the others might not. That means: - **Failover must compare positions.** The controller has to query each candidate's replication position and promote the most advanced one. Promoting an arbitrary node can lose acknowledged transactions, which destroys the guarantee you paid for. - **Divergent standbys must be repaired.** After promotion, the standbys that were ahead of the new primary on some records must be rewound or re-seeded before they can follow it. - **Fencing matters.** The old primary must be prevented from accepting writes after promotion; otherwise two nodes accept writes and the acknowledged-set argument collapses. This is why quorum commit is a package deal with an automated failover controller, not a standalone setting. ## Sizing k and n Start from the durability requirement: how many simultaneous node losses must not lose an acknowledged transaction? That fixes k. Then choose n larger than k by the number of concurrent failures or maintenance events that must not block writes. Common shapes: k = 1, n = 2 for zero loss on primary failure with one spare; k = 1, n = 3 spread across availability zones so a whole zone can go away; k = 2, n = 3 when a correlated double failure must not lose data, accepting that only one node may be down. Watch the cross-zone latency: with `ANY 1` over three zones, commit latency is the nearest zone's round trip, which is usually acceptable, whereas `ANY 2` forces the second-nearest and can meaningfully raise write latency. ## What quorum commit is not It is not a consensus protocol. The primary is still the single writer, and the quorum is only about how widely a commit record is replicated before acknowledgement; there is no voting on which transactions exist, and no automatic leader election unless a separate controller provides one. That is precisely why fencing and position-based promotion have to be handled outside the replication setting. It is also not a read-freshness mechanism. A standby that acknowledged a write may not have replayed it, so reads there can still be stale unless the acknowledgement level is apply. ## How to present it Contrast the two configurations on three axes (durability, availability, latency), then volunteer the consequence interviewers are listening for: with a quorum you no longer know which standby is current, so failover must compare replication positions, fence the old primary, and repair divergent standbys.
- With ANY 1 of three standbys, how does the failover controller pick which one to promote?It queries each candidate's replication position and promotes the most advanced one, because the acknowledged transaction may exist on only one of them. It must also fence the old primary so it cannot keep accepting writes, and then rewind or re-seed any standby that had received records the new primary lacks.
- Does increasing k from 1 to 2 improve availability?No, it reduces it. Requiring more acknowledgements means fewer standbys may be unavailable before commits block, and commit latency rises to the second-fastest acknowledgement. Increasing k buys durability against correlated multi-node failure; availability comes from increasing n relative to k.
saying these in an interview costs you the question
- Describing quorum commit as a consensus or voting protocol
- Promoting an arbitrary standby after failover despite a quorum configuration
- Believing more synchronous standbys automatically means higher write availability, regardless of k
- Forgetting to fence the old primary, allowing two writable nodes
- Assuming a standby that acknowledged the commit can already serve it to readers