Why does automated database failover require a majority quorum, and what is a witness (or arbiter) node used for in a cluster that would otherwise have an even number of members?
answer
- floor(N/2)+1, only one majority exists
- two nodes cannot fail over safely
- odd sizing: 4 buys nothing over 3
- witness votes, stores no data
- third independent failure domain
basics
~20 sOnly a majority of a fixed member set may elect a primary, because at most one majority can exist, so a minority partition can never promote. A witness is a cheap vote-only member that makes the member count odd without storing data.
solid answer
~60 sQuorum turns a subjective judgement (is the primary dead?) into an arithmetic one (do I have more than half the votes?). Define a fixed membership of N members; any decision requires floor(N/2)+1 of them. Two disjoint groups can never both hold a majority, so at most one side of any partition can elect a primary. The minority side must refuse to promote and, if it holds the old primary, demote itself. That is why a two-node cluster cannot fail over safely: after a partition each node sees one vote out of two, so either both promote (split-brain) or neither does (no availability). The fix is a third vote. A witness or arbiter is a lightweight member that participates in elections but stores no database data, so it can run on a small VM in a third availability zone. Critically, the witness must be placed where it fails independently. A witness in the same rack, host, or datacentre as one database node just makes that site's failure the cluster's failure.
go deeper
Know that promotion needs more than half the members to agree, and that this is why clusters have three nodes rather than two.
Do the arithmetic out loud, explain the witness as a vote-only tiebreaker, and note that even member counts add no tolerance.
Focus on placement and failure domains, on what the cluster does when quorum is lost, and on the gap between winning an election and stopping the old writer.
Own the availability-versus-divergence tradeoff explicitly: state that quorum loss means read-only by design, and define the manual override procedure and who is authorised to run it.
## The arithmetic Quorum-based failover fixes membership up front: the cluster is exactly these N members. Any state change that matters (electing a leader, granting a lease) requires votes from a strict majority, floor(N/2)+1. The key property is trivial but decisive: two disjoint sets cannot both be majorities of the same N. Therefore, however the network partitions, at most one partition can elect a primary. Split-brain becomes impossible by counting rather than by hoping detection was accurate. The price is availability. A partition that leaves no side with a majority leaves the cluster with no writable primary. Quorum systems deliberately choose unavailability over divergence. ## Fault tolerance by member count - N=2: majority is 2. Either node's loss stops writes; a partition gives each side one vote. Tolerates zero failures for automated failover. - N=3: majority is 2. Tolerates one failure. - N=4: majority is 3. Still tolerates only one failure, while adding a node that can fail. Even counts buy nothing. - N=5: majority is 3. Tolerates two failures. This is why clusters are sized odd. Going from three to four members increases failure surface without increasing tolerance. ## What a witness is A witness (arbiter, quorum device, tiebreaker) is a member that votes but carries no data replica and can never be promoted. Its purpose is purely to make N odd and to sit in a third failure domain. Because it stores no data it needs almost no disk or CPU, so a small instance or container is enough. Typical shapes: - Two data nodes plus one witness: majority is 2, so the surviving data node plus the witness can promote while a lone isolated node cannot. - A three-node consensus store (etcd/ZooKeeper/Consul) holding the leadership state, with database nodes as clients. Here the quorum lives in the store rather than in the database processes. Be precise about which layer holds the quorum. In the Patroni pattern the quorum is etcd's; Patroni agents merely read and write the leader key. Losing etcd quorum means no agent can hold or acquire leadership, which is the intended safe state. ## Placement is the part people get wrong Quorum only helps if members fail independently. Common mistakes: - Witness on the same hypervisor, rack, or power feed as a data node, so one hardware failure removes two votes. - Witness in one of two datacentres, making that datacentre's loss fatal to the whole cluster while the other site's loss is survivable. That is an asymmetric HA design people rarely intend. - Witness on a network path that shares fate with the replication link, so a link failure removes it exactly when it is needed. The useful rule: place the tiebreaker in a third, independent failure domain, and check whether the majority still exists after each single-domain loss you claim to survive. ## What quorum does not give you Quorum decides who may become primary. It does not by itself stop the old primary from writing. A partitioned old primary may be entirely unaware it lost the election, and clients pinned to it keep writing. That is why quorum is paired with self-demotion on lease expiry and with fencing, and why the routing layer must be updated as part of promotion rather than after it. Quorum also says nothing about data currency. A member can win an election while lagging; choosing a leader and choosing an up-to-date leader are separate concerns, which is why managers additionally compare replay positions and may refuse to promote a candidate beyond a configured lag threshold. ## Interview framing State the arithmetic, give the two-node failure case, explain the witness as a vote-only third failure domain, note odd sizing, and close with the caveat that quorum prevents a second election but not a second writer, which is what fencing is for.
- Your cluster is three nodes and two of them are lost. What should the survivor do?It must refuse to be primary. With one vote out of three it has no majority, so promoting would risk split-brain if the other two are alive behind a partition. The correct behaviour is to stay read-only or stop accepting writes until an operator either restores quorum or performs a deliberate, documented manual override that accepts the risk.
- Does adding a fourth node make the cluster more available?No. With four members the majority is three, so it still tolerates only one failure, and you have added another component that can fail or partition. Even-sized member sets also create the possibility of a 2-2 split where neither side can act. Add nodes in odd increments, or add a vote-only witness instead of a fourth data node.
saying these in an interview costs you the question
- Believing a two-node cluster can fail over automatically and safely
- Adding a fourth node to improve fault tolerance
- Placing the witness in one of the two datacentres it is supposed to arbitrate between
- Assuming quorum alone prevents the old primary from continuing to accept writes
- Confusing quorum for leadership with synchronous-commit quorum for durability