skip to content

You must run a highly available relational database across two data centres with automatic failover. How do you place voting members so that losing one site does not either strand the cluster read-only or allow two primaries, and what tradeoff would you accept?

level: principalimportance: should knowfreq 28%

answer

  1. two domains, no symmetric majority
  2. 2-1 split: one site is decisive
  3. third domain for a vote-only witness
  4. or explicit manual, fenced cross-site promotion
  5. leadership quorum is not sync-commit durability

basics

~20 s

Two sites cannot give symmetric automatic failover: whichever side holds the majority survives, the other cannot promote. Either add a third independent site or region for the tiebreaking vote, or accept an asymmetric design where one site is primary-capable and the other requires a manual, fenced promotion.

solid answer

~1 min

With exactly two failure domains, any majority is unbalanced. If votes split 2-1 across the sites, losing the two-vote site leaves the survivor unable to promote; losing the one-vote site is survivable. There is no placement that makes both site losses automatically recoverable, and pretending otherwise is how people end up disabling quorum and inviting split-brain. The realistic options: 1. **Third domain for the tiebreaker.** Put the witness or the third consensus-store member in a third site, region, or cheap cloud VM. This restores symmetric automatic failover and is nearly always the right answer, since the witness needs almost no resources, only independence and modest latency. 2. **Asymmetric by design.** Site A holds the majority and runs automatic failover within itself; site B is a DR target promoted only by a deliberate, human-authorised, fenced procedure with a documented RPO. Availability inside a site is automatic; cross-site is a decision, not an event. 3. **Never**: two votes each side, a witness co-located with a database node, or automation configured to promote without quorum. I would also separate concerns: quorum placement decides who may write; synchronous commit decides how much data survives, and cross-site latency usually pushes that to asynchronous plus a stated RPO.

go deeper

for a junior

Recognise that three votes across two sites is unbalanced and that a third independent location is the standard fix.

for a middle

Enumerate the placements and their failure outcomes, and explain why a co-located witness defeats the purpose.

for a senior

Produce a failure matrix per domain, cover fencing and rebuild after a site returns, and separate leadership quorum from synchronous-commit durability.

for a principal

Own the tradeoff and communicate it: pick third-domain or explicit asymmetry, attach RTO/RPO/rebuild numbers, and forbid quorum overrides as an incident-time practice.

## Why two sites is the hard case Quorum works because two disjoint groups cannot both be majorities. With three independent failure domains that maps cleanly onto geography: one member each, and any single site loss leaves two of three. With two domains, every allocation is lopsided: - 2 members in A, 1 in B: losing B is fine; losing A leaves one vote out of three, so no promotion. Site A is a single point of failure for automated write availability. - 2 and 2: any site loss leaves 2 of 4, not a majority, so a site loss stops writes entirely; also a partition gives a 2-2 tie. - 1 and 1: neither side can ever act alone. This is arithmetic, not a product limitation, and no cluster software escapes it. The frequent bad response is to disable quorum checks or lower the required votes during an incident, which converts a partition into two live primaries. ## Option 1: buy a third failure domain The cheapest genuine fix is a vote-only member somewhere independent: a third small site, a colo, or a small cloud instance in a different provider or region. It stores no data, so cost and bandwidth are negligible; what matters is that its connectivity does not share fate with the inter-site link. Verify that: a witness reachable only through the same circuit that connects A and B is not a third domain. Latency matters modestly. The witness participates in consensus writes, so its round trip is added to leadership operations, not to user transactions, provided it never holds data. Tens of milliseconds is usually acceptable; if the store's election timeouts are tighter than the link's jitter, you will get leadership churn. Check the failure matrix explicitly: for each of the three domains, does a majority survive its loss, and can the surviving majority reach a data-carrying, sufficiently current candidate? A quorum that survives but has only the witness and a badly lagging replica is a promotion you may not want to happen automatically. ## Option 2: accept asymmetry deliberately If a third domain is genuinely impossible, be explicit rather than clever. Site A carries the quorum and runs automated failover between its own nodes, covering the common failures: a host dies, a disk fails, a process crashes. Site B is a replica for disaster recovery, promoted only by a human-run runbook that includes verifying A is truly gone, fencing it (power or network, and at minimum revoking its route), and accepting the RPO implied by the replication mode. The honest statement to stakeholders is that host failure is automatic and site failure is a decision with a stated time to restore and a stated data-loss bound. That is a much better place to be than automation that silently promotes across sites on a link flap. ## Separating the two questions people conflate Leadership quorum and durability are different knobs. Placing votes decides who may write. Synchronous commit decides whether a transaction is acknowledged before a remote copy has it. Across a WAN, synchronous commit adds the round trip to every write, so most systems run cross-site asynchronously and publish an RPO of seconds. If the business demands zero data loss cross-site, that cost lands on transaction latency, and you should say so numerically rather than promising both. ## Failback and dual-site operations Also plan for what happens after: the failed site returns with a stale, possibly diverged primary. It must not rejoin as a writer, and it will need a rewind or rebuild. If your inter-site bandwidth cannot rebuild a multi-terabyte replica in an acceptable window, that constraint should shape the design before the incident, for example by keeping a local backup at each site to seed from. ## Interview framing State the arithmetic first so it is clear the constraint is structural, then present exactly two credible designs (third domain, or explicit asymmetry), then name what you refuse to do (equal votes per site, co-located witness, quorum overrides in automation). Finish by separating leadership placement from durability mode and attaching numbers: RTO, RPO, and rebuild time.

  • Management refuses a third location. What do you put in the design document?
    That intra-site failover is automatic and cross-site failover is manual, with named approvers and a runbook that mandates fencing the old site before promotion. I would state the resulting numbers: expected RTO for a site loss, the RPO implied by asynchronous cross-site replication, and the time to rebuild the failed site afterwards. Making the asymmetry explicit is what prevents someone from quietly enabling automatic cross-site promotion later.
  • Someone proposes lowering the required votes during an outage so the surviving site can promote. What is your response?
    It is exactly the action that produces split-brain, because the reason the votes are missing may be a partition rather than a dead site. If a human genuinely must override, it has to follow a procedure that first proves or forces the other side down, by power, network isolation, or storage revocation, and it must be a one-way decision with the old site rebuilt afterwards rather than a config knob left in place.

saying these in an interview costs you the question

  • Claiming a two-site cluster can have symmetric automatic failover
  • Splitting votes evenly between two sites
  • Placing the tiebreaker in one of the two sites it arbitrates between
  • Planning to lower the quorum requirement during an incident
  • Promising both zero data loss and low write latency across a WAN without stating the cost

context