skip to content

questions

5

In a replicated database cluster with one primary and several standbys, what is split-brain, how does it arise during a failover, and what damage does it do to the data?

level: juniorimportance: must knowfreq 55%

answer

  1. two writable primaries at once
  2. no heartbeat != dead
  3. partition, freeze, GC pause
  4. divergent history, no merge
  5. quorum + fencing, not better timeouts

basics

~20 s

Split-brain is two nodes both believing they are primary and accepting writes, usually after a network partition hid a still-running primary. Their histories diverge, and reconciling them means discarding writes the application already saw committed.

solid answer

~50 s

Split-brain means the cluster ends up with two writable primaries at once. Failure detection is only observation: standbys and monitors see heartbeats stop, but no heartbeat can mean the primary crashed, the network partitioned, or the primary froze under load, a long checkpoint, or an IO stall. If the monitoring side promotes a standby while the old primary is alive and still serving the clients on its side of the partition, both nodes accept transactions. From that moment the two nodes have divergent committed histories: different rows, different sequence values, conflicting unique keys. Single-primary replication has no merge semantics, so recovery is not a merge, it is choosing one node as truth and throwing the other side's writes away even though users were told they committed. That is silent data loss plus broken invariants: duplicate order numbers, double payments. The cure is not better detection. It is requiring a majority quorum to promote and fencing the old primary so it cannot keep writing.

go deeper

for a junior

Be able to define it crisply: two primaries accepting writes at once after a partition, leading to diverging data that cannot be merged.

for a middle

Explain why detection is ambiguous (partition vs freeze vs crash) and name the concrete corruption: sequence collisions, unique-key conflicts, lost acknowledged commits.

for a senior

Move quickly to prevention and cleanup: quorum-gated promotion, leases, fencing, and a realistic recovery plan including rebuilding standbys that followed the loser.

for a principal

Frame it as a durability and correctness risk traded against availability, and set the policy: what the cluster does when it cannot reach quorum, and who signs off on the recovery choice.

## The setup that makes it possible A relational HA cluster has one primary that accepts writes and one or more standbys that replay its change stream (WAL/redo/binlog). Only one node may accept writes, because the replication stream is a single linear history: every standby assumes it is replaying one authoritative sequence of changes. Nothing in the engine merges two write streams. Failover means choosing a standby, stopping its replay, and making it writable. The dangerous part is the decision, not the promotion. ## Why detection is fundamentally ambiguous Nodes decide a peer is dead by absence of evidence: heartbeats stop, a TCP connection breaks, a health query times out. All of these are indistinguishable between: - the primary process actually crashed or the host died; - the network between observer and primary partitioned, while clients on the primary's side still reach it happily; - the primary is alive but unresponsive: swap storm, saturated IO, a long checkpoint, a VM pause, a stop-the-world GC in the layer in front of it. The second and third cases are the split-brain generators. In case two, the primary keeps committing user transactions while the observers conclude it is gone. In case three, the primary may unfreeze seconds after promotion and resume writing, unaware anything happened. ## What actually breaks Call the old primary A and the promoted standby B. After promotion both accept writes: - Rows diverge. A row updated on A and differently on B has two committed values, both durable. - Surrogate keys collide. Sequences and auto-increment counters continue independently, so both nodes hand out id 5001 to different orders. - Uniqueness is violated cluster-wide even though each node's constraint held locally. - Application side effects double up. If the app on each side charged a card or emitted an event, that already left the database. When the partition heals there is no automatic reconciliation in single-primary replication. Operators pick a survivor and rebuild the loser from a fresh base backup, so the loser's committed transactions are lost. Recovering them means manual forensic extraction from WAL/binlog and hand-merging, which is slow and often impossible without violating business invariants. Meanwhile any standby that was still following A now has a history incompatible with B and must also be rewound or rebuilt. A second, subtler harm: clients. Old connections, cached DNS, and a proxy pointing at A can keep routing writes to the demoted node long after promotion, so split-brain can persist even if the cluster software thinks it resolved. ## Why it is not solved by better monitoring No timeout distinguishes slow from dead; that is a property of asynchronous networks, not a tuning gap. Longer timeouts reduce false promotions but raise downtime, and shorter ones do the opposite. The workable answers change the shape of the problem: 1. Make promotion require agreement from a majority of a fixed member set, so a minority partition can never elect anyone. 2. Make leadership a time-bounded lease that the primary must keep renewing, so a primary that loses contact demotes itself before anyone else is promoted. 3. Fence the old primary: power it off, revoke its storage access, or drop it out of the routing layer before the new primary opens for writes. In practice these are combined: a majority quorum store holds a leader lease, the leader demotes itself on lease loss, and a fencing action covers the case where it cannot. ## What to say in an interview Define it in one sentence, name the partition/freeze cause, state that the damage is divergent committed history with no merge path, and then name quorum plus fencing as the structural fix. Candidates who only say we should monitor better have missed the point.

  • If both nodes accepted writes for two minutes, how do you recover?
    You choose one node as authoritative, usually the one clients were actually routed to or the one with more valuable transactions, and rebuild the other from a base backup or rewind it to the divergence point. The loser's committed transactions are gone unless you manually extract them from its WAL/binlog and replay them as compensating business transactions, which requires application-level judgement about duplicates. Any standby that followed the losing node must be rebuilt too.
  • Would synchronous replication have prevented the split-brain?
    Not by itself. Synchronous replication bounds data loss on a clean failover, but a partitioned primary that still has its synchronous standby on its side can keep committing, and a primary configured to fall back to asynchronous when standbys are lost can keep committing alone. Synchronous commit constrains durability, not who is allowed to be leader; only quorum-based election plus fencing constrains that.

Two air traffic controllers who lose contact with each other and both keep clearing planes onto the same runway. Neither is malfunctioning; each is doing its job with incomplete information, and the collision comes from the missing agreement, not the missing skill.

saying these in an interview costs you the question

  • Claiming better monitoring or shorter heartbeat timeouts can eliminate split-brain
  • Believing the two diverged nodes can be merged automatically when the partition heals
  • Assuming a frozen or overloaded primary is harmless because it is not really up
  • Thinking a two-node cluster is safe because each node can just check the other
  • Treating split-brain as only a downtime issue rather than acknowledged-write loss and constraint violation

context

open as a page

Why does automated database failover require a majority quorum, and what is a witness (or arbiter) node used for in a cluster that would otherwise have an even number of members?

level: middleimportance: must knowfreq 48%

basics

~20 s

Only a majority of a fixed member set may elect a primary, because at most one majority can exist, so a minority partition can never promote. A witness is a cheap vote-only member that makes the member count odd without storing data.

open as a page

What does fencing mean when a database standby is promoted, and what mechanisms are used to fence the old primary, including the STONITH (shoot the other node in the head) approach?

level: seniorimportance: must knowfreq 42%

basics

~20 s

Fencing means guaranteeing the old primary can no longer commit writes or be reached by clients before the new primary opens. It is done by killing the node (STONITH via power or hypervisor), revoking its storage or network access, or removing it from the routing layer.

open as a page

Explain lease-based leadership in an automated relational failover setup such as Patroni backed by etcd: what the leader key and its TTL are, what each agent does on every loop, and how this arrangement stops two primaries from existing.

level: seniorimportance: should knowfreq 34%

basics

~20 s

Leadership is a short-lived key in a quorum-backed store (etcd) that the primary's agent must keep renewing. If it cannot renew before the TTL expires, it demotes its own database; only after the key expires can another agent create it and promote. One key, one writer.

open as a page

You must run a highly available relational database across two data centres with automatic failover. How do you place voting members so that losing one site does not either strand the cluster read-only or allow two primaries, and what tradeoff would you accept?

level: principalimportance: should knowfreq 28%

basics

~20 s

Two sites cannot give symmetric automatic failover: whichever side holds the majority survives, the other cannot promote. Either add a third independent site or region for the tiebreaking vote, or accept an asymmetric design where one site is primary-capable and the other requires a manual, fenced promotion.

open as a page