In an in-memory store replicated as one primary and two replicas, why is promotion gated on a majority of the deciding members?
answer
- two writable primaries, one keyspace
- a partition, not a crash
- who is permitted to promote
- more than half of one membership
- an even split promotes nobody
basics
~20 sA majority gate guarantees that only one side of a network partition can promote, because two separated sides cannot both hold more than half of one membership. Without it, both halves promote, both accept writes, and the two keyspaces diverge.
solid answer
~50 sSplit brain is the state where two nodes each accept writes as primary for the same keyspace. It is caused by a network partition rather than a crash: a crashed primary writes nothing, but a primary that is merely unreachable keeps serving the callers that can still reach it while the other side promotes a replica. The gate is that a side may promote only if it can reach more than half of the **deciding members**, and since two disjoint sides cannot both hold more than half of one membership, at most one side is ever permitted to promote. The arithmetic behind a strict majority is consensus material and you borrow the conclusion; what belongs to the store is *which members count* — the store nodes themselves, a set of dedicated observer processes, or a controller outside the data plane — because that fixes the smallest deployment that is actually safe.
go deeper
Recall that promotion is not automatic and not free: a replica is a lagging copy, and something has to decide it may take over. Be able to say that if two nodes both take over, writes go to two places.
Explain the gate mechanically. More than half of the deciding members, two separated sides cannot both have that, so at most one may promote. Then say who the deciding members are in the deployment you are describing.
Show that you know what the gate does not cover: the old primary keeps writing, the promoted replica may be missing acknowledged writes, and the smaller side is deliberately unavailable. Name the third failure domain as the structural fix for a two-site layout.
Treat it as a placement decision rather than a setting. Decide where deciding members live, which site is permitted to survive alone, and what the organisation will do during an outage instead of lowering the threshold, since that reaction is how most real split brains are created.
## The failure this rule exists to prevent A replicated in-memory store puts a copy of the **whole keyspace** on each node. One node is the **primary** and accepts writes; the others are **replicas** and receive them. When the primary stops answering, something promotes a replica so writes can continue. **Split brain** is the state where that promotion happens while the old primary is still alive and still accepting writes. Two nodes each believe they own the same keyspace, and callers that can reach either one are told their writes succeeded. It is worth being precise about the cause, because candidates usually name the wrong one: a crash does not produce split brain, because a crashed node writes nothing. A **network partition** does — the link between two groups of machines fails while every machine keeps running, each group sees the other as dead, and the group holding replicas has every reason to promote one. ## What the majority gate actually says The rule is simple to state: **a side may promote only if it can reach more than half of the deciding members.** Two disjoint sides cannot both hold more than half of one membership, so at most one side of any split is ever permitted to promote. That is the whole of the guarantee. The arithmetic of why a strict majority is the right quorum, rather than some fixed count, is consensus material. In an interview about a store you cite it as a conclusion and move on. What is specific to this tier is **which members are being counted**, because that varies by product and by deployment, and each arrangement makes a different deployment unsafe. | Who the deciding members are | Smallest shape that survives losing one | What the isolated side sees | |---|---|---| | The store nodes themselves | Three nodes in three independent failure domains | Its peers unreachable; it must act on that alone | | Dedicated observer processes | Three observers, placed apart from the store nodes | Observers unreachable; the store node is not itself a voter | | A controller outside the data plane | Depends on the controller's own availability, not on node count | Whatever the controller does to it, which may include cutting it off at the network | A candidate who names all three, and says which one their deployment uses, is demonstrating the thing this question is for. A candidate who assumes the peers vote has silently ruled out the other two arrangements. ## The deployment shapes this rule condemns - **Two nodes.** One primary, one replica, nothing else. Neither side of a split holds more than half of two, so nobody may promote. Teams then either lose writes entirely or disable the gate and manufacture the split brain it existed to prevent. - **An even number of deciding members.** Split evenly, no side promotes; and an even count buys no additional failure tolerance over the odd count below it, so it is cost without benefit. - **Every deciding member inside the two sites you are protecting against each other.** Whichever site holds the larger share is the only site that can ever promote alone. That may be a deliberate choice, but it should be a choice, not a discovery made during an incident. - **Observers co-located with the store nodes they watch.** Losing the machine takes the watcher with it, so the membership shrinks exactly when you need it. The repair for the first three is the same: put one deciding member in a **third failure domain**. It holds no keyspace and serves no traffic; it exists so that a two-site deployment has an odd membership and a tiebreaker that survives either site. ## What the gate does not do This is where most answers overreach. The gate governs **who may be promoted**, and nothing more. 1. It does not stop the old primary. An isolated primary is still running, still reachable by some callers, and still accepting writes. Stopping it is a separate obligation that falls on that node itself, or on whatever can still reach it. 2. It does not preserve the writes already accepted. Under acknowledge-then-propagate the promoted replica is missing whatever had not yet reached it, and on this tier those writes are simply gone rather than replayable from a durable record. Some stores in this class instead withhold acknowledgment until at least one copy holds the write, and some let a caller choose that per write, which changes the size of the gap but not the existence of the gate. 3. It does not keep the minority side available. By design, the smaller side gives up writes. That is the trade being bought. ## The trap worth naming out loud The most common way teams get a real split brain is by reacting to the *availability* failure. The link drops, no side has a majority, nothing promotes, and someone lowers the required threshold or promotes by hand on the smaller side to restore service. Both are the same mistake: they permit two primaries. The honest responses are to add a deciding member in a third place, to accept that one site cannot survive alone, or to promote manually only after confirming the other side is genuinely stopped rather than merely unreachable.
- The deciding members are split evenly across two data halls and the link between them fails. Which hall promotes?Neither. No side reaches more than half, so nothing is promoted and the tier is unavailable for writes on the side that lost its primary. That is an availability failure, not a split brain, and it is the outcome the gate is supposed to produce. The fix is structural: an odd membership with one deciding member in a third failure domain, so one hall can always form a majority without the other.
- Does the majority gate stop the isolated old primary from continuing to accept writes?No. It only decides who may be promoted. The old primary is running, some callers can still reach it, and nothing on the majority side can tell it to stop, since by definition it cannot be reached from there. Ending the writes is a separate obligation carried by that node itself, or by a controller outside the data plane that still has a path to it.
- Does a majority gate exist in every product of this class?No, and that is the point of asking which members decide. Some deployments run peer nodes that vote, some run dedicated observer processes with their own membership, and some managed offerings promote from a control plane outside the data plane where the node count is not the relevant quantity at all. Treat the gate as a property you verify for the deployment you actually have, not as a guarantee the class provides.
Two duty officers share one watch and one rulebook. The rule is that you may take command only if you can raise more than half the watch on the radio. When the radio splits the watch into two groups, at most one officer can raise a majority, so at most one takes command. The other has to put the logbook down rather than keep writing entries nobody will ever read — the rule does not make the second officer stop on its own, it only decides which one is allowed to start.
saying these in an interview costs you the question
- Believes both sides can promote safely and reconcile the keyspaces afterwards
- Thinks the majority gate also stops the isolated old primary accepting writes
- Counts the store nodes as voters when dedicated observer processes decide
- Places every deciding member inside the two sites being protected from each other
- Lowers the required threshold during an outage to get promotions working again
- Assumes an even number of deciding members is safer than the odd number below it