In a replicated in-memory store cut in two by a network partition, why must the minority side refuse writes or demote itself unprompted?
answer
- no message is coming
- isolation and mass failure look identical
- act against yourself on a timer
- refuse writes below healthy copies
- narrows the window, never closes it
basics
~20 sBecause the majority side cannot reach it to tell it anything. Safety depends on the isolated primary acting against itself on a local timer, refusing writes or demoting, since a promotion is very likely already happening on the other side.
solid answer
~50 sThe majority gate only decides who may be promoted. It says nothing about the node that was already primary and is now isolated, and that node is the one still accepting writes. The awkward part is that it cannot tell the two cases apart: from inside, *I am cut off* and *everyone else died* look identical, and in the second case continuing to serve is correct. So the only safe rule is one it applies to itself — after going longer than some interval without confirmation from enough current copies or enough deciding members, it stops accepting writes and steps down. The usual mechanism is a **minimum-healthy-copies write gate**: refuse writes when fewer than some number of copies are current within a lag bound. It **narrows** the loss window and does not close it, because the writes accepted in the seconds before it trips are already unpropagated.
go deeper
Recall that a node cut off from the network keeps running and keeps answering whoever can still reach it. Being unreachable is not the same as being stopped.
Explain why the instruction cannot arrive: the failure is precisely the inability to communicate, so the isolated side has to act on a local rule. Name refusing writes and stepping down as the two things it can do.
Demonstrate the judgment: pick a gate setting, defend it against ordinary propagation-lag spikes, and state clearly that it narrows rather than closes the window. Tie the choice to whether the keys are derived or the only copy.
Decide the standing posture and who may change it. Whether the organisation tolerates silently discarded acknowledged writes at all, for which keyspaces, and whether you require an arrangement where something outside the data plane can stop a node that will not stop itself.
## Why the obligation runs inwards During a network partition the two sides are symmetric in exactly the way that matters: neither can reach the other. The majority side may promote a replica, and that is what the majority gate governs. But the node that was already primary is on the other side, still running, still holding connections from the callers that share its half of the network, and still answering writes with yes. Nothing on the majority side can stop it. That is not an implementation gap, it is the definition of the partition — if the majority side could reach the old primary to demote it, there would be no partition. **So the only party able to end the old primary's writes is the old primary itself**, or something that still has a path to it, which in practice means a controller outside the data plane rather than a peer. ## What the isolated node can and cannot know From inside the isolated primary, the observable fact is silence: no replica is acknowledging, no observer is checking in. That single observation is produced by at least three different situations. - Every other node crashed, and this node is the last one standing. Continuing to serve is correct. - The link failed, and the other side is fine and is about to promote. Continuing to serve is a split brain. - The other side is slow rather than gone. Continuing to serve is correct and stepping down is an unnecessary outage. No local observation distinguishes them, and waiting longer does not resolve the ambiguity, it only makes the window wider. Since the consequence of guessing wrong in the second case is divergence that will be thrown away, the safe default is to assume the expensive case and act against yourself. ## The mechanism The concrete form is a **minimum-healthy-copies write gate**: the primary refuses new writes unless at least some number of copies are known to be current, where current means their **propagation lag** is inside a bound. Isolation makes that condition false, so the gate trips on its own, without anybody sending it a message. Two things must be said about it honestly, because both are routinely overstated. 1. **It narrows the loss window, it does not close it.** The gate can only fire after the primary has noticed the condition. Every write accepted between the moment the link dropped and the moment the gate trips has already been acknowledged and is already unpropagated. Those writes are gone when the other side promotes. 2. **Whether you have one at all varies.** Some stores in this class offer such a gate; some offer no self-demotion behaviour whatsoever, and an isolated primary will keep serving writes for as long as callers keep arriving. Some managed offerings resolve it from outside instead, by having the control plane cut the node off at the network, which is a stronger guarantee precisely because it does not depend on the node's own cooperation. ## The cost of tuning it | Setting | What you get | What it costs | |---|---|---| | Gate tight — several copies, small lag bound | Writes stop quickly after isolation; the smallest divergence | Ordinary propagation-lag spikes stop writes when nothing is actually wrong | | Gate loose — one copy, generous lag bound | Writes survive routine lag and single-replica maintenance | A wider band of writes is accepted and then discarded | | No gate at all | Maximum availability on the isolated side | The isolated primary serves writes for the whole partition, and all of them are lost | The tuning question is therefore not *how do I avoid downtime*, it is *which is worse for this keyspace: a refused write or an accepted write that will be silently thrown away?* And the answer depends on something the mechanism cannot see. ## Why the state decides the setting If the affected keys hold **derived state** — something another system can reproduce — an acknowledged write that later vanishes costs recomputation and a slow period. A refusal, by contrast, is an immediate error the caller must handle. A loose gate is defensible. If the keys hold **sole-copy ephemeral state** — a claim, a deduplication record, a quota counter, a session — an acknowledged write that vanishes is an incident: the deduplication record disappears and the operation runs twice, the quota resets and a customer is overcharged, the claim vanishes and two workers hold what should be one. Here a refused write is strictly better, because the caller at least knows. Tighten the gate, and be prepared to argue it against the availability number. ## What a strong answer sounds like Say the obligation runs inwards and why: no message is coming, because the inability to receive one is the failure. Say the isolated node cannot distinguish isolation from mass failure, so it must assume the worse case. Name the gate, then immediately say that it narrows rather than closes the window, and that whether the product offers it at all is something you check rather than assume.
- Why can a minimum-healthy-copies write gate not close the loss window completely?Because it can only fire after the primary has noticed that too few copies are current, and noticing takes at least one detection interval. Every write accepted between the link dropping and the gate tripping was already acknowledged and never propagated. The gate bounds that band, it does not remove it; only refusing to acknowledge until a copy holds the write attacks the window itself, and that trade is a different one.
- Is there an arrangement in which the minority side does not have to police itself?Yes, where a controller outside the data plane still has a path to the isolated node. Because it is not on either side of the split in the same sense, it can cut the node off at the network or stop the process, and the guarantee no longer depends on the node's cooperation or its clock. That is one reason the identity of the deciding party matters so much in this tier.
- How tight should the gate be for a keyspace holding deduplication records?Tight. A deduplication record is usually sole-copy: if it vanishes, the operation it was suppressing runs a second time, and the caller was told the record was stored. A refused write surfaces as an error the caller can retry or fail on, which is far cheaper than a silent double execution. Derived data justifies the opposite setting, which is why the gate is a per-keyspace argument rather than a default.
saying these in an interview costs you the question
- Assumes the majority side can message the isolated primary to stand down
- Thinks an isolated primary can tell isolation apart from every peer crashing
- Treats a minimum-healthy-copies write gate as closing the loss window
- Sets the gate so tight that ordinary propagation lag stops writes
- Believes every store in this class offers a self-demotion behaviour