Three copies of a quorum-based store all reported available, the drain honoured its cap, and the store still lost quorum - how?
answer
- the cap trusts one signal
- available to serve, not caught up
- membership is the store's own fact
- three members, majority of two
- report unavailable while catching up
basics
~20 sThe cap counts copies that report themselves available, and a copy can report available before it has rejoined the store's own membership. Two of three were counted, only one was really voting, so removing the third left no majority.
solid answer
~50 sA disruption cap has exactly one input: how many copies report themselves available to serve. It knows nothing about what those copies are to each other. A store that accepts writes only while a majority of its members agree has a second, invisible notion of membership, and a copy that restarted minutes ago can be answering its readiness signal while still catching up and not yet counted by its peers. The cap then sees three available, grants a removal because two would remain at the floor, and the store is left with one real member out of three. The arithmetic was right and the input was wrong. The fix is a truthful signal - report unavailable until caught up and counted - plus sizing the cap from the majority rather than from the copy count.
go deeper
The takeaway is that a platform only knows what a copy tells it. If a copy says it is ready while it is still catching up, every decision built on that answer - including which copies may be taken down - is made on a false number.
Explain the two separate notions of membership: the platform's availability count and the store's own voting set. Then walk the arithmetic, showing that a cap of one-down on three members is only safe when all three counted copies are really members.
Show you would fix the input rather than the cap: make the readiness signal reflect membership, size the budget from the majority, and wait on the signal between removals instead of on a fixed sleep.
The general rule is that any safety budget is only as good as the signal it counts. Where loss of quorum is expensive, buy margin with more members rather than a stricter cap, since a stricter cap protects nothing against the host that fails mid-drain.
## What the cap counts A disruption cap compares one number against a floor: how many copies of the workload the platform believes are **available**. That belief comes from a per-copy signal - the copy answering that it is ready to serve. The cap has no other input. It does not read the workload's internal state, it does not know what the copies are to one another, and it has no idea that these three are members of a store which accepts writes only while a majority of them agree. That is the whole gap. **Available to serve** and **counted as a member by its own peers** are two different facts, and the cap only ever sees the first. ## The sequence, with numbers Three copies, A, B and C, one per host. The store needs a majority - **2 of 3** - to accept writes. The cap says at most one copy may be voluntarily down, so the floor is 2. 1. **14:02** - B's host had an unrelated problem and B restarted. It comes back up and begins replaying its log to catch up with the others. 2. **14:05** - B answers its readiness signal affirmatively: the process is up and its port is open. The platform now counts 3 available. The store, meanwhile, still has only A and C as voting members, because B has not finished catching up. 3. **14:06** - the drain of C's host begins. The cap checks: 3 available, removing one leaves 2, which is at the floor. **Granted.** 4. **14:06** - C stops. The platform sees 2 available and is satisfied. The store sees one voting member of three, loses its majority and stops accepting writes. No rule was breached and no arithmetic was wrong. The input was wrong. ## What the platform believed against what was true | | Platform's view | Reality | |---|---|---| | A | available | voting member | | B | available | up, still replaying, not yet a member | | C | available, safe to remove | the second of only two voting members | | Members left after the removal | 2 | 1 | ## Making the signal honest The fix is not a larger cap; it is a truthful input. A copy of a membership-based store should report itself **unavailable** until it is caught up and counted by its own cluster, not merely until its process is listening. Then the platform would have counted 2 available at 14:06, seen that removing C leaves 1 - below the floor of 2 - and refused the removal until B had rejoined. The drain would have waited a few minutes and completed safely. Three practices follow: - **Derive the signal from membership, not from the process being up.** A copy that is running but not yet counted by its peers must say so, and saying so costs it nothing but a slower drain. - **Size the cap from the majority, not from the copy count.** Three members tolerate one loss and five tolerate two. That number, not a round percentage, is the cap. - **Wait for the signal to recover between removals** rather than for a fixed sleep. A fixed pause is only ever accidentally long enough, and it stops being long enough the day the store gets bigger. ## What the cap still cannot do Even with an honest signal, the cap governs only the removals the platform is asked to perform. If a host fails on its own thirty seconds after a granted removal, the store can lose its majority with nobody having broken a rule: the budget was spent legitimately, and then reality spent it again. For a store whose loss of quorum is expensive, that argues for **more members rather than a stricter cap** - five members leave a margin three do not, and that margin is what absorbs the failure which arrives while you are draining. It is also why the interesting question in a review is not 'what is the cap' but 'what does the availability signal actually mean for this workload'. Two workloads with identical caps can be safe and unsafe respectively, entirely on the strength of what their copies say about themselves.
- How should the cap be sized for a five-member quorum store?From the majority rather than the member count: five members tolerate two losses, so at most two may be down. The budget is shared with whatever is already gone involuntarily, so if a host failure has taken one member, an honest availability count leaves room for one more removal, not two.
- Does draining one host at a time make this safe?Not on its own. One host at a time bounds how many copies you take, but says nothing about whether the copies left behind are real members. The wait between removals is what helps, and only if the thing you wait for is the signal meaning 'caught up' rather than 'process started'.
saying these in an interview costs you the question
- If the platform says a copy is available, its cluster has it as a member
- A cap of one-down is always safe for a three-member store
- Quorum is just a count of running processes
- A restarted member is caught up as soon as it accepts connections
- Losing quorum here is the drain's fault, not the signal's