In a volatile tier that fails over without a human, which parties can be entitled to declare the primary dead, and why is one observer never enough?
answer
- name the deciding party, not just the outcome
- peers, dedicated observers, external controller
- one vantage point sees only silence
- dead node and dead link look identical
- observers off the data hosts
basics
~20 sThree parties can own the decision: the surviving peer nodes, a set of dedicated observer processes, or a controller outside the data plane. A single observer is never enough — it cannot tell a dead node from its own severed link.
solid answer
~50 sThere are three arrangements, and a deployment has exactly one of them. The **surviving peers** decide among themselves, which means the data nodes carry the monitoring role and you need enough independent voices left after a loss for agreement to be possible at all. A set of **dedicated observer processes** sits beside the tier, deployed on hosts that do not share a failure domain with the data nodes, agrees among itself, performs the reconfiguration, and is usually also what callers ask for the current address. A **controller outside the data plane** — the managed arrangement — decides for you, typically behind one stable address it re-points, at the cost of windows you cannot see or tune. And no arrangement trusts a single voice, because one observer cannot distinguish "the node is dead" from "my link to the node is dead", and it is itself a single point of failure.
go deeper
Know that something outside the failed node decides it is gone, and that this something is a component someone deployed — not a property the store has by itself.
Be able to name the three arrangements as a set and say how each one's caller learns the new address. Avoid describing one of them as though it were how all stores work.
Give the smallest safe deployment for the arrangement you run, place the observers off the data hosts and across failure domains, and argue the single-voice problem as an evidence problem rather than as redundancy.
Decide which arrangement the organisation standardises on and what it is allowed to depend on. An external controller removes tuning but also removes visibility, and that trade should be made once, deliberately, not per team.
## Why the question has to be asked at all "The tier fails over automatically" names an outcome, not a component. Something concrete concluded that the primary was gone and performed the reconfiguration, and which something it was changes the smallest safe deployment, what happens on the isolated side of a network problem, and how a caller ends up addressing the new node. A stem that quietly assumes one of the three marks the other two wrong, which is why this is worth asking explicitly. ## The three deciding parties **1. The surviving peer nodes.** The data nodes themselves carry the monitoring role: they probe each other, and the survivors reach agreement that one of them is gone before one is promoted. Nothing extra to deploy, and the deciding party dies exactly when the data does. The sizing consequence is the one people get wrong: you need enough independent voices that agreement is still possible after losing one, so two nodes is never a safe count — neither survivor can tell being alone from being cut off from a peer that is fine. In practice this means at least three participating nodes, and deployments commonly use odd counts so that a split of the participants has an unambiguous larger side. **2. A set of dedicated observer processes.** Separate processes whose only job is to watch, agree and reconfigure. They are deployed beside the tier but not on it — an observer sharing a host with the primary dies with the primary and its vote goes missing at the worst possible moment, which also argues for spreading them across whatever your failure domains are, racks or availability zones. Because they already hold the current view, they are typically also the thing a caller asks "who is the primary now?", which collapses part of the rediscovery problem. The cost is a second thing to deploy, patch and monitor, and a deployment where the observers are healthy but the data nodes are not is its own operational state. **3. A controller outside the data plane.** In a managed offering the decision is made by the provider's control plane, which is not part of the request path at all. Usually the caller keeps one stable address that the controller re-points, so failover is nearly invisible from the caller side. The cost is that the silence threshold, the candidate-selection rule and the resulting windows are the provider's: you measure them, you do not set them. | Arrangement | Smallest safe deployment | How the caller gets the new address | What it costs you | |---|---|---|---| | Surviving peers | At least three participating nodes, spread across failure domains | The caller asks a node, or holds a map it refreshes | The data nodes carry the monitoring role and the decision dies with them | | Dedicated observers | Several observers, off the data hosts and spread across failure domains | The caller asks the observers for the current primary | A second component to deploy, patch and monitor | | External controller | Whatever the provider runs; you deploy nothing | One stable address the controller re-points | Windows and rules you cannot inspect or tune | ## Why one voice is never enough Two independent reasons, and candidates usually have only the second. 1. **A single observer cannot tell the two failures apart.** From one vantage point, "the primary is dead" and "my link to the primary is cut" produce exactly the same evidence: silence. A single severed link between an otherwise-healthy observer and an otherwise-healthy primary would, if that observer could act alone, trigger a promotion nobody needed — and the healthy primary keeps serving whatever callers can still reach it. Requiring several independent voices to agree before acting is what makes the evidence mean something, because a cut affecting one observer's path does not silence the others. 2. **A single observer is a single point of failure.** If the one thing entitled to decide is itself down when the primary fails — or was on the same host, or behind the same failed switch — nothing decides, and you have an automatic failover mechanism that produces a manual failover. This is also why observer count and data-node count are separate sizing exercises. Adding a fourth replica does not improve the decision; adding a third observer, on an independent host, does. ## What to say in an interview Name the three arrangements as a set, say which one the system you are describing uses, and give its smallest safe deployment and its address-discovery story. Then give the single-voice argument in the epistemic form — one vantage point cannot separate a dead node from a dead link — rather than only as "you need more than one for redundancy". The first answer shows you know why agreement is required; the second shows only that you know it is.
- Why is a two-node tier with peer-based deciding unsafe no matter how it is tuned?Because after losing contact, each survivor sees exactly one thing — silence — and there is no second independent voice to corroborate it. Neither can distinguish a dead peer from a cut link, and no timeout value changes that. You need a third participating voice before an automatic decision means anything.
- Where should dedicated observers run?On hosts that do not share a failure domain with the data nodes or with each other: not co-located with the primary, and spread across whatever boundary actually fails for you — hosts, racks, zones. An observer on the primary's host is a vote that always disappears exactly when you need it.
- If a managed controller handles all of this, what is left for you to decide?What you depend on. You cannot set the provider's thresholds, so your job is to measure the failover a drill actually produces, confirm that callers recover without a restart, and decide whether the workload's state can tolerate those windows. Treat the numbers as an observed property, not a configured one.
saying these in an interview costs you the question
- Describes failover as something that just happens, naming no deciding party
- Assumes the surviving peers always vote, in every deployment
- Thinks a single monitoring process is sufficient if it is reliable
- Runs the observers on the same hosts as the data nodes
- Sizes the deciding party by counting replicas rather than voices
- Believes a two-node deployment is safe with a long enough timeout