A five-server Consul datacenter loses three of its servers, so there is no leader. What continues to work for service discovery, what stops, and how do the `stale` and `consistent` read modes change the answer?
answer
- majority gone means no commits
- gossip does not need a quorum
- freshness is a per-request choice
- the lookups keep working, which is the trap
- alert on the leader, not on lookup errors
basics
~20 sWithout a quorum there is no leader, so all writes fail: registrations, deregistrations and check-state updates. Reads depend on mode — stale reads are answered by any surviving server from its own replicated state, so DNS keeps resolving, while default and consistent reads fail.
solid answer
~50 sLosing three of five servers loses the majority Raft needs, so no leader can be elected. Every write path stops: new registrations, deregistrations, and the anti-entropy syncs that carry check-state changes into the catalog. What happens to reads is a per-request choice. A **`consistent`** read is verified with the leader before answering, so it fails outright. The **`default`** mode is answered by the leader without that verification, so it also fails with no leader. A **`stale`** read is answered by *any* server out of its own replicated copy, so it still succeeds — with data frozen at whatever that server last replicated. This is why discovery degrades gracefully rather than dying: Consul's DNS allows stale reads by default, so lookups keep returning the last known healthy set. The danger is precisely that: the answers look normal while being blind to every change since the outage began, so an instance that died during the outage keeps resolving.
code
bash · 8 lines# Survives leader loss: any server answers from its own replica
curl -s 'http://127.0.0.1:8500/v1/health/service/web?passing&stale'
# Fails without a leader: freshness is verified with the leader first
curl -s 'http://127.0.0.1:8500/v1/health/service/web?passing&consistent'
# The signal that actually matters during a server-tier incident
curl -s 'http://127.0.0.1:8500/v1/status/leader'go deeper
Know that Consul's servers need a majority to accept writes, and that losing it means new registrations cannot be recorded.
Explain the three read modes and which one each interface uses by default, including why the DNS path stays available when the leader is gone.
Show the incident judgment: lookups keep succeeding while the catalog silently freezes, so instances that died during the outage still resolve — and the alerting has to be on leader and peer health.
Own the trade the whole platform inherits: a consistent catalog with deliberately stale discovery reads, what maximum staleness you accept, and a rehearsed outage-recovery runbook instead of improvisation.
## What quorum loss actually breaks Consul's servers replicate the catalog with Raft. Raft commits only with a majority of the peer set, so five servers need three; losing three leaves two, which cannot elect a leader and cannot commit. The cluster is not *down* — the surviving servers are running and hold a full replica of the state as of the last successful replication — but it has no write path. **Stops immediately:** - New service registrations reaching the catalog. - Deregistrations. - Check-state transitions being recorded centrally, because anti-entropy syncs are catalog writes. - Any operation requiring leader-verified freshness. **Keeps working:** - Client agents keep running their local health checks and keep serving their local API for local state. - Gossip continues; membership and failure detection are a separate protocol from Raft and need no quorum. - Reads that are allowed to be stale. ## The three consistency modes Consul lets the caller choose freshness per request, and the choice is exactly a latency-versus-staleness trade. **`consistent`** — the server checks with the leader that it is still the leader before answering. Strongest guarantee, one extra round trip, and it is the mode that fails first when leadership is in doubt. Use it when a stale answer would be a correctness bug. **`default`** — the leader answers from its own state without re-verifying leadership. Almost always correct; the theoretical window is an old leader that has been deposed but does not know it yet, bounded by the leader-lease timeout. This is the everyday mode. **`stale`** — *any* server answers from its local replica. No leader involvement at all, so it survives leader loss, spreads read load across every server instead of concentrating it on the leader, and returns data that may lag. Consul tracks how far behind a stale answer is and can refuse to serve one that is too old, governed by a maximum-staleness setting. ``` curl 'http://127.0.0.1:8500/v1/health/service/web?passing&stale' curl 'http://127.0.0.1:8500/v1/health/service/web?passing&consistent' ``` ## Why DNS keeps answering Consul's DNS interface allows stale reads by default (`dns_config.allow_stale`, on by default in modern versions, with `max_stale` bounding how far behind an answer may be). That is a deliberate availability choice for the discovery path: an application resolving a service name during a server-tier incident gets the last known good answer rather than SERVFAIL. If you need DNS to reflect strictly current state, you turn `allow_stale` off — and accept that DNS then fails whenever the leader is unavailable. ## The failure mode this creates The outage is quiet, and that is the problem. Callers keep resolving names and keep connecting successfully, so nothing looks wrong from the application's side. Meanwhile: - An instance that crashed *during* the outage still resolves, because its check-state change never reached the catalog. Callers connect to it and fail. - An instance that came up during the outage never resolves, because its registration never landed, so a scale-up or a deploy silently adds no capacity. - The catalog's picture is frozen; the longer the outage, the more wrong it gets. So the operational signal you need is not "are lookups failing" — they are not — but leader health itself: alert on the absence of a leader and on Raft peer count, not on discovery error rates. ## Recovering Bring servers back and the surviving replicas resume normal Raft operation as soon as a majority exists again; the returning peers catch up from the leader. When the lost servers are gone for good and you cannot reach a majority — the classic "three of five nodes destroyed" case — Consul provides an outage-recovery procedure that rewrites the peer set on the survivors so a smaller cluster can elect a leader again. That is a last resort with real risk of losing recent writes, and it should be a documented, rehearsed runbook rather than something improvised during the incident. Autopilot's dead-server cleanup handles the ordinary case of replacing failed servers one at a time, which is why you never want to be doing manual recovery in the first place. ## The design lesson Consul chooses consistency for the catalog: it would rather refuse a write than accept a conflicting one. Discovery *reads* are then deliberately allowed to be stale so that the availability cost of that choice does not fall on every application in the fleet. Knowing which of your calls sits on which side of that line is the whole answer.
- If lookups keep succeeding during a quorum outage, what should you actually alert on?Leader presence and Raft peer count, not discovery error rates. Because DNS serves stale reads, the application-visible symptom of a leaderless cluster is nothing at all until a registration or a health transition is missed. Alerting on an empty leader endpoint, on peers below the quorum threshold, and on anti-entropy sync failures at the agents catches it while the catalog is still nearly current.
- When is it worth paying for a `consistent` read rather than accepting the default mode?When acting on a stale answer would be a correctness bug rather than a performance annoyance — leader-election-adjacent logic, or a decision that must not be made twice on two different views. It costs an extra round trip to the leader per request and concentrates load there, so it is wrong as a blanket default for discovery, where the data is inherently a moment behind reality anyway.
- Why does gossip-based failure detection keep working when Raft has lost its quorum?They are independent protocols. Gossip is a peer-to-peer membership and failure-detection protocol with no leader and no consensus requirement, so agents keep learning which nodes are alive. What breaks is recording those facts in the catalog, since that is a Raft write. The cluster still *knows* a node died; it just cannot durably publish it.
saying these in an interview costs you the question
- Thinking discovery goes fully dark when the leader is lost
- Believing gossip needs a Raft quorum to work
- Treating stale reads as identical to default reads
- Assuming successful lookups mean the catalog is current
- Improvising manual peer recovery during an incident