How does Elasticsearch's voting configuration decide whether a master can be elected after nodes leave?
answer
- A named subset of nodes whose votes count
- More than half, never merely half
- Two of them is the worst possible number
- One setting exists only for the very first start
- It shrinks itself, but only down to a floor
basics
~20 sElasticsearch maintains a voting configuration: the set of master-eligible nodes whose votes count. An election needs a strict majority of that set, so the cluster survives losing fewer than half of it. Below that it has no master and rejects cluster-state changes.
solid answer
~50 sElasticsearch keeps an explicit **voting configuration** — the subset of master-eligible nodes whose votes are counted — inside the cluster state, and an election requires a strict majority of it. The cluster manages that set itself; the old `discovery.zen.minimum_master_nodes` setting was removed in 7.0 precisely because operators kept setting it wrong and creating split brains. As master-eligible nodes leave, `cluster.auto_shrink_voting_configuration` (default true) removes them from the set, but never shrinks it below three master-eligible nodes, so a three-node master quorum stays at two votes rather than collapsing to one. `cluster.initial_master_nodes` is used **only** to bootstrap a brand-new cluster the first time it starts, and must be removed afterwards — leaving it in place risks bootstrapping a second, separate cluster. With no elected master, searches on already-allocated shards keep working, but nothing that changes cluster state does: index creation, mapping updates, and shard allocation all fail.
code
bash · 1 lineGET _cluster/state?filter_path=metadata.cluster_coordination.last_committed_configgo deeper
Recall that Elasticsearch elects one master from the master-eligible nodes and that a majority is required, which is why three is the usual count.
Explain the strict-majority arithmetic for three, five and two nodes, and that Elasticsearch manages the voting configuration itself rather than through a hand-set quorum setting.
Diagnose a masterless cluster: what still serves traffic, what fails, and which configuration errors — bootstrap settings left in place, even node counts, unreachable transport ports — produce it.
Own the topology decision: where master-eligible nodes sit relative to failure domains, when a voting-only tiebreaker is warranted, and the procedure for safely shrinking a master quorum without an outage.
## The voting configuration Every Elasticsearch cluster stores, as part of its cluster state, a *voting configuration*: the list of master-eligible nodes whose votes count towards electing a master and towards committing cluster-state updates. It is usually the same as the set of master-eligible nodes, but not always — the cluster adjusts it deliberately as membership changes. You can read it directly: ``` GET _cluster/state?filter_path=metadata.cluster_coordination.last_committed_config ``` ## Why a strict majority A master is elected, and a cluster-state update is committed, only with votes from more than half of the voting configuration. Requiring a strict majority means two disjoint groups of nodes can never both assemble one, which is what stops two masters from existing on either side of a network partition and diverging the cluster state. The arithmetic that follows is the part interviewers probe: a voting configuration of three tolerates one failure, five tolerates two, and **two tolerates none** — a majority of two is two, so losing either node leaves the cluster without a master. That makes a two-master-eligible-node cluster strictly worse than a one-node cluster for availability, and it is why odd sizes are the rule. ## Automatic management, and why minimum_master_nodes is gone Before 7.0 you configured a quorum by hand with `discovery.zen.minimum_master_nodes`, and every time the cluster grew or shrank you had to update it. Setting it too low permitted split brain; forgetting to update it caused unnecessary outages. Elasticsearch 7.0 replaced that whole subsystem: the cluster now maintains the voting configuration itself, and the setting no longer exists. `cluster.auto_shrink_voting_configuration` (default `true`) controls what happens when a master-eligible node leaves. With the default, the departed node is removed from the voting configuration automatically — as long as at least three master-eligible nodes remain. That floor matters: it stops a five-node master quorum from shrinking all the way down to one node during a rolling failure, which would make a single remaining node authoritative. Below three, the configuration is left alone and the operator must intervene by using the voting configuration exclusions API (`POST /_cluster/voting_config_exclusions`) when *deliberately* shrinking a cluster — and must clear those exclusions afterwards. ## Bootstrapping A brand-new cluster has no cluster state and therefore no voting configuration, so it cannot elect anything. `cluster.initial_master_nodes` breaks that circularity by naming the master-eligible nodes that form the first voting configuration. Three properties matter operationally: 1. It applies **only** to the very first startup of a brand-new cluster. Once bootstrapped, it is ignored on that node's data path. 2. It must be identical on every node you list, and must name nodes by their `node.name` (or address) exactly. 3. It must be **removed** from the configuration afterwards. If a node later starts with an empty data directory and this setting still present, it can bootstrap a *second* cluster of its own rather than joining the existing one — the classic accidental split. ## Voting-only nodes A node configured with `node.roles: [ master, voting_only ]` participates in elections and in cluster-state commits but can never be elected master itself. It is the standard tiebreaker: a small, cheap machine that gives an otherwise even-sized master set an odd vote count, which is exactly what a two-data-centre deployment needs in a third location. ## What breaks with no master When a majority is unreachable, the cluster has no elected master. What still works: searches and gets against shards that are already allocated and started, because they need no cluster-state change. What stops: creating or deleting indices, mapping and settings updates, shard allocation and rebalancing, and anything that publishes state. Client requests that need the master fail with a `master_not_discovered_exception`, and `GET _cluster/health` itself may hang or time out. Indexing into existing shards may continue briefly but degrades as soon as anything needs a cluster-state update. ## Diagnosis and operational rules When a cluster loses its master, check the master-eligible nodes' logs for repeated election attempts and for the coordination subsystem's messages naming which nodes it can and cannot discover — they explicitly report the nodes it is waiting on. Verify that all master-eligible nodes agree on `cluster.name` and can reach each other on the transport port, and that `discovery.seed_hosts` lists the master-eligible nodes. The rules worth stating in an interview: keep an odd number of master-eligible nodes, three for most clusters and five for very large ones; never run exactly two; remove `cluster.initial_master_nodes` after bootstrap; when permanently removing a master-eligible node from a small cluster, use the voting configuration exclusions API rather than just switching the machine off.
- Why is a two-master-eligible-node Elasticsearch cluster worse than a one-node cluster for availability?A majority of two is two, so both nodes must be up for an election or a cluster-state commit to succeed. Losing either one leaves the cluster without a master. A single-node cluster is its own majority and keeps working. Two is the worst possible count — add a third master-eligible node, or a voting-only tiebreaker.
- What still works in an Elasticsearch cluster that has lost its elected master?Searches and gets against shards that are already allocated and started continue, since they need no cluster-state change. Everything that publishes state stops: index creation and deletion, mapping and settings updates, shard allocation and rebalancing. Requests needing the master fail with master_not_discovered_exception, and cluster health may not answer at all.
- What is a voting-only master-eligible node used for?Declared as node.roles: [ master, voting_only ], it votes in elections and in cluster-state commits but can never be elected master. It is the cheap tiebreaker that gives an otherwise even master set an odd vote count — typically a small instance placed in a third zone so a two-zone deployment can survive losing one.
The voting configuration is the committee roll, not the room. A vote passes only with more than half of the names on the roll, so no breakaway group in a separate room can ever pass one too.
saying these in an interview costs you the question
- Still recommends setting discovery.zen.minimum_master_nodes
- Says two master-eligible nodes tolerate one failure
- Leaves cluster.initial_master_nodes in the config permanently
- Thinks losing the master makes all searches fail immediately
- Believes more master-eligible nodes always raises fault tolerance