How would you lay out an Elasticsearch cluster across three availability zones so losing one zone keeps it writable?
answer
- Three separate things must survive, not one
- Spread the votes before spreading the data
- Nodes need a label before placement rules mean anything
- Survivors must not silently absorb the missing share
- Two thirds of the fleet carries the whole load
basics
~20 sPut one master-eligible node in each zone so a majority survives losing one, spread data nodes evenly, tag every node with a zone attribute, enable allocation awareness on that attribute so shard copies land in different zones, and keep at least one replica per index.
solid answer
~40 sThree things must survive the loss of a zone: the master quorum, at least one copy of every shard, and enough capacity to serve the traffic. Place **one master-eligible node per zone** — a majority of three is two, so one zone can go without losing the master. Spread data nodes evenly across the zones, set `node.attr.zone` on each, and enable `cluster.routing.allocation.awareness.attributes: zone` so Elasticsearch spreads the copies of each shard across distinct zone values. With `number_of_replicas: 1` you survive a zone loss; with `2` you survive a zone loss *and* a node failure. Add **forced awareness** (`cluster.routing.allocation.awareness.force.zone.values`) so that when a zone disappears the surviving nodes do not try to absorb all its copies and blow their disk. Then size for the loss: two thirds of the fleet must serve the whole load.
code
yaml · 6 lines# on every node in zone a
node.attr.zone: eu-west-1a
# cluster-wide
cluster.routing.allocation.awareness.attributes: zone
cluster.routing.allocation.awareness.force.zone.values: eu-west-1a,eu-west-1b,eu-west-1cgo deeper
Recall that Elasticsearch can be told which zone each node is in, and that it then keeps a shard's primary and replica in different zones.
Explain the node attribute plus awareness settings, and why one master-eligible node per zone across three zones lets a majority survive losing one.
Design and defend the whole layout: master placement, replica count, forced awareness, and what happens operationally when a zone drops and later returns.
Own the tradeoffs end to end — capacity headroom versus cost, one replica versus two per data tier, the two-zone tiebreaker decision, and the boundary where zone redundancy stops helping and snapshots or a second cluster start.
## Name the failure you are designing for "Survives a zone loss" is three separate guarantees, and a design that only delivers one of them fails in production: 1. **Coordination survives** — a master can still be elected, so shards can be reallocated and mappings can change. 2. **Data survives** — every shard still has an assigned copy, so the cluster is not red. 3. **Capacity survives** — the surviving zones can carry 100% of the traffic that three zones were carrying. Most designs get the first two and quietly fail the third. ## Master-eligible placement Election requires a strict majority of the voting configuration. Three master-eligible nodes, one per zone, gives a majority of two, so any single zone can vanish and the remaining two still elect a master. Five (2/2/1) also works and tolerates two node failures, at the cost of a larger state-publication group. What does not work is concentrating masters: two master-eligible nodes in zone A and one in zone B means losing zone A leaves one vote out of three and a cluster that cannot elect anything. This is also the argument that decides the **two-zone** question, which comes up constantly because two data centres are what many organisations actually have. Two zones cannot host a majority that survives either one going away — whichever zone holds two of three masters becomes a single point of failure. The standard answer is a third location holding one small node with `node.roles: [ master, voting_only ]`: it votes, it can never be elected, it holds no data, and it costs almost nothing. Without it, honest two-zone designs accept manual intervention on failover. ## Data node placement and awareness Tag each node with an attribute, conventionally `node.attr.zone: eu-west-1a`, and turn on awareness: ``` cluster.routing.allocation.awareness.attributes: zone ``` The allocator then treats zone values as failure domains and will not place a primary and its replica in the same one. It also balances shard counts per zone rather than only per node. Keep the number of data nodes per zone equal — an uneven layout produces uneven per-node shard counts, since awareness balances within each zone independently. ## Choosing the replica count With three zones and awareness on, `number_of_replicas: 1` means two copies of every shard in two different zones. Lose a zone and every shard still has a copy: not red, but temporarily without redundancy, and the surviving copy of some shards is now a single point of failure until the cluster rebuilds. `number_of_replicas: 2` puts a copy in each zone: after a zone loss you still have two copies, and you can lose a node as well. The cost is a third more disk and a third more indexing work on every write. That tradeoff — a whole extra copy of the dataset against surviving a compound failure — is exactly the judgement call a lead is expected to make explicitly rather than by default, and it is legitimately different for a billing index and for a 7-day log index that can be re-ingested. ## Forced awareness Plain awareness only spreads copies across the zones that currently exist. If a zone disappears, the allocator happily rebuilds its copies in the survivors — which is often precisely what you do not want, because you sized those nodes for two thirds of the data and they will fill up, cross the high watermark, and start refusing allocations. Forced awareness fixes this: ``` cluster.routing.allocation.awareness.force.zone.values: eu-west-1a,eu-west-1b,eu-west-1c ``` Now Elasticsearch knows the full set of expected zone values and simply leaves the missing zone's copies unassigned until that zone returns. The cluster runs yellow, deliberately, with stable disk usage — and recovers by rebuilding into the returning zone rather than shuffling twice. The tradeoff is that a zone that is gone for good needs an operator to remove it from the forced list. ## Capacity, not just copies If three zones each run at 70% CPU, losing one leaves two zones asked to do 150% of what they were doing. Designing for a zone loss means running each zone at roughly two thirds of its usable capacity, or accepting degraded latency during the event and saying so in the SLO. The same applies to heap and disk: the watermark maths must hold with one zone's worth of nodes missing, which is another reason forced awareness — which stops the rebuild that would consume that headroom — is part of the design rather than an optimisation. ## The recovery burst When the zone comes back, all of its shard copies recover at once. Recovery is throttled per node (`indices.recovery.max_bytes_per_sec`, and the concurrent-recovery limits), and those defaults are conservative on modern hardware: the choice is between a long yellow period and a recovery that competes with live traffic. Decide it in advance and write it into the runbook, because it is not a decision anyone makes well during an incident. ## What this does not cover Zone redundancy protects against infrastructure failure in one location. It does nothing about a bad mapping change, a delete-by-query run against the wrong index, or a corrupt bulk load — those replicate faithfully to every zone. Snapshots to object storage remain the backstop, and a region-level disaster needs a separate cluster with cross-cluster replication rather than more zones.
- Why does a two-data-centre Elasticsearch deployment usually need a node in a third location?A majority of master-eligible nodes cannot be split evenly across two locations. Whichever site holds two of three votes becomes a single point of failure, and a 2/2 split can never form a majority at all. A small master plus voting-only node in a third location holds no data, costs little, and lets either main site survive alone.
- What does forced awareness add over plain allocation awareness in Elasticsearch?Plain awareness spreads copies across the zones that exist right now, so a missing zone's copies get rebuilt in the survivors — filling disks that were sized for their own share. Forced awareness declares the full set of expected zone values, so the missing zone's copies stay unassigned and the cluster runs deliberately yellow with stable capacity until that zone returns.
- With three zones, when is number_of_replicas: 2 worth the extra copy of the data?When you must survive a zone loss and a node failure together, or when a zone-loss event must not leave any shard with a single copy. It costs a third more disk and a third more indexing work per write. For a critical transactional index that is usually worth it; for high-volume logs that can be re-ingested, one replica plus snapshots is the better trade.
saying these in an interview costs you the question
- Puts two of three master-eligible nodes in one zone
- Assumes plain awareness alone prevents zone overload after a failure
- Sizes zones so survivors cannot carry the full load
- Treats zone redundancy as a substitute for snapshots
- Believes a two-zone cluster can fail over automatically without a tiebreaker