skip to content

What happens to a SolrCloud cluster when its ZooKeeper ensemble loses quorum?

level: seniorimportance: should knowfreq 55%

answer

  1. Solr caches something locally, and it helps
  2. Reads and writes fare very differently
  3. Election needs a writable store
  4. Ephemeral registrations disappear on session loss
  5. Nothing self-heals until the majority returns

basics

~20 s

Solr keeps serving queries from its cached cluster state, but indexing stops: without a writable ZooKeeper it cannot update collection state, elect leaders or register recovering replicas. Solr treats ZooKeeper as the authority, so writes fail rather than proceeding blindly.

solid answer

~50 s

ZooKeeper is SolrCloud's source of truth: it holds each collection's `state.json` (shards, hash ranges, replicas, states, leaders), the configsets under `/configs`, the ephemeral `/live_nodes` registrations, aliases, and the leader-election and shard-term znodes. A ZooKeeper ensemble needs a majority of its members to serve writes. Lose that majority — two of three nodes gone, a partition, or disks full — and ZooKeeper stops accepting writes and drops client sessions on the minority side. For Solr this means: **reads mostly survive, writes do not.** Nodes still hold a cached view of cluster state and can answer queries against replicas that remain up. But updates need a valid leader and the ability to publish state changes, and leader election itself needs writable ZooKeeper — so indexing requests fail, replicas that go down cannot recover, and no failover happens. Nothing self-heals until quorum returns; then nodes re-register, elections settle, and recovery resumes.

code

bash · 3 lines
bash
# Point Solr at a three-node ensemble with a chroot, and upload a configset
export ZK_HOST="zk1:2181,zk2:2181,zk3:2181/solr"
bin/solr zk upconfig -n products_config -d /opt/configs/products -z "$ZK_HOST"

go deeper

for a junior

Know that SolrCloud cannot run without ZooKeeper and that ZooKeeper holds cluster metadata and configuration, not the index itself.

for a middle

Explain what lives in ZooKeeper — live_nodes, state.json, configsets, election znodes — and why write paths depend on it while an already-open searcher does not.

for a senior

Be able to triage the incident: identify that queries look healthy while indexing fails, know the session-expiry and recovery-storm behaviour, and get the ensemble back to a majority before touching Solr.

for a principal

Own the topology decision — ensemble size and placement across failure domains, dedicated disks, collection-count limits driven by state size — and the monitoring that makes quorum loss visible before users notice.

## What SolrCloud keeps in ZooKeeper ZooKeeper is not an optional add-on to SolrCloud; it is where the cluster's identity lives. - `/live_nodes` — one **ephemeral** znode per running Solr node. Ephemeral means it disappears when that node's ZooKeeper session ends, which is how the cluster detects node loss. - `/collections/<name>/state.json` — the per-collection topology: shards, their hash ranges, every replica with its node, core name, type and state (`active`, `recovering`, `down`), and which replica is leader. - `/configs/<name>` — configsets: `solrconfig.xml`, the schema, stopword and synonym files. Every replica of a collection loads the same configset from here. - `/aliases.json`, `/security.json`, the overseer work queues, and per-shard election and **shard term** znodes used by leader election. Solr nodes do not poll this; they set watches and cache what they need locally, which is why a healthy cluster can absorb short ZooKeeper hiccups. ## What quorum means A ZooKeeper ensemble commits a write only when a majority of its members acknowledge it. A three-node ensemble tolerates one failure; a five-node ensemble tolerates two. Below that threshold ZooKeeper does not degrade into a read-only-but-happy mode for its clients — the minority side has no leader, cannot commit, and expires sessions. That is deliberate: it is what stops two halves of a partitioned cluster from both believing they own the same shard. ## The effect on Solr **Queries largely keep working.** A Solr node that already knows the collection layout can route a search to replicas it believes are up and return results. Nothing about serving a query from an existing searcher requires ZooKeeper. What degrades is accuracy of the view: if a node died during the outage, others may not learn it is gone, so some shard requests fail and the aggregator may return partial results. **Indexing stops.** Update requests need a shard leader, and Solr will not accept updates it cannot safely coordinate. State transitions — publishing a replica as `recovering` or `active`, updating shard terms after a failed replica update — are ZooKeeper writes. With no writable ZooKeeper, those cannot happen, so updates error out instead of being applied on a subset of replicas. **No failover, no recovery.** Leader election is implemented with ZooKeeper znodes. If a leader is lost during a quorum outage, no replacement can be elected; the shard is simply leaderless. A replica that falls behind cannot register for recovery either. **Sessions expire.** If ZooKeeper is unreachable for longer than the configured `zkClientTimeout`, a node's session expires, its `/live_nodes` entry vanishes, and on reconnect it must re-register and re-establish its watches — often triggering a wave of recovery work exactly when the cluster is already stressed. ## Recovery Bring the ensemble back to a majority and the cluster reassembles itself: nodes recreate their ephemeral `/live_nodes` entries, elections run for any leaderless shard, replicas that missed updates compare shard terms against the new leader and either replay from their transaction log or pull a full index copy from the leader. Expect a burst of recovery traffic; it is normal, but on a large cluster it can be heavy enough to look like a second incident. ## Operating rules that follow - Run **three or five** dedicated ZooKeeper nodes across failure domains. Two is worse than one — it tolerates zero failures and doubles the chance of hitting one. - Do not co-locate ZooKeeper with heavy Solr nodes. ZooKeeper is latency- and fsync-sensitive; a GC pause or an I/O storm from a merging Solr instance turns into session expiries cluster-wide. - Never run the embedded ZooKeeper (started implicitly when you bring up Solr with `-c` and no `-z`) in production. It shares the Solr JVM's fate and gives you an ensemble of one. - Give ZooKeeper its own fast disk for the transaction log, and keep autopurge configured so snapshots do not fill the volume — a full disk is one of the most common causes of a lost ensemble. - Watch the ZooKeeper znode size limit. Very large clusters with many collections can produce big `state.json` documents; keeping collection counts sane is a real design constraint, not a detail. - Alert on ZooKeeper quorum and on Solr's `/live_nodes` count, not just on Solr's own health endpoints. Query traffic can look green while indexing has been broken for an hour.

  • Why do queries survive a ZooKeeper outage but indexing does not?
    Serving a query only needs an open searcher on a replica plus a routing decision, and every Solr node caches the cluster state it needs for that. Indexing needs a coordinating leader and the ability to record state transitions — replica states, shard terms, leader identity — all of which are ZooKeeper writes. Rather than apply an update it cannot coordinate safely, Solr fails it.
  • Why is a two-node ZooKeeper ensemble worse than a single node?
    A majority of two is still two, so a two-node ensemble tolerates zero failures — exactly like a single node — while doubling the number of machines that can fail and take quorum with them. Ensembles should have an odd count: three tolerates one failure, five tolerates two. Adding a node only increases fault tolerance when it moves the majority threshold.
  • What should you expect immediately after quorum is restored on a large cluster?
    A recovery storm. Every node re-registers in /live_nodes, leaderless shards run elections, and replicas compare shard terms with their new leader and either replay their transaction log or pull a full index copy. On a big cluster that means substantial network and disk I/O concurrently with normal query traffic, so throttle indexing while it settles and watch for replicas cycling through recovery.
  • How can a query still return wrong results during a ZooKeeper outage?
    Cluster state goes stale. If a Solr node dies while ZooKeeper is unavailable, other nodes may not learn it is gone, so the aggregator still sends shard requests to it. Depending on the shards.tolerant setting, the request either fails or returns partial results marked as such — a silently incomplete result set is the dangerous case, because totals and facet counts are then simply wrong.

saying these in an interview costs you the question

  • Says Solr keeps indexing normally without ZooKeeper
  • Thinks ZooKeeper stores the actual index data
  • Runs a two-node ensemble believing it adds fault tolerance
  • Uses the embedded ZooKeeper in production
  • Expects leader election to work while quorum is lost

context