You are laying out a Redis Cluster across three availability zones. Which placement and replication decisions determine whether the cluster keeps serving — and keeps its writes — when an entire zone disappears?
answer
- Majority of slot-owning primaries → no zone may hold ≥ half
- Two zones cannot fail over automatically
- Replica never in its primary's zone
- Spare replicas + migration barrier re-cover orphaned primaries
- cluster-require-full-coverage no = partial service
basics
~20 sSpread primaries so no zone holds half or more of them, since failover needs a majority of primaries; never place a replica in its primary's zone; keep spare replicas so a promoted node is not left bare; and decide explicitly whether uncovered slots should take the whole cluster down.
solid answer
~60 sThree decisions dominate. **1. Primary quorum placement.** Both FAIL agreement and promotion require a majority of slot-owning primaries. With three primaries, one per zone, losing a zone leaves two of three — a majority, so failover proceeds. Two zones cannot work: losing the zone with most primaries leaves a minority that can neither promote nor accept writes. Never let one zone hold half or more of the primaries. **2. Replica placement and count.** A replica must not share a zone with its primary, or the zone loss takes both. One replica per primary survives a zone loss, but afterwards each promoted primary is bare; `cluster-migration-barrier` lets a primary with spare replicas donate one, so a few extra replicas make the cluster self-heal. Cross-zone replication latency widens the async loss window. **3. Availability policy.** `cluster-require-full-coverage no` keeps the surviving slots serving when a shard cannot recover; `yes` fails the whole cluster. Pair that with `cluster-node-timeout` tuned to cross-zone jitter, `min-replicas-to-write` per data class, and client libraries that refresh topology and retry promptly.
code
text · 17 lines# does every primary still have a replica, and where do nodes live?
$ redis-cli --cluster check 10.0.1.11:6379
M: <id> 10.0.1.11:6379 slots:0-5460 (5461 slots) 1 replicas
S: <id> 10.0.2.21:6379 replicates <id-of-10.0.1.11>
...
[WARNING] Node 10.0.3.13:6379 has no replica
# key settings (identical on every node)
cluster-node-timeout 5000
cluster-require-full-coverage no # keep serving covered slots
cluster-migration-barrier 1 # donate a spare replica to an orphan
cluster-replica-validity-factor 10 # 0 = always allow promotion
# during a zone loss
> CLUSTER INFO
cluster_state:ok
cluster_slots_ok:16384go deeper
Focus on the two placement rules: at least three zones with primaries spread out, and a replica never in the same zone as its primary.
Add why the spread matters — failover needs a majority of primaries — and that after a failover the promoted node may have no replica left.
Bring in cluster-node-timeout against cross-zone latency and stalls, replica migration with the migration barrier, cluster-require-full-coverage, and client topology refresh.
Present it as a stated availability and durability contract per data class — quorum placement, memory headroom for the failure state, acceptable loss window, partial-service policy — and insist it be verified by a rehearsed zone-loss exercise rather than assumed.
## The governing constraint: majority of primaries Everything in Redis Cluster failover — escalating PFAIL to FAIL, and authorising a replica's promotion — needs agreement from more than half of the **primaries that own hash slots**. Replicas do not vote. That single sentence dictates the topology. - **Three zones, primaries spread evenly.** Three primaries (one per zone) or six (two per zone) both leave a majority after any single zone fails: 2 of 3, or 4 of 6. Note that 4 of 6 is a majority, so an even primary count across three zones still works — the danger is not evenness, it is concentration. - **Two zones cannot self-heal.** Whichever zone holds the majority is a single point of failure for the entire cluster's ability to fail over; lose it and the survivors are a minority that can neither promote nor accept writes. If only two zones exist, be honest that automatic failover across a zone loss is not achievable, and plan an operator-driven recovery (`CLUSTER FAILOVER TAKEOVER`, accepting the data-loss risk) rather than pretending otherwise. - **Never concentrate.** A layout of four primaries as 2/1/1 survives losing either small zone but not the big one (2 of 4 is not a majority). Audit placement as a property of the deployment, not a hope about the scheduler. ## Replica placement A replica in the same zone as its primary provides no protection against zone loss — the correlated failure takes both, and the shard's slots go uncovered. Enforce anti-affinity at the orchestration layer, and verify it after every rebalance: `redis-cli --cluster check` reports whether replicas are co-located with their primaries, and cluster creation tooling can be told to spread by host, but zone-awareness must generally be enforced by whatever schedules the pods or instances. Redis itself has no notion of zones. ## How many replicas One replica per primary is the minimum for automatic failover. After a zone loss, however, every promoted primary in the surviving zones is running **without** a replica, so a second failure in that window is unrecoverable and the cluster is fragile exactly when it is under stress. Options: - **Two replicas per primary** — expensive (each replica holds the full shard dataset in memory) but leaves coverage intact after one zone loss. - **Replica migration** — a primary with more replicas than `cluster-migration-barrier` (default 1) automatically donates one to an orphaned primary. A handful of spare replicas placed to survive a zone loss lets the cluster re-cover itself without operator action. This is why an "extra" replica somewhere is better than none. Also budget memory for the failure state: after a zone loss the surviving zones host the promoted primaries as well as their own, so their instances must have the headroom to be primaries (client connections, output buffers, replication backlogs) rather than quiet replicas. ## Timeouts and cross-zone latency `cluster-node-timeout` must sit comfortably above the worst realistic cross-zone round trip *and* above ordinary process stalls — fork for snapshotting, page-cache pressure, swap — or you get spurious failovers. But every extra second is also a second an isolated primary may accept writes that will be discarded. Measure the sources of pause and pick a value from data; the 15 s default is conservative and often fine, and 5 s is defensible on a low-latency inter-zone network with disciplined snapshotting. Keep the value identical on every node. Cross-zone replication also widens the asynchronous loss window: a replica two milliseconds away is two milliseconds of writes behind at best, more under load. If a data class cannot tolerate that, add `min-replicas-to-write` for that shard and `WAIT` on critical writes, and accept the availability cost. ## Partial availability policy `cluster-require-full-coverage` is the explicit choice between consistency of the cluster's advertised capability and partial service: - `yes` (default): if any hash slot has no serving primary, the cluster reports itself down and refuses queries even for healthy slots. Predictable and blunt. - `no`: keep serving the covered slots and error only on requests for the uncovered ones. For a cache tier, `no` is nearly always right — losing 1/3 of a cache is a degraded hit ratio, not an outage. For a store whose partial answers would be wrong, `yes` may be the honest setting. Decide it deliberately. ## Client-side The cluster can be perfect and the outage still long if clients do not refresh topology. Require: a client that refreshes slot maps on redirection and on connection errors, bounded socket and command timeouts consistent with the node timeout, retry policy that does not amplify load during the failover, and idempotent write shapes so a retry after an ambiguous timeout is safe. Whether the application reads from replicas is a separate decision with its own staleness consequences. ## What to test Rehearse it: kill a whole zone in a game day, measure time to `cluster_state:ok`, count lost writes with a write-and-verify harness, confirm replica migration re-covered the orphaned primaries, and check that memory headroom in the survivors was sufficient. A topology that has never been zone-failed is a hypothesis, not a design.
- Why is a two-zone Redis Cluster fundamentally unable to survive a zone loss automatically?Failover requires PFAIL to be escalated to FAIL and a promotion to be authorised, both by a majority of slot-owning primaries. Split over two zones, one zone necessarily holds at least half the primaries; when that zone is lost, the survivors are not a majority and can neither declare FAIL nor grant votes. The remaining primaries also stop accepting writes once they cannot reach a majority, so recovery requires an operator forcing a takeover and accepting the risk.
- After a zone loss, several promoted primaries have no replica. What in Redis addresses that, and what are its limits?Replica migration: a primary that has more replicas than cluster-migration-barrier automatically hands one to an orphaned primary, restoring coverage without operator action. Its limits are that donors must actually have spare replicas — with exactly one replica per primary there is nothing to donate — and that the migrated replica must perform a synchronisation with its new primary, which costs network and memory at a bad moment. The design implication is to provision extra replicas placed to survive the zone loss.
- What does setting cluster-require-full-coverage to no actually change?By default, if any of the 16384 hash slots has no serving primary, the entire cluster reports itself down and refuses queries even for slots that are perfectly healthy. Setting it to no keeps the covered slots serving and returns errors only for requests that map to uncovered ones. For a cache tier that is usually the right posture, since a fraction of misses is far better than a total outage; for a store where partial answers would be silently wrong, the default is safer.
saying these in an interview costs you the question
- Spreading replicas across zones but concentrating half or more of the primaries in one zone
- Believing Redis is zone-aware and will place replicas sensibly by itself
- Running one replica per primary and assuming the cluster is still redundant immediately after a failover
- Leaving cluster-require-full-coverage at its default without deciding whether partial service is preferable
- Sizing surviving instances only for their steady-state role and not for hosting promoted primaries