A Redis Cluster starts rejecting commands with CLUSTERDOWN and CLUSTER INFO reports cluster_state:fail. Explain what hash-slot coverage means, how you identify which slots have no owner, and what the cluster-require-full-coverage setting changes.
answer
- 16384 slots, each owned by exactly one reachable master
- cluster_slots_assigned < 16384 = holes
- default require-full-coverage yes = refuse everything
- no = serve covered slots, cache-friendly
- redis-cli --cluster check / fix
basics
~20 sCoverage means all 16384 slots are claimed by a reachable master. If any slot has no owner, the cluster marks itself failed and, with cluster-require-full-coverage yes (the default), refuses every command - even for slots that are fine. Set it to no to keep serving the covered slots.
solid answer
~50 sA cluster is healthy only when the 16384 slots are each owned by exactly one reachable master. `CLUSTER INFO` shows this: `cluster_slots_assigned` should be 16384, `cluster_slots_ok` should match, and `cluster_slots_pfail`/`cluster_slots_fail` should be 0. When slots are uncovered - a master down with no replica to promote, an aborted resharding that left slots unassigned, or a node reset - `cluster_state` becomes `fail`. With `cluster-require-full-coverage yes` (the default) a node in that state rejects *all* commands with `CLUSTERDOWN The cluster is down`, on the reasoning that partial data silently serving stale or missing results is worse than an outright error. Setting it to `no` lets each master keep serving the slots it owns, and only requests for uncovered slots fail. Diagnose with `CLUSTER NODES` on several nodes (look for `fail` flags and missing ranges) and `redis-cli --cluster check` / `--cluster fix`.
code
text · 13 lines> CLUSTER INFO
cluster_state:fail
cluster_slots_assigned:10923 # holes: 5461 slots unowned
cluster_slots_ok:10923
cluster_slots_fail:0
$ redis-cli --cluster check 127.0.0.1:7000
[ERR] Not all 16384 slots are covered by nodes.
# cache-style tradeoff: keep serving the slots that are healthy
> CONFIG SET cluster-require-full-coverage no
OK
# (persist it in redis.conf on every node too)go deeper
Know that all 16384 slots must be owned, and that CLUSTERDOWN means some are not.
Read CLUSTER INFO fields correctly and explain that the default configuration refuses all commands, not just those for missing slots.
Show the triage sequence, distinguish a genuine hole from a minority view, and argue the require-full-coverage trade for cache versus authoritative data.
Frame it as an availability policy decision per data class, tie it to replica placement and migration-barrier settings, and require alerting on cluster_state and slot counters rather than on client errors.
## What coverage means Redis Cluster partitions the keyspace into 16384 slots. The cluster is *fully covered* when every slot number 0..16383 is claimed by exactly one master that is reachable and considered healthy by the majority. Coverage is a property of the slot-to-node map, not of the data: a slot can be covered and empty, or uncovered while its data still exists on a node that the cluster considers failed. There are three ways coverage breaks: 1. **A master fails with no replica available to promote.** Its slots become unowned from the cluster's point of view. 2. **An aborted or half-finished resharding.** Slots left in `importing`/`migrating` state, or explicitly unassigned by a failed operation, leave holes. 3. **Administrative accidents** - `CLUSTER RESET` on a live node, a node re-added with an empty config, deleting a node that still owned slots. ## Reading the state `CLUSTER INFO` is the fastest probe: - `cluster_state:ok|fail` - the node's verdict on whether the cluster can serve. - `cluster_slots_assigned` - how many of the 16384 have any owner at all. Below 16384 means holes. - `cluster_slots_ok` - assigned slots whose owner is believed healthy. - `cluster_slots_pfail` / `cluster_slots_fail` - slots whose owner is suspected (pfail) or agreed (fail) to be down. - `cluster_known_nodes` and `cluster_size` (masters serving at least one slot) help spot a node that vanished. `CLUSTER NODES` shows the per-node detail: flags including `fail?`/`fail`, and the slot ranges each node claims. Summing the claimed ranges of healthy masters tells you exactly which numbers are missing. `redis-cli --cluster check <host:port>` does this arithmetic for you, printing coverage and flagging slots that are unassigned or claimed by more than one node; `--cluster fix` will attempt to close the holes and finish stuck migrations. Always ask more than one node. Each node reports its own gossiped belief, so during a partition the minority side may declare `fail` while the majority side is happily serving. ## What cluster-require-full-coverage does The default is `yes`, and it is a deliberately conservative availability stance: **if any slot is uncovered, every master refuses every command** with `CLUSTERDOWN The cluster is down` - including commands for slots that are perfectly healthy. The rationale is correctness in the common case where Redis backs a system that assumes the whole keyspace is present: an application that silently gets "key not found" for a shard's worth of data may take wrong actions (re-registering users, re-charging, recomputing and caching wrong values), whereas an explicit error usually triggers a clean fallback or a page. Setting `cluster-require-full-coverage no` flips the trade: each master continues to serve the slots it owns, and only keys hashing into uncovered slots fail (with `CLUSTERDOWN Hash slot not served`). That is the right choice when Redis is a pure cache whose misses are cheap and safe - losing 1/3 of a cache is far better than losing 100% of it. It is the wrong choice when Redis holds authoritative state, or when a partial view can produce silent data corruption downstream. The setting is applied per node, so change it everywhere; it can be set live with `CONFIG SET cluster-require-full-coverage no` and must also be persisted in the config file to survive a restart. A related knob is `cluster-allow-reads-when-down`, which permits reads even while the node considers the cluster down, again trading consistency for availability. ## The response, in order 1. Confirm scope: `CLUSTER INFO` on several nodes; is it a genuine hole or one node's minority view? 2. Identify the missing ranges: `CLUSTER NODES` / `redis-cli --cluster check`. 3. Find the cause. Down master with no replica? Promote or restart the node so its slots return. Stuck migration? Finish or roll it back with `--cluster fix` (or `CLUSTER SETSLOT ... NODE`) rather than hand-assigning slots blindly. 4. Only assign orphaned slots to a fresh node once you accept that the data in them is gone - `CLUSTER ADDSLOTS` on an empty node restores availability at the cost of that slice of data. 5. Afterwards, fix the structural cause: every master should have at least one replica, `cluster-migration-barrier` should allow spare replicas to move to a master that lost its own, and alerting should watch `cluster_state` and `cluster_slots_ok` directly rather than only client error rates. ## Nuance worth stating Coverage is about ownership, not about node count: a single master can legally own all 16384 slots, and a cluster with dozens of nodes is still `fail` if one range is unclaimed. Conversely, `cluster_state:ok` says nothing about whether every master has a replica - a fully covered cluster with no redundancy is one crash away from being uncovered.
- When is cluster-require-full-coverage no a good default, and when is it dangerous?It is good when Redis is a pure cache: a miss is cheap, so serving two thirds of the keyspace beats serving none. It is dangerous when Redis holds authoritative state or when the application cannot distinguish 'absent because the shard is down' from 'absent because it was never written' - that ambiguity leads to duplicate work, wrong writes, or cache poisoning with values computed from incomplete data.
- The uncovered slots belong to a master whose disk is gone. How do you get the cluster back to ok?Either restore that shard - bring up a replacement node and let it take over the slots from a replica or a backup - or accept the data loss and assign the orphaned slots to a live node with CLUSTER ADDSLOTS (redis-cli --cluster fix automates the common cases). The second path restores availability immediately but permanently discards the keys that hashed into those slots.
saying these in an interview costs you the question
- Thinking cluster_state:fail means the whole cluster process is down rather than that some slots are unowned
- Assuming healthy shards keep serving by default when slots are missing
- Blindly running --cluster fix or ADDSLOTS without realising it can declare data lost
- Trusting a single node's CLUSTER NODES output during a partition
- Believing cluster_state:ok implies every master has a replica