skip to content

You need to add a fourth master to a live three-master Redis Cluster and give it a fair share of the keyspace. Describe the operational procedure with redis-cli's cluster tooling and how you verify the result.

level: seniorimportance: should knowfreq 42%

answer

  1. add-node joins with zero slots
  2. reshard --cluster-from/--cluster-to/--cluster-slots
  3. rebalance needs --cluster-use-empty-masters
  4. --cluster-slave for the new master's replica
  5. check: 16384 covered, no open slots

basics

~20 s

Join the node with redis-cli --cluster add-node (it starts with zero slots), then move slots to it with --cluster reshard (or --cluster rebalance for an even split), attach a replica with --cluster add-node --cluster-slave, and verify with --cluster check and CLUSTER INFO showing cluster_state:ok and 16384 slots covered.

solid answer

~50 s

1. **Join.** `redis-cli --cluster add-node <new-host>:<port> <any-existing-node>:<port>`. The node joins the gossip mesh owning **zero slots**, so it takes no traffic yet — joining is safe and reversible. 2. **Move slots.** `redis-cli --cluster reshard <any-node>:<port> --cluster-from <id1>,<id2>,<id3> --cluster-to <new-id> --cluster-slots 4096 --cluster-yes`. Under the hood the tool runs the SETSLOT IMPORTING/MIGRATING + GETKEYSINSLOT + MIGRATE loop per slot and commits with SETSLOT NODE. `--cluster-rebalance --cluster-use-empty-masters` is the hands-off alternative that computes an even distribution itself. 3. **Add a replica** for the new master: `--cluster add-node <replica> <new-master> --cluster-slave --cluster-master-id <new-id>`; a master without a replica is a single point of data loss. 4. **Verify.** `redis-cli --cluster check <node>` must report all 16384 slots covered and **no open slots**; `CLUSTER INFO` must show `cluster_state:ok`; `--cluster info` shows keys per node. Pace it: reshard in chunks, watch p99 and `SLOWLOG`, and never leave a migration half-finished.

code

text · 16 lines
text
# 1. join (zero slots, no traffic yet)
redis-cli --cluster add-node 10.0.0.9:6379 10.0.0.1:6379

# 2. move a quarter of the keyspace, paced in chunks
redis-cli --cluster reshard 10.0.0.1:6379 \
  --cluster-from <id-a>,<id-b>,<id-c> --cluster-to <id-new> \
  --cluster-slots 1024 --cluster-pipeline 10 --cluster-yes
# ... repeat 4x, checking latency between chunks

# 3. give it a replica
redis-cli --cluster add-node 10.0.0.10:6379 10.0.0.9:6379 \
  --cluster-slave --cluster-master-id <id-new>

# 4. verify
redis-cli --cluster check 10.0.0.1:6379
redis-cli -h 10.0.0.1 cluster info | grep -E 'state|slots_assigned'

go deeper

for a junior

Know that adding a node and giving it slots are two separate steps, and name add-node plus reshard.

for a middle

Give the full command sequence including the replica, and explain that reshard drives the SETSLOT/MIGRATE loop underneath.

for a senior

Add pacing, --cluster-pipeline, large-key risk, cluster-node-timeout interaction, and the verification checklist including open slots.

for a principal

Weigh scale-out against vertical growth and key-design fixes, define the change-window policy and rollback story, and decide the availability stance behind cluster-require-full-coverage.

## What scaling out actually means here Redis Cluster does not rebalance itself. Slots are assigned explicitly and stay put until someone moves them, so "add a shard" is two independent steps: make the node a cluster member, then hand it slots. Keeping those steps separate is a feature — a joined-but-empty master serves no traffic, so you can add it well ahead of the migration window and roll back by simply never resharding to it. ## Step 1 — join ``` redis-cli --cluster add-node 10.0.0.9:6379 10.0.0.1:6379 ``` The first address is the newcomer, the second is any existing cluster member used as the entry point. The tool checks the new instance is empty and has `cluster-enabled yes`, then issues `CLUSTER MEET`. Gossip propagates the membership. `CLUSTER NODES` will now list it as a master with an empty slot range. ## Step 2 — move slots Two interfaces to the same primitive loop: **Targeted reshard**, when you know exactly what you want: ``` redis-cli --cluster reshard 10.0.0.1:6379 \ --cluster-from <id-a>,<id-b>,<id-c> \ --cluster-to <id-new> \ --cluster-slots 4096 \ --cluster-yes ``` With four masters an even split is 4096 slots each, drawn proportionally from the three incumbents. Run interactively (omit the flags) and the tool prompts for the same values, printing the slot plan before executing. **Rebalance**, when you want the tool to compute the plan: ``` redis-cli --cluster rebalance 10.0.0.1:6379 --cluster-use-empty-masters ``` `--cluster-use-empty-masters` is essential — without it the new zero-slot master is ignored, which is the single most common reason "rebalance did nothing". Useful modifiers: `--cluster-weight <id>=<w>` for heterogeneous nodes, `--cluster-threshold <pct>` for how far off balance is tolerable, `--cluster-pipeline <n>` for the number of keys per MIGRATE batch (default 10; raising it speeds the copy but lengthens each blocking call), and `--cluster-timeout` for the MIGRATE timeout. For each slot the tool performs the manual protocol: `CLUSTER SETSLOT <slot> IMPORTING` on the target, `MIGRATING` on the source, then `CLUSTER GETKEYSINSLOT` + `MIGRATE ... KEYS ...` until the slot is empty, then `CLUSTER SETSLOT <slot> NODE <target>` on the masters. Traffic keeps flowing; clients see ASK redirects for keys already moved and MOVED once ownership commits. ## Step 3 — replica for the new master ``` redis-cli --cluster add-node 10.0.0.10:6379 10.0.0.9:6379 \ --cluster-slave --cluster-master-id <id-new> ``` Until this runs, a quarter of your keyspace has no failover target. If `cluster-require-full-coverage` is `yes` (the default), losing that master takes the **whole cluster** down for writes, not just its slots — so the replica is not optional in production. ## Step 4 — verify - `redis-cli --cluster check 10.0.0.1:6379` — the important line is `[OK] All 16384 slots covered` and the absence of `[WARNING] Node ... has slots in migrating state`. Open slots mean an aborted migration; `--cluster fix` resolves them. - `redis-cli --cluster info` — keys and slot counts per node; confirms the new node actually received data. - `CLUSTER INFO` on any node — `cluster_state:ok`, `cluster_slots_assigned:16384`, `cluster_known_nodes` matching your inventory. - Application side: watch error rates for `MOVED`/`TRYAGAIN`, and confirm clients converged (a client that never refreshes topology will keep paying a double round trip; see its metrics or latency). ## Pacing and risk Resharding is a live-traffic operation and its cost is dominated by `MIGRATE`, which blocks both the source and target event loops per batch. Practical discipline: move slots in chunks rather than 4096 at once; keep `--cluster-pipeline` modest if you have large values; run during a trough; watch `SLOWLOG` and p99 latency between chunks; and know that a very large single key (a multi-gigabyte hash) can stall two masters for seconds, so find and split those before you start. Also ensure `cluster-node-timeout` is comfortably larger than your worst MIGRATE stall, otherwise a blocked master can be mistaken for a failed one and trigger a failover in the middle of your reshard. ## Scaling in The mirror image: reshard the departing node's slots away, then `redis-cli --cluster del-node <any-node> <id>`. `del-node` refuses to remove a master that still owns slots, which is a useful guard rail.

  • You ran redis-cli --cluster rebalance and it reported nothing to do, even though the new master is empty. Why?
    By default rebalance only considers masters that already own slots, so a freshly added, zero-slot master is excluded from the plan. Adding --cluster-use-empty-masters includes it. The same flag matters after a failed node replacement, where the replacement master starts empty.
  • How do you remove a master from the cluster safely?
    Reshard all of its slots to the remaining masters first — redis-cli --cluster reshard with --cluster-from the departing node's id and --cluster-to the survivors — then remove it with redis-cli --cluster del-node <entry> <id>. del-node refuses to remove a master that still owns slots, and you should remove or repoint its replica as well so the cluster does not keep an orphan.
  • Should cluster-require-full-coverage stay at its default during a scale-out?
    Its default yes means the whole cluster stops serving if any slot is uncovered, which is a strong consistency-of-availability stance and is unaffected by a healthy reshard, since slots are always owned by someone. It matters if a master dies with no replica mid-operation: with yes the entire cluster goes read-error, with no only the affected slots do. Changing it is a product decision about partial availability, not something to flip just to get through a migration.

saying these in an interview costs you the question

  • Believing Redis Cluster rebalances slots automatically when a node joins
  • Expecting the new node to take traffic immediately after add-node
  • Running rebalance without --cluster-use-empty-masters and concluding the tool is broken
  • Leaving the new master without a replica
  • Resharding the entire keyspace in one unpaced run during peak traffic

context