skip to content

Walk through how a single hash slot is moved from one Redis Cluster master to another while the cluster keeps serving traffic. Which commands are involved and what does each node do during the window?

level: seniorimportance: must knowfreq 48%

answer

  1. IMPORTING on target first, then MIGRATING on source
  2. GETKEYSINSLOT + MIGRATE ... KEYS in batches
  3. MIGRATE = DUMP + RESTORE + DEL, atomic, blocking
  4. source serves or ASKs; target needs ASKING
  5. SETSLOT ... NODE commits, bumps config epoch

basics

~20 s

Mark the slot IMPORTING on the target and MIGRATING on the source with CLUSTER SETSLOT, then loop CLUSTER GETKEYSINSLOT plus MIGRATE to copy keys in batches while the source answers ASK for already-moved keys, and finish with CLUSTER SETSLOT <slot> NODE <target> on every master.

solid answer

~60 s

Four phases. 1. **Mark.** On the target: `CLUSTER SETSLOT <slot> IMPORTING <source-id>`. On the source: `CLUSTER SETSLOT <slot> MIGRATING <target-id>`. Order matters — target first, so no key can be transferred into a node that is not expecting it. 2. **Copy.** Loop `CLUSTER GETKEYSINSLOT <slot> <count>` on the source and pass the batch to `MIGRATE <host> <port> "" <db> <timeout> KEYS k1 k2 ...`. MIGRATE is atomic per key: it serialises the value, RESTOREs it on the target, and deletes it locally only after an OK. 3. **Serve during the window.** The source still owns the slot: keys it still has are served normally, missing keys get `-ASK <slot> <target>`, and multi-key commands split across both nodes get `-TRYAGAIN`. The target only serves the slot to connections that sent `ASKING`. 4. **Commit.** `CLUSTER SETSLOT <slot> NODE <target-id>`, issued to both nodes and ideally all masters. The target bumps its config epoch, gossip spreads ownership, and the source starts replying MOVED. In practice `redis-cli --cluster reshard` does all of this; knowing the primitives is what lets you repair a half-finished migration.

code

text · 22 lines
text
# ids
A=07c37dfe...  # source
B=e7d1eecc...  # target

# 1. mark (target first)
redis-cli -p 7002 cluster setslot 866 importing $A
redis-cli -p 7000 cluster setslot 866 migrating $B

# 2. copy in batches until empty
while true; do
  keys=$(redis-cli -p 7000 cluster getkeysinslot 866 100)
  [ -z "$keys" ] && break
  redis-cli -p 7000 migrate 127.0.0.1 7002 "" 0 5000 KEYS $keys
done

# 3. commit on every master
for p in 7000 7001 7002; do
  redis-cli -p $p cluster setslot 866 node $B
done

# verify
redis-cli -p 7000 cluster countkeysinslot 866   # -> 0

go deeper

for a junior

Recall that slots are moved key by key with MIGRATE while the source keeps serving and redirects with ASK.

for a middle

Name the four phases and the exact commands, and explain why the target is marked IMPORTING first.

for a senior

Add the blocking nature of MIGRATE, batch sizing as a pacing knob, TRYAGAIN semantics, and how to detect and repair an open slot.

for a principal

Discuss migration as an availability event: the latency budget it consumes, large-key risk, pacing policy, and whether the topology should change at all versus provisioning differently.

## The constraint Redis Cluster has no data movement layer running in the background and no two-phase commit. A slot must change owner while both nodes keep serving, without ever letting two nodes believe they own the same slot for the same key. The design solves this by making the **source authoritative for the whole window** and having it hand out per-key redirects. ## Phase 1 — mark the endpoints ``` # on target CLUSTER SETSLOT <slot> IMPORTING <source-node-id> # on source CLUSTER SETSLOT <slot> MIGRATING <target-node-id> ``` These flags are **local to the two nodes** and are not gossiped: no other node's view of ownership changes. `IMPORTING` tells the target "you may accept keys of this slot from the source, and you may serve them to clients that sent ASKING". `MIGRATING` tells the source "if a key of this slot is missing, do not reply nil — reply ASK". Setting IMPORTING first is deliberate. If you flagged the source first and started migrating, the target could receive a RESTORE for a slot it does not own and reject it. ## Phase 2 — copy the keys ``` CLUSTER COUNTKEYSINSLOT <slot> CLUSTER GETKEYSINSLOT <slot> 100 MIGRATE <target-host> <target-port> "" 0 5000 KEYS key1 key2 ... key100 ``` `MIGRATE` is the workhorse. For each key it runs, on the source, the equivalent of `DUMP`, sends a `RESTORE` to the target over an internal connection, waits for `+OK`, and only then `DEL`s the local copy. It is **atomic per key and synchronously blocking**: both instances are blocked for the duration of the transfer of that batch. That is the operationally important fact — migrating a single 500 MB key or a batch containing one blocks the event loop of *two* masters for as long as serialisation plus transfer takes, which is how resharding produces latency spikes. Options: `COPY` keeps the source copy, `REPLACE` overwrites an existing target key, `AUTH`/`AUTH2` supply credentials, and the empty key argument `""` plus `KEYS ...` is the multi-key form. The loop repeats until `CLUSTER COUNTKEYSINSLOT` returns 0. Batch size is the pacing knob: smaller batches mean more round trips but shorter blocking per batch. ## Phase 3 — behaviour of live traffic during the window On the **source** (owner, MIGRATING): - key present → serve normally; - key absent → `-ASK <slot> <target>`; - multi-key command, all keys present → serve; - multi-key command, some keys already migrated → `-TRYAGAIN`, because atomicity cannot be honoured. On the **target** (IMPORTING): - plain command for that slot → `-MOVED <slot> <source>` (it does not own the slot); - `ASKING` then the command → served. Writes are not lost: a write to a key already migrated goes through ASK to the target; a write to a key not yet migrated lands on the source and is copied later by a subsequent MIGRATE batch. A key created on the source *after* its batch has run is still in the slot and will be picked up by the next `GETKEYSINSLOT` sweep, which is why the loop keeps going until the count is genuinely zero. ## Phase 4 — commit ownership ``` CLUSTER SETSLOT <slot> NODE <target-node-id> ``` Issued on the target it clears IMPORTING and claims the slot, bumping the node's **config epoch** so its claim wins in gossip. Issued on the source it clears MIGRATING and, since the source no longer holds keys of the slot, it starts replying MOVED. Redis's own tooling sends SETSLOT to **all masters** so that every node's view converges immediately rather than waiting for gossip; on 7.x this is the documented practice and reduces the window of stale MOVEDs. If a node is unreachable at commit time, gossip still converges it once it returns. ## Failure and repair If the operator or the tool dies mid-migration the slot is left flagged. `redis-cli --cluster check` reports it as open; `redis-cli --cluster fix` completes or rolls it back, and `CLUSTER SETSLOT <slot> STABLE` clears a stray flag manually. A cluster left with open slots will keep emitting ASK/TRYAGAIN and will refuse to pass a consistency check — a common cause of "random TRYAGAIN errors" weeks after a resize. ## Why you should know the primitives Day to day you run `redis-cli --cluster reshard` or `--cluster rebalance`, which implement exactly this loop. The primitives matter when a migration aborts, when you need to move a specific hot slot rather than a range, or when you need to explain a latency spike: the answer is nearly always a large key blocking both masters inside MIGRATE.

  • Why can MIGRATE cause a latency spike, and what would you do about a slot containing one huge key?
    MIGRATE serialises the value, ships it, waits for the target's RESTORE and then deletes locally — all synchronously, blocking the event loop on both the source and the target for the duration. A multi-gigabyte hash or list can therefore stall two masters for seconds. Mitigations are to migrate during a low-traffic window, use small batch sizes so no batch bundles the big key with others, raise the MIGRATE timeout so it does not fail halfway, and structurally to split the big key (shard the hash under a hash tag) before resharding.
  • A resharding job was killed halfway. What state is the cluster in and how do you recover?
    The slot is left flagged MIGRATING on the source and IMPORTING on the target, so clients keep getting ASK for keys that moved and TRYAGAIN for split multi-key commands, and redis-cli --cluster check reports an open slot. Recovery is redis-cli --cluster fix, which finishes the transfer and issues SETSLOT NODE, or manual cleanup: finish the key copy, commit with CLUSTER SETSLOT <slot> NODE <target-id> on all masters, or abandon it with CLUSTER SETSLOT <slot> STABLE on both nodes.
  • Are writes lost if a client writes to a key that has just been migrated?
    No. Once the key is gone from the source, the source replies ASK, so the write is routed to the target after an ASKING and applied there. If the write instead lands on a key that has not migrated yet, it is applied on the source and picked up by a later GETKEYSINSLOT/MIGRATE batch, since the loop runs until COUNTKEYSINSLOT is zero.

saying these in an interview costs you the question

  • Claiming the cluster blocks writes to the slot during migration
  • Saying MIGRATE copies keys asynchronously in the background
  • Setting MIGRATING on the source before IMPORTING on the target
  • Forgetting the final CLUSTER SETSLOT ... NODE, leaving the slot permanently open
  • Believing gossip alone moves the data, with no explicit key transfer

context