skip to content

Resharding & Redirects

You will learn how slots migrate between nodes while the cluster stays live, and the MOVED vs ASK redirects clients must follow mid-migration. Interviewers use the MOVED/ASK distinction as the canonical 'have you actually operated Redis Cluster' probe.

part ofRedisoverview, primer and where to startread it →
on this pageshow

questions

5

In Redis Cluster, a client can receive either a -MOVED or an -ASK error reply. What does each one mean, and how must a correct client react differently to them?

level: middleimportance: must knowfreq 62%

answer

  1. MOVED = permanent, update map
  2. ASK = one key, one shot, don't update map
  3. ASKING precedes the redirected command
  4. MIGRATING source / IMPORTING target
  5. partial multi-key during migration → TRYAGAIN

basics

~20 s

MOVED means the slot permanently belongs to another node: follow it and update the cached slot map. ASK means only this key has already moved during an in-progress migration: send ASKING then the command to the target once, and do not update the map.

solid answer

~60 s

Both replies redirect the client, but they differ in **permanence**. `-MOVED <slot> <host>:<port>` says the whole slot is now owned by that node. The client should follow the redirect **and refresh its slot map**, because every future key in that slot goes there too. MOVED is what you see after a resharding completes or after a failover promotes a replica. `-ASK <slot> <host>:<port>` is emitted only while a slot is being migrated: the source node still owns the slot but this particular key has already been transferred. It is a **one-shot** redirect for that command only. The client must send `ASKING` to the target first (the target rejects non-ASKING traffic for a slot it is only importing, replying MOVED back to the source), then the command, and must **not** update its map — the slot's owner has not changed yet. Getting this wrong is a real bug: caching an ASK target flips the map mid-migration and causes redirect storms; ignoring `ASKING` bounces the client back and forth until the redirect cap trips.

code

text · 17 lines
text
# source node, slot 866 is MIGRATING, key already transferred
> GET foo
(error) ASK 866 127.0.0.1:7002

# target node, slot 866 is only IMPORTING -> plain command is bounced back
> GET foo
(error) MOVED 866 127.0.0.1:7000

# correct sequence on the target
> ASKING
OK
> GET foo
"bar"

# after CLUSTER SETSLOT 866 NODE <target-id>
> GET foo            # sent to the old source
(error) MOVED 866 127.0.0.1:7002

go deeper

for a junior

Know the one-line distinction: MOVED = permanent, refresh your map; ASK = temporary, one key, retry there once.

for a middle

Explain the MIGRATING/IMPORTING pair, why ASKING is required, and why caching an ASK target is a bug.

for a senior

Add TRYAGAIN for partial multi-key commands, redirect caps, and the latency cost of redirects during a large reshard.

for a principal

Discuss it as the cost of clientside routing without a proxy: correctness pushed into every client library, so client maturity is a real availability dependency during topology change.

## Why two different redirects exist Redis Cluster owns 16384 hash slots, each served by exactly one master. Clients cache a slot-to-node map so that, in steady state, a command costs one round trip. Two situations invalidate that cache, and they have different lifetimes — hence two error replies. ## MOVED: the ownership changed, permanently ``` -MOVED 15495 10.0.0.7:6379 ``` The node is saying: *slot 15495 is not mine; it belongs to 10.0.0.7:6379, and it will keep belonging to it.* Ownership changes are propagated between nodes by gossip and arbitrated by config epochs, so by the time you see MOVED the cluster has already agreed. Causes: a resharding finished (`CLUSTER SETSLOT <slot> NODE <target>`), or a replica was promoted and inherited its master's slots. Correct client behaviour: retry the command against the named node **and invalidate/refresh the slot map**. Most clients update just the one slot immediately and schedule a full `CLUSTER SHARDS` refresh, because one MOVED usually means a whole range moved. A client that follows MOVED but never updates its map still works — every command simply costs two round trips forever, which is a silent 50% throughput loss and a classic production mystery. ## ASK: this key moved, the slot has not During a live migration the operator marks the slot `MIGRATING` on the source and `IMPORTING` on the target. The slot still belongs to the source; keys are copied over in batches while traffic continues. So the source may hold some keys of the slot and the target the rest. The source therefore answers per key: - key exists locally → serve it normally; - key does not exist locally and the slot is MIGRATING → reply `-ASK <slot> <target>`, meaning "it may already be over there, go look, just this once"; - multi-key command with some keys present and some not → reply `-TRYAGAIN`, because it cannot serve the command atomically; the client backs off and retries. The target refuses ordinary commands for a slot it is only IMPORTING — it would reply MOVED and bounce the client back to the source, producing an infinite loop. The escape hatch is the `ASKING` command: it sets a one-command flag on that connection meaning "I was sent here by an ASK redirect, serve this even though the slot is only importing". So the client must send `ASKING` immediately before the redirected command; pipelining them together is the normal implementation. In a `MULTI` transaction, `ASKING` goes before the `MULTI` or is queued as the first command, depending on the client. Crucially, **ASK must not update the slot map**. The slot's real owner is still the source; if the client rewrites its map to the target, every subsequent key in that slot goes to the importing node, gets MOVED back, and the client flaps. Clients that treat ASK as MOVED show up as redirect storms and `Too many redirects` errors during resharding — a genuine, frequently-seen bug. ## The end of a migration When the last key has been transferred, the operator issues `CLUSTER SETSLOT <slot> NODE <target-id>`, ideally to every master (redis-cli's resharding tool does this). The target bumps its config epoch and claims the slot; gossip spreads the new ownership. From that instant the source answers MOVED instead of ASK, and clients converge onto the new map. The window in which ASK is possible is exactly the migration window. ## Operational and client-side implications - Redirects cost an extra round trip, so a large resharding under heavy load raises p99 latency even when nothing is broken. Pacing the migration (fewer slots at a time) keeps the ASK/MOVED rate bounded. - Clients cap redirect chains (commonly 5, configurable) and then surface an error; a badly stale map plus an in-flight migration can exhaust that cap, so many clients also react to redirects by triggering an *adaptive* topology refresh. - `redis-cli -c` implements both redirects, including `ASKING`, which makes it a good way to sanity-check behaviour during a migration. - If you see ASK replies when no migration is running, the cluster is in an inconsistent state — `redis-cli --cluster check` will report slots still marked migrating/importing, usually because a resharding aborted midway; `CLUSTER SETSLOT <slot> STABLE` clears the flags.

  • What happens to an MSET whose keys are all in one slot, while that slot is being migrated and only half the keys have been copied?
    The source node cannot serve it atomically — some keys are local, some are already on the target — so it replies -TRYAGAIN instead of ASK. A correct client waits a short interval and retries; after the migration finishes the command works again, either on the source or, once ownership flips, on the target after a MOVED. TRYAGAIN is therefore a transient, migration-only error and should be retried with backoff rather than surfaced to the user.
  • A client library treats ASK exactly like MOVED. What symptom would you expect in production?
    Redirect storms during any resharding. The client caches the importing node as the slot owner, but that node replies MOVED back to the source for keys not yet migrated, so requests ping-pong until the redirect cap trips and the command fails with a 'too many redirects' style error. Latency spikes and intermittent errors appear only while a migration is running, which makes it look like a cluster fault rather than a client bug.

MOVED is a change-of-address card: file it, all future post goes to the new address. ASK is a neighbour saying 'that one parcel is already next door' — you fetch it there, but you do not rewrite the address book.

saying these in an interview costs you the question

  • Saying ASK and MOVED are the same thing with different wording
  • Updating the cached slot map on an ASK redirect
  • Forgetting the ASKING command and assuming the target will just serve the key
  • Claiming the node forwards the request internally instead of redirecting the client
  • Thinking MOVED indicates an error or an unhealthy node

context

open as a page

Walk through how a single hash slot is moved from one Redis Cluster master to another while the cluster keeps serving traffic. Which commands are involved and what does each node do during the window?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Mark the slot IMPORTING on the target and MIGRATING on the source with CLUSTER SETSLOT, then loop CLUSTER GETKEYSINSLOT plus MIGRATE to copy keys in batches while the source answers ASK for already-moved keys, and finish with CLUSTER SETSLOT <slot> NODE <target> on every master.

open as a page

A Redis Cluster client caches which node owns which hash slot. After a resharding or a failover that cache is stale — how does the client find out, and what goes wrong if it never refreshes?

level: middleimportance: should knowfreq 45%

basics

~20 s

A MOVED reply tells the client the slot moved; the client should update that slot and re-fetch the map with CLUSTER SHARDS. Many clients also refresh periodically or on triggers. Without refresh every command pays two round trips, and redirect caps eventually turn it into errors.

open as a page

You need to add a fourth master to a live three-master Redis Cluster and give it a fair share of the keyspace. Describe the operational procedure with redis-cli's cluster tooling and how you verify the result.

level: seniorimportance: should knowfreq 42%

basics

~20 s

Join the node with redis-cli --cluster add-node (it starts with zero slots), then move slots to it with --cluster reshard (or --cluster rebalance for an even split), attach a replica with --cluster add-node --cluster-slave, and verify with --cluster check and CLUSTER INFO showing cluster_state:ok and 16384 slots covered.

open as a page

You must double the shard count of a heavily loaded Redis Cluster that backs a latency-sensitive service, with no maintenance window. How do you plan and pace the slot migration, and what would make you refuse to do it online at all?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Survey key sizes first, add empty masters plus replicas ahead of time, then move slots in small paced chunks with a modest MIGRATE batch size, watching p99 and SLOWLOG between chunks. Refuse online if huge keys would block masters past cluster-node-timeout and trigger spurious failovers.

open as a page