skip to content

Replication & Cluster

You will learn how Redis scales beyond one box: asynchronous leader–replica replication and Redis Cluster's 16384 hash slots with their resharding and failover machinery. Interviewers push here to see whether you understand stale replica reads, cross-slot limits, and what really happens when a master dies.

part ofRedisoverview, primer and where to startread it →
on this pageshow

explore

questions

25

In Redis Cluster, how does the system decide which node stores a given key? Walk through the computation from the key string to the node that serves it.

level: juniorimportance: must knowfreq 62%

answer

  1. CRC16(key) mod 16384 = slot
  2. two hops: key to slot, slot to node
  3. 16384 = 2^14, fixed, not configurable
  4. hash tag hashes only the braces content
  5. rebalancing moves slot ownership, never rehashes keys

basics

~20 s

Redis Cluster computes CRC16 of the key modulo 16384, giving a hash slot. Each of the 16384 slots is owned by exactly one master node, so the slot decides the node. Keys are never hashed against node addresses.

solid answer

~50 s

Redis Cluster shards by **hash slot**, not directly by node. The keyspace is statically divided into 16384 slots. For any key the server computes `CRC16(key) mod 16384`, where `key` is the whole key name unless it contains a hash tag, in which case only the substring inside the first `{...}` is hashed. That yields a slot number in 0..16383. Separately, slot *ranges* are assigned to master nodes (e.g. node A owns 0-5460, B owns 5461-10922, C the rest). Routing is therefore two lookups: key to slot (a pure function, identical on every node and every client, forever) and slot to node (cluster state, gossiped on the cluster bus and cached by smart clients). Because the key-to-slot function never changes, rebalancing moves slot *ownership* rather than rehashing keys. `CLUSTER KEYSLOT foo` shows a key's slot; `CLUSTER SHARDS` shows which node owns it.

code

text · 8 lines
text
127.0.0.1:7000> CLUSTER KEYSLOT user:1000:profile
(integer) 1649
127.0.0.1:7000> CLUSTER KEYSLOT "user:{1000}:profile"
(integer) 5288
127.0.0.1:7000> CLUSTER KEYSLOT "orders:{1000}"
(integer) 5288
127.0.0.1:7000> CLUSTER COUNTKEYSINSLOT 5288
(integer) 2

go deeper

for a junior

Recall the formula and the two-step mapping: CRC16 mod 16384 gives the slot, and each slot belongs to one master.

for a middle

Add that hash tags override which substring is hashed, and that rebalancing moves slot ownership rather than rehashing keys.

for a senior

Explain why fixed slots beat a consistent-hashing ring operationally - enumerable ownership, a well-defined migration unit, verifiable full coverage - and how clients cache the map.

for a principal

Frame it as a client-assisted sharding design with no proxy in the data path, and discuss what the fixed slot granularity implies for imbalance and for keyspace design.

## The two-level mapping Redis Cluster deliberately does not hash keys onto nodes. It inserts a fixed intermediate layer called the **hash slot**. There are exactly 16384 slots, numbered 0 to 16383, and the entire keyspace is partitioned across them. Routing has two independent steps: 1. **key to slot** - a pure function, `CRC16(key) mod 16384`. It depends on nothing but the key bytes, so every node, every client library and `redis-cli` compute the same answer, and the answer never changes over the life of the cluster. 2. **slot to node** - cluster state. Each master claims a set of slot ranges; that ownership map is shared between nodes over the cluster bus (a separate binary gossip protocol on port + 10000) and handed to clients on request. Separating the two is the whole design. Adding or removing a node changes only step 2: some slots change owner and their keys are migrated. Step 1 is untouched, so no key ever needs to be recomputed, and a client that knows the key can compute the slot offline without talking to anyone. ## The hash function Redis uses CRC16 in the XMODEM/CCITT variant over the raw key bytes, then masks to 14 bits (`mod 16384` is exactly `& 16383` since 16384 = 2^14). CRC16 is not a cryptographic hash; it was chosen because it is extremely cheap and distributes ordinary key names - which share long common prefixes like `user:1000:profile` - evenly across slots. Uniformity in the low bits is what matters here, not collision resistance. Two different keys landing in the same slot is normal and harmless: a slot holds an arbitrary number of keys. ## The hash-tag exception If the key contains a `{`, and after it a `}` with at least one byte between them, Redis hashes **only that inner substring**. `user:{1000}:profile` and `orders:{1000}` both hash `1000` and therefore share a slot. Everything else about the key is ignored for routing purposes. This is the single escape hatch that lets an application place related keys together deliberately. ## From slot to node Every master in the cluster owns a subset of the 16384 slots; in a healthy cluster the union of all owned slots is the full range and no slot has two owners. Replicas own no slots of their own - they mirror their master's. The mapping is not a range partition of the key space (keys are not ordered), just an arbitrary assignment of slot numbers, which is why an operator can move individual slots to correct imbalance. If a client sends a command to the wrong node, the node does not proxy it. It answers with a redirect telling the client which node owns that slot, and a well-behaved client refreshes its cached slot map and retries. Redis Cluster is therefore a *client-assisted* sharding design: there is no router process in the data path, which is why latency stays close to single-instance Redis. ## Why not consistent hashing A classic consistent-hashing ring maps keys onto node identities via a hash of the node address, so the set of keys a node owns changes implicitly whenever membership changes, and ownership is hard to state exactly. Fixed slots give explicit, enumerable ownership: a slot is either yours or not, migration is a well-defined unit of work, and the cluster can verify that all 16384 slots are covered. It also makes hash tags meaningful - co-location is a property of the slot, and the slot is what moves. ## Practical consequences - Multi-key commands only work when all keys are in one slot, because a single node must be able to execute them locally. - Slots are the unit of migration and of imbalance. A single slot cannot be split, so a slot that holds a disproportionate share of data (usually because of an over-broad hash tag) cannot be relieved by rebalancing. - `CLUSTER KEYSLOT <key>` asks the server to run the same function for you; `CLUSTER COUNTKEYSINSLOT <slot>` tells you how many keys are in a slot; `CLUSTER SHARDS` (Redis 7.0+, replacing `CLUSTER SLOTS`) prints ownership. - Because the key-to-slot function is public and stable, clients precompute slots to route commands and to group pipelined commands per node. ## Common confusion The number of slots is fixed at 16384 and is not configurable, and it is unrelated to the number of nodes. A three-node cluster and a thirty-node cluster both have 16384 slots; they differ only in how many slots each node owns. Similarly, slots are not databases: cluster mode supports only database 0.

  • If the key-to-slot function is fixed, how does adding a fourth node to a three-node cluster change anything?
    Only slot ownership changes. The operator assigns a subset of slots from the existing masters to the new node, and the keys currently living in those slots are migrated to it. No key changes its slot number, and clients keep computing slots exactly as before - they just refresh the slot-to-node map.
  • Do replicas own slots?
    No. Slot ownership belongs to masters; a replica serves as a copy of its master's slots and is advertised alongside it in CLUSTER SHARDS. Replicas answer reads only if the client opts in explicitly, and on failover the promoted replica takes over the master's slot ownership.

Slots are like the 16384 numbered pigeonholes in a mailroom: the address on a letter always maps to the same pigeonhole, and reorganising staff just reassigns which clerk empties which pigeonholes.

saying these in an interview costs you the question

  • Saying Redis Cluster uses consistent hashing with a ring of node hashes
  • Thinking the slot count is tuned to the number of nodes or is configurable
  • Believing a node proxies a command for a key it does not own
  • Assuming keys with the same prefix automatically land on the same node
  • Confusing hash slots with the 16 numbered databases of standalone Redis

context

open as a page

What does the Redis REPLICAOF command do to the instance it is run on, what can that instance still serve, and how do you detach it again?

level: juniorimportance: must knowfreq 50%

basics

~20 s

REPLICAOF host port makes the instance a replica: it discards its own dataset, copies the primary's, and then applies the primary's write stream. It serves reads but rejects writes (replica-read-only is on by default). REPLICAOF NO ONE detaches it into an independent primary, keeping the data it has.

open as a page

In Redis Cluster, how do the nodes decide that a primary is actually down rather than just slow, and what does the `cluster-node-timeout` setting control in that process?

level: middleimportance: must knowfreq 55%

basics

~20 s

Nodes ping each other over the cluster bus and gossip what they see. A node unreachable for longer than cluster-node-timeout is flagged PFAIL (possibly failed); once a majority of primaries report PFAIL for it, one marks it FAIL and broadcasts that, which triggers replica election.

open as a page

What does wrapping part of a Redis key in curly braces, as in user:{1000}:profile, do to where the key is placed in a Redis Cluster, and when would you deliberately use that?

level: middleimportance: must knowfreq 55%

basics

~20 s

Curly braces form a hash tag: Redis hashes only the substring inside the first {...} instead of the whole key. Keys sharing a tag share a slot and therefore a node, which is how you co-locate keys that must be read or written together.

open as a page

An MGET over three keys that worked against a single Redis instance now fails with "CROSSSLOT Keys in request don't hash to the same slot" after a move to Redis Cluster. Explain the rule being enforced and the realistic ways to make that read work.

level: middleimportance: must knowfreq 58%

basics

~20 s

In cluster mode every key of a single command must live in one hash slot, because one node executes it locally with no cross-node coordination. Fix it by hash-tagging the keys into one slot, by issuing separate GETs pipelined per node, or by restructuring the data into one key.

open as a page

In Redis Cluster, a client can receive either a -MOVED or an -ASK error reply. What does each one mean, and how must a correct client react differently to them?

level: middleimportance: must knowfreq 62%

basics

~20 s

MOVED means the slot permanently belongs to another node: follow it and update the cached slot map. ASK means only this key has already moved during an in-progress migration: send ASKING then the command to the target once, and do not update the map.

open as a page

Walk through what happens on both the primary and the replica during a Redis full synchronization, when a replica connects for the first time.

level: middleimportance: must knowfreq 55%

basics

~20 s

The replica sends PSYNC with no known history; the primary replies FULLRESYNC with its replication ID and offset, produces an RDB snapshot from a forked child (to disk or streamed directly), and buffers all writes made meanwhile in that replica's output buffer. The replica flushes its data, loads the RDB, then applies the buffered and ongoing stream.

open as a page

A Redis Cluster primary that owns a range of hash slots has been marked as failed. Describe how one of its replicas becomes the new owner of those slots.

level: seniorimportance: must knowfreq 48%

basics

~20 s

Eligible replicas wait a short delay ranked by how far behind they are, then request votes for a new configuration epoch. Primaries each vote once per epoch; a replica winning a majority bumps its config epoch, claims the slots and announces itself, and the old primary rejoins as its replica.

open as a page

In Redis Cluster, a primary that owns a range of hash slots can keep replying OK to client writes for seconds after the cluster has already promoted one of its replicas for those same slots. Explain why that window exists, what happens to those writes when the old primary rejoins the cluster, and how the `cluster-node-timeout` setting sizes that window against the risk of unnecessary failovers.

level: seniorimportance: must knowfreq 50%

basics

~20 s

An isolated primary stops accepting writes only after failing to reach a majority of primaries for cluster-node-timeout, so it keeps serving its side meanwhile. On rejoin the promoted replica's higher configEpoch wins; the old primary demotes and discards those writes.

open as a page

What extra constraints does Redis Cluster impose on Lua scripts executed with EVAL and on transactions opened with MULTI, compared with running the same code against a single Redis instance?

level: seniorimportance: must knowfreq 46%

basics

~20 s

Both run entirely on one node. Every key a script or transaction touches must hash to one slot, all script keys must be declared in KEYS (Redis refuses undeclared or non-local key access), and a cross-slot command inside MULTI aborts the transaction.

open as a page

Walk through how a single hash slot is moved from one Redis Cluster master to another while the cluster keeps serving traffic. Which commands are involved and what does each node do during the window?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Mark the slot IMPORTING on the target and MIGRATING on the source with CLUSTER SETSLOT, then loop CLUSTER GETKEYSINSLOT plus MIGRATE to copy keys in batches while the source answers ASK for already-moved keys, and finish with CLUSTER SETSLOT <slot> NODE <target> on every master.

open as a page

Redis replicates asynchronously. When a client receives +OK for a SET command, what has actually been guaranteed, and what does the WAIT command add on top?

level: seniorimportance: must knowfreq 48%

basics

~20 s

+OK means only that the write was applied in the primary's memory. Replicas are sent the write after the reply, so an acknowledged write can vanish if the primary dies before propagating and a replica is promoted. WAIT numreplicas timeout blocks until that many replicas acknowledge the offset and returns how many did; it never rolls anything back.

open as a page

What do the Redis settings `min-replicas-to-write` and `min-replicas-max-lag` actually protect against, and what do they fail to guarantee?

level: middleimportance: should knowfreq 35%

basics

~20 s

They make a primary refuse writes unless at least N replicas were acknowledging it within the last few seconds. That shrinks the window of writes that exist on only one node, but it is a check made before the write, not an acknowledgement after it — so accepted writes can still be lost.

open as a page

Which Redis commands and features behave differently, or stop working entirely, when an application moves from a single Redis instance to Redis Cluster mode?

level: middleimportance: should knowfreq 45%

basics

~20 s

Only database 0 exists, so SELECT beyond 0, SWAPDB and MOVE are gone. Multi-key commands, scripts and transactions must stay within one slot. KEYS, SCAN, DBSIZE, FLUSHALL and RANDOMKEY act per node and need fan-out. Classic Pub/Sub broadcasts cluster-wide unless you use sharded channels.

open as a page

A Redis Cluster client caches which node owns which hash slot. After a resharding or a failover that cache is stale — how does the client find out, and what goes wrong if it never refreshes?

level: middleimportance: should knowfreq 45%

basics

~20 s

A MOVED reply tells the client the slot moved; the client should update that slot and re-fetch the map with CLUSTER SHARDS. Many clients also refresh periodically or on triggers. Without refresh every command pays two round trips, and redirect caps eventually turn it into errors.

open as a page

Your service sends read traffic to Redis replicas. What consistency properties can a read from a replica be relied on for, and which ones can it not? Explain what the `master_repl_offset`, `slave_repl_offset` and `master_link_status` fields of the Redis `INFO replication` section let you observe, and what changes when the replica is configured with `replica-serve-stale-data no` and its link to the primary drops.

level: middleimportance: should knowfreq 45%

basics

~20 s

A replica read reflects the primary at an earlier offset — staleness tracks replication lag but is never bounded by a promise. No read-your-writes after a primary write, and values can appear to move backwards between replicas. INFO replication shows the offsets and master_link_status; replica-serve-stale-data no makes a disconnected replica error instead of answering.

open as a page

A Redis Cluster starts rejecting commands with CLUSTERDOWN and CLUSTER INFO reports cluster_state:fail. Explain what hash-slot coverage means, how you identify which slots have no owner, and what the cluster-require-full-coverage setting changes.

level: seniorimportance: should knowfreq 34%

basics

~20 s

Coverage means all 16384 slots are claimed by a reachable master. If any slot has no owner, the cluster marks itself failed and, with cluster-require-full-coverage yes (the default), refuses every command - even for slots that are fine. Set it to no to keep serving the covered slots.

open as a page

How do you inspect which hash slots each Redis Cluster node owns, verify that two specific keys will be served by the same node, and how does a client library know where to send a command?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Use CLUSTER SHARDS (or the older CLUSTER SLOTS) for slot-range ownership, CLUSTER NODES for the raw view, and CLUSTER KEYSLOT on each key to compare slots. Smart clients fetch that map once, cache it, compute slots locally with CRC16, and refresh when the server says a slot moved.

open as a page

A colleague proposes putting every key belonging to a customer inside a hash tag such as {tenant:42} so that multi-key commands and Lua scripts keep working after a move to Redis Cluster. What does that buy, and what are the operational risks?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It buys legal multi-key commands and single-node atomicity per tenant. It costs balance: all of a tenant's data and traffic sit in one indivisible slot, so a large or busy tenant creates a hotspot that resharding cannot fix, since slots can be moved but never split.

open as a page

You need to add a fourth master to a live three-master Redis Cluster and give it a fair share of the keyspace. Describe the operational procedure with redis-cli's cluster tooling and how you verify the result.

level: seniorimportance: should knowfreq 42%

basics

~20 s

Join the node with redis-cli --cluster add-node (it starts with zero slots), then move slots to it with --cluster reshard (or --cluster rebalance for an even split), attach a replica with --cluster add-node --cluster-slave, and verify with --cluster check and CLUSTER INFO showing cluster_state:ok and 16384 slots covered.

open as a page

A Redis replica reconnects after a 20-second network blip and performs a full resynchronization instead of a partial one. Which mechanism should have prevented that, and how do you tune it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Partial resync uses the primary's replication backlog, a fixed-size circular buffer of the recent write stream (repl-backlog-size, default 1mb). If the replica's offset has fallen out of it, or the replication ID no longer matches, the primary must send everything again. Size the backlog as peak write bytes per second times the outage you want to survive.

open as a page

You are laying out a Redis Cluster across three availability zones. Which placement and replication decisions determine whether the cluster keeps serving — and keeps its writes — when an entire zone disappears?

level: principalimportance: should knowfreq 28%

basics

~20 s

Spread primaries so no zone holds half or more of them, since failover needs a majority of primaries; never place a replica in its primary's zone; keep spare replicas so a promoted node is not left bare; and decide explicitly whether uncovered slots should take the whole cluster down.

open as a page

You are moving an application that leans heavily on multi-key Redis commands - set intersections, MSET batches, and Lua scripts touching several keys - onto Redis Cluster. How do you decide, per access path, between co-locating keys with hash tags, restructuring the data, or moving the work into the application?

level: principalimportance: should knowfreq 32%

basics

~20 s

Classify each access path by whether it truly needs atomicity. Paths that do get a narrow hash tag; paths that merely batch get split and pipelined per node; recurring groups get collapsed into one key. Reject any co-location group that is unbounded or uneven.

open as a page

Redis Cluster uses exactly 16384 hash slots. Why that number rather than, say, 1024 or 65536?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

16384 balances two forces: each cluster-bus heartbeat carries a bitmap of the sender's slots, so 16384 bits is 2 KB per packet while 65536 would be 8 KB; and since clusters are not expected to exceed roughly 1000 nodes, 16384 slots still give fine-grained rebalancing.

open as a page

You must double the shard count of a heavily loaded Redis Cluster that backs a latency-sensitive service, with no maintenance window. How do you plan and pace the slot migration, and what would make you refuse to do it online at all?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Survey key sizes first, add empty masters plus replicas ahead of time, then move slots in small paced chunks with a modest MIGRATE batch size, watching p99 and SLOWLOG between chunks. Refuse online if huge keys would block masters past cluster-node-timeout and trigger spurious failovers.

open as a page