A Redis Cluster client caches which node owns which hash slot. After a resharding or a failover that cache is stale — how does the client find out, and what goes wrong if it never refreshes?
answer
- MOVED = refresh signal; ASK = never cache
- CLUSTER SHARDS re-fetch (7.0+)
- reactive + periodic + adaptive triggers
- Lettuce refresh options are off by default
- stale map = 2 RTTs, then redirect-cap errors
basics
~20 sA MOVED reply tells the client the slot moved; the client should update that slot and re-fetch the map with CLUSTER SHARDS. Many clients also refresh periodically or on triggers. Without refresh every command pays two round trips, and redirect caps eventually turn it into errors.
solid answer
~50 sThe authoritative signal is a `-MOVED` reply: the node says the slot now belongs elsewhere. A correct client follows the redirect, updates the slot entry, and usually schedules a full topology re-fetch via `CLUSTER SHARDS` (7.0+; `CLUSTER SLOTS`/`CLUSTER NODES` on older servers), because one MOVED usually means a whole range moved. Good clients combine three mechanisms: **reactive** refresh on MOVED/ASK, **periodic** refresh on a timer, and **adaptive** triggers such as persistent reconnects or a node being seen as unreachable. Several libraries ship with periodic and adaptive refresh **disabled by default** — Lettuce's `ClusterTopologyRefreshOptions` is the well-known example — so it must be enabled explicitly. If the client never refreshes: every command to a moved slot costs one extra round trip (redirect + retry), so throughput drops and p99 rises while nothing looks broken; connections to decommissioned nodes are retained; and after a failover the redirect cap (typically 5) can be exhausted, turning silent slowness into hard errors.
code
text · 13 linesClusterTopologyRefreshOptions refresh = ClusterTopologyRefreshOptions.builder()
.enablePeriodicRefresh(Duration.ofSeconds(30))
.enableAllAdaptiveRefreshTriggers() // MOVED, ASK, reconnects, unknown node
.adaptiveRefreshTriggersTimeout(Duration.ofSeconds(10))
.build();
clusterClient.setOptions(ClusterClientOptions.builder()
.topologyRefreshOptions(refresh)
.maxRedirects(5)
.build());
// Without this block the client follows redirects per command
// but may keep a stale slot map for a long time.go deeper
Know that the client caches a slot map and that a MOVED reply is the signal to update it.
Describe reactive, periodic and adaptive refresh, name CLUSTER SHARDS, and state that these options are often off by default.
Quantify the failure mode — extra round trip per command, redirect-cap errors after failover — and describe the metrics you would watch to prove convergence.
Treat client library maturity and configuration as an availability dependency of the cluster, and set standards for refresh settings, redirect caps and redirect-rate SLOs across services.
## Where the map comes from A cluster client bootstraps by connecting to a seed node and asking for topology. On Redis 7.0+ that is `CLUSTER SHARDS`, which returns, per shard, the owned slot ranges plus every node's endpoint, role and health; older code uses `CLUSTER SLOTS` (deprecated in 7.0) or parses `CLUSTER NODES`. The result is an array of 16384 slot → node entries, plus a connection pool per node. From then on the client computes `CRC16(key) mod 16384` locally and sends each command straight to the owner — one round trip, no proxy. ## What invalidates it - **Resharding.** Slots are explicitly moved from one master to another; when the migration commits with `CLUSTER SETSLOT <slot> NODE <id>`, ownership changes. - **Failover.** A replica is promoted and inherits all of its master's slots, so many slots change endpoint at once. - **Scale in/out.** Nodes join or leave, so the endpoint set itself changes. The server does not push notifications about any of this to ordinary clients; the map is refreshed by the client, driven by errors it observes. ## The three refresh mechanisms **Reactive (mandatory).** A `-MOVED <slot> <host>:<port>` reply is a definitive statement of new ownership. The minimum correct behaviour is to retry against the named node and update that slot's entry. The usual improvement is to also mark the whole map dirty and re-fetch, since resharding and failover move contiguous ranges — updating one slot at a time means one extra round trip per slot in the range. Note that `-ASK` must **not** update the map: it is a per-key, one-shot redirect during a live migration, and caching its target causes flapping. But an ASK is still useful as a *hint* that the topology is in flux; some clients use it as an adaptive trigger. **Periodic.** A background timer re-fetches topology every N seconds (30–60 s is typical). This bounds staleness even when traffic is too low for a MOVED to be observed, and cleans up connections to nodes that have left. The cost is a small constant load of `CLUSTER SHARDS` calls. **Adaptive.** Refresh triggered by specific events: MOVED redirect, ASK redirect, persistent reconnects to a node, an unknown node in a reply, or a connection attempt failing. This gives fast convergence after failover without a tight polling loop. ## The default-off trap The practically important detail is that these are frequently **not on by default**. Lettuce, for instance, requires an explicit `ClusterTopologyRefreshOptions` with `enablePeriodicRefresh(...)` and/or `enableAllAdaptiveRefreshTriggers()` on the `ClusterClientOptions`; without it the client relies purely on per-command redirect following and can hold a stale view for a long time. Other libraries differ, and the only reliable move is to read the client's cluster options rather than assume. This is the concrete configuration answer an interviewer is often fishing for. ## Symptoms of never refreshing 1. **Silent latency tax.** Every command to a moved slot becomes redirect + retry: two round trips instead of one, roughly halving throughput on the affected fraction of traffic. Nothing errors, so dashboards show "Redis got slower" with no cause. This is the most common production manifestation. 2. **Redirect exhaustion.** Clients cap redirect chains (commonly 5). Combine a stale map with an in-flight migration (MOVED to a node that then replies ASK, etc.) and the cap trips, surfacing errors like "too many cluster redirections". 3. **Dead connections.** After scale-in, pools keep entries for nodes that no longer exist, producing connection errors and slow failure paths. 4. **Reads to the wrong role.** Clients doing replica reads with `READONLY` need role information; a stale map can keep sending reads to a node that has been demoted, which then redirects them. ## Verifying it in practice Instrument the client: most expose counters for redirects followed and topology refreshes. A healthy cluster in steady state should show a redirect rate near zero, with a burst during a reshard or failover that decays within seconds. A persistently non-zero redirect rate means the map is not converging — check whether periodic/adaptive refresh is actually enabled, and whether ASK is being mistaken for MOVED.
- Why is it wrong for a client to refresh its slot map on an ASK redirect the way it does on MOVED?ASK does not mean ownership changed; the slot is still owned by the source node and only that one key has already been transferred. Rewriting the map to point at the importing node makes every subsequent key in that slot go there, where it is bounced back with MOVED, so the client flaps between nodes until the redirect cap trips. ASK may legitimately be used as a hint to schedule a refresh, but the map entry itself must not change.
- Your cluster is stable, no resharding is running, yet client metrics show a steady 3% redirect rate. What would you look at?A steady redirect rate with a stable topology points at the client, not the cluster. Check whether periodic and adaptive refresh are enabled at all, whether the client updates only the single redirected slot rather than re-fetching the range, and whether some instances were started before the last topology change and never converged. Also run redis-cli --cluster check for slots left in migrating or importing state from an aborted reshard, which produces persistent ASK and TRYAGAIN traffic.
saying these in an interview costs you the question
- Assuming the server pushes topology changes to clients automatically
- Caching the ASK target as the new slot owner
- Thinking a stale map is harmless because redirects are followed anyway
- Believing every client library refreshes topology periodically out of the box
- Confusing the client's slot map with the cluster's own gossip state