In ElastiCache for Valkey or Redis OSS, what is the difference between a replication group with cluster mode disabled and one with cluster mode enabled, and what does that choice force on the application's client library?
answer
- one shard versus many shards
- replicas add reads, not writes
- the client must follow redirections
- configuration endpoint, not primary endpoint
- multi-key work must stay on one shard
basics
~20 sCluster mode disabled means one shard holding the whole keyspace, scaled up by node size. Cluster mode enabled splits the keyspace across many shards, each with its own primary, and requires a cluster-aware client that follows slot redirections.
solid answer
~50 sWith **cluster mode disabled**, a replication group is a single shard: one primary that holds the entire keyspace, plus up to five read replicas. You scale memory and write throughput only by moving to a larger node type; replicas add read capacity and failover targets, nothing else. Clients connect to a primary endpoint and a reader endpoint and can be any ordinary client. With **cluster mode enabled**, the keyspace is divided across multiple shards, each with its own primary and replicas, and you scale horizontally by adding shards and resharding online. Clients connect to a **configuration endpoint**, and the library must speak the cluster protocol: fetch the slot-to-node map, keep it cached, and follow redirections when a slot moves. Multi-key commands, transactions and scripts must touch keys that live on the same shard. So the practical rule is: ship a cluster-capable client first, then enable cluster mode.
code
bash · 10 lines# Add a shard to a live cluster-mode-enabled group (online resharding)
aws elasticache modify-replication-group-shard-configuration \
--replication-group-id app-cache \
--node-group-count 4 \
--apply-immediately
# Inspect the resulting shard layout
aws elasticache describe-replication-groups \
--replication-group-id app-cache \
--query 'ReplicationGroups[0].NodeGroups[].NodeGroupId'go deeper
Recall that cluster mode disabled is one primary holding everything, while cluster mode enabled splits the data across several primaries. Know that replicas are copies, not extra capacity for writes.
Explain the endpoints each shape exposes, how you scale each one, and why a sharded cluster requires a client that caches a slot map and follows redirections.
Show the judgment: size the working set and write rate against a single node, weigh the partial-failure benefit of sharding, and sequence a migration so the client change lands before the topology change.
Own the default for the platform — whether teams start sharded for headroom or unsharded for simplicity — and the standard client configuration that makes either choice safely reversible.
## Two shapes of the same replication group ElastiCache calls both of these a *replication group*, which hides how differently they behave. The switch is `cluster mode`, and it decides whether the cache is one shard or many. ## Cluster mode disabled A cluster-mode-disabled replication group has exactly **one shard**: - one primary node holds the **entire** keyspace; - up to five read replicas hold copies; - the group exposes a **primary endpoint** (always the current primary) and a **reader endpoint** (spread across the replicas). What this means for capacity is the part candidates get wrong. Replicas do not add write throughput and do not add memory — every replica holds the same full dataset. The only way to hold more data or absorb more writes is a **larger node type**, which ElastiCache performs online but with a failover along the way. The whole working set must fit in one node's memory, minus the headroom the engine needs. What you get in exchange is simplicity: any ordinary client works, multi-key commands and scripts have no placement restriction, and there is exactly one topology to reason about. ## Cluster mode enabled A cluster-mode-enabled replication group has **many shards** (AWS calls them node groups). The keyspace is partitioned across 16,384 hash slots, and each shard owns a contiguous range of them. Every shard has its own primary and its own replicas, and every shard fails over independently. Capacity now scales in two dimensions: memory and write throughput both grow as you add shards, because each primary owns only its slice. ElastiCache supports **online resharding** — adding or removing shards and rebalancing slot ownership while the cluster keeps serving — and you can drive shard or replica counts with ElastiCache Auto Scaling target-tracking policies. ## What it forces on the client This is the operational heart of the question: 1. **The library must support cluster mode.** It connects to the configuration endpoint, retrieves the slot-to-node map, and opens connections to every shard. When a slot moves — a resharding, a failover — the server answers with a redirection and the client must follow it and refresh its cached map. A plain non-cluster client pointed at the configuration endpoint will work by accident for keys that happen to land on the node it reached, and fail for everything else. 2. **Multi-key work is constrained.** A command touching several keys, a transaction, or a script must operate on keys that hash to the same slot; otherwise the server rejects it. Applications that batch reads across unrelated keys need reworking before the move. 3. **Connection count grows.** Instead of one pool to a primary, the client maintains connections per shard, which matters for connection-heavy runtimes. ```bash # Three shards, one replica each — cluster mode enabled aws elasticache create-replication-group \ --replication-group-id app-cache \ --replication-group-description "sharded app cache" \ --engine valkey \ --cache-node-type cache.m7g.large \ --num-node-groups 3 \ --replicas-per-node-group 1 \ --automatic-failover-enabled \ --multi-az-enabled ``` `--num-node-groups` greater than one is what makes it a sharded cluster; the same call with `--num-node-groups 1` and no sharding intent is the cluster-mode-disabled shape. ## How to choose Start from the dataset and the write rate: - **Cluster mode disabled** when the working set fits comfortably in one node type you are willing to pay for, writes fit within one primary, and the application uses multi-key operations freely. It is less to operate and less to get wrong. - **Cluster mode enabled** when the dataset is larger than a single node, is growing unpredictably, or the write rate exceeds one primary. Also when you want failure to be partial rather than total: one shard's failover disturbs a fraction of the keyspace instead of all of it. ## The migration path You are not locked in. ElastiCache supports migrating a cluster-mode-disabled replication group to cluster mode enabled online. The ordering that matters is the client: deploy a cluster-capable client and fix multi-key access paths **before** the migration, because the day the topology changes is the wrong day to discover a library that cannot follow a redirection. ## What candidates get wrong The two classic errors are believing that adding replicas increases write capacity, and assuming any client works against a sharded cluster. A third, subtler one is treating cluster mode enabled as the safe default: for a modest cache it adds client requirements and access-pattern restrictions in exchange for scale you do not need yet.
- You started with cluster mode disabled and now need more write throughput than one primary can absorb. What are your options, in order?First, scale up the node type — online, but it involves a failover and eventually hits the largest node available. Second, migrate the group to cluster mode enabled and add shards with online resharding. The migration is supported live, but the ordering matters: deploy a cluster-capable client and fix multi-key access paths first, then change the topology.
- Which application behaviours break when you move from cluster mode disabled to cluster mode enabled?Anything that assumes all keys live on one node: multi-key commands, transactions and scripts spanning keys in different slots are rejected, and full-keyspace iteration now has to be run per shard. Connection handling also changes, since the client maintains a pool per shard instead of one pool to a single primary.
- Does a cluster-mode-enabled replication group with a single shard behave like a cluster-mode-disabled one?Operationally it looks similar — one primary, its replicas — but it still speaks the cluster protocol, so clients must be cluster-capable and slot restrictions on multi-key commands still apply. The advantage is that you can add shards later without a topology migration, so it is a reasonable starting point when growth is expected.
saying these in an interview costs you the question
- Thinks cluster mode disabled means there are no replicas
- Says adding read replicas increases write throughput
- Assumes any client library works against a sharded cluster
- Treats cluster mode enabled as the default for every cache
- Believes each replica holds only part of the keyspace