Explain how splitting one heavily-read cache entry into several Redis keys with different suffixes (for example `product:123:0` through `product:123:9`) relieves load, and what that technique costs you.
answer
- Change the name to change the slot
- No hash tag — braces would co-locate the copies
- Read 1, write N, memory N
- Stagger TTLs or you rebuild the stampede
- Counter version: shard the INCR, sum on read
basics
~20 sYou store N identical copies under different key names. Because the names hash to different slots, the copies land on different cluster nodes, and each reader picks one at random — so read traffic divides by N. You pay N times the memory, an N-key write fan-out that is not atomic, and a window where copies disagree.
solid answer
~1 minWrite the same value under `product:123:0 … product:123:9` (deliberately **no** hash tag, so `CRC16` spreads them across different slots and therefore different masters). A reader picks a suffix — random, or derived from the client/instance id for better local-cache reuse — and issues a normal `GET`. Read load on any one node drops to roughly 1/N. The costs are real: - **Memory ×N** — only worth it for a handful of known-hot keys, never as a default key scheme. - **Write fan-out** — every update or invalidation must touch all N keys, and since they are in different slots you cannot wrap them in `MULTI` or a single Lua script in cluster mode; it is N independent commands that can partially fail. - **Inconsistency window** — during a fan-out some readers see old and some see new. - **Correlated TTLs** — if all copies are written together with the same TTL, they expire together and you get a stampede on N keys at once; stagger them. - **Operational complexity** — you need a list of which keys are split, ideally driven by hot-key detection rather than hard-coded guesses. And note: on a *single* Redis instance splitting buys almost nothing, since all copies still sit behind the same command thread. It pays off across cluster nodes or across separate instances.
code
text · 11 lines# write path: fan out N copies with jittered TTLs
SET product:123:0 "{...}" EX 300
SET product:123:1 "{...}" EX 306
SET product:123:2 "{...}" EX 312
# read path: pick one copy
GET product:123:2
# WRONG: braces make Redis hash only "product:123",
# so all copies land in the SAME slot on the SAME node
SET {product:123}:0 "{...}" EX 300go deeper
Know the shape: store several identical copies under different key names and read a random one.
Explain why different names mean different slots and nodes, and that writes must fan out to every copy.
Discuss suffix-selection strategy, non-atomic cross-slot invalidation, TTL staggering, and when splitting loses to a near cache or replica reads.
Treat it as a last resort with an operational cost; require measurement-driven selection of split keys, a bounded N, and an explicit statement of the consistency the product is buying relief with.
## The idea A hot key cannot be load-balanced because placement is a pure function of the key name. So change the name. Instead of one `product:123`, keep N replicas of the same value under N different names, and have each reader consult one of them. ``` write: SET product:123:0 <v> EX 300 SET product:123:1 <v> EX 307 ... (N copies) read: GET product:123:<random 0..N-1> ``` This is sometimes called suffix sharding, key salting, or fan-out caching. ## Why it works — and the condition that makes it work In Redis Cluster the slot is `CRC16(key) mod 16384`. Adding a suffix changes the hash, so the copies scatter across slots and — as long as slots are distributed over several masters — across nodes. Each node then has its own command thread, its own NIC and its own CPU, so aggregate capacity for that logical entry multiplies. The critical condition: **do not use a hash tag**. Redis Cluster hashes only the substring inside `{...}` when a key contains braces, so `{product:123}:0` and `{product:123}:1` deliberately land in the *same* slot — the exact opposite of what you want here. Hash tags exist to co-locate keys for multi-key operations; using them while splitting a hot key silently defeats the whole exercise. Conversely, if you are on a single non-clustered instance, splitting has almost no effect on CPU: same process, same thread. It can still slightly help by breaking up per-key contention in some client-side structures, but the honest answer is that splitting is a cluster technique (or a technique for spreading across several independent Redis instances). ## Choosing the suffix at read time Two strategies: - **Random per request** — perfectly even spread, but every application instance touches every copy, so no locality and each instance's near cache (if any) sees all N. - **Sticky per instance/thread** — e.g. `hash(instance_id) mod N`. Traffic is even in aggregate when instance count ≫ N, and each instance always talks to the same copy, which pairs nicely with a local cache and gives more predictable connection routing. The risk is uneven spread when N is close to the number of instances. Choose N from the arithmetic, not from a round number: if one node can serve 100k rps for this value and you need 400k, N=5 or 6 gives headroom. Bigger N multiplies memory and write cost for nothing. ## The write path is where it gets expensive Updating or invalidating means touching all N keys. - In cluster mode the copies are in different slots, so a single `MULTI`/`EXEC` or a single `EVAL` cannot cover them — cross-slot multi-key commands are rejected. You issue N separate commands (pipelined per node, ideally). - That is not atomic. A crash or a partial failure mid fan-out leaves a mix of old and new copies, and a reader picking at random gets a coin flip. For a cache this is usually acceptable (it heals at the next write or TTL) but it must be a conscious decision — never split a key whose readers require a consistent view across successive requests. - Deleting is the same problem: an invalidation that misses one copy leaves stale data alive until its TTL. Because writes cost N× and reads cost 1×, splitting is right for read-hot, write-rare entries and wrong for write-hot ones. ## TTL correlation If all N copies are written in the same instant with the same TTL, they expire in the same instant, and the very key you split because it was hot now produces N simultaneous misses and N recomputes. Add a per-copy jitter to the TTL (a few percent) so expiries scatter, and combine with a recompute guard on the miss path. ## Hot *write* keys The mirror-image technique exists for counters: instead of `INCR views:123`, do `INCR views:123:<random 0..N-1>` and sum the N shards on read (`MGET` plus addition, or a Lua script if they are co-located — here you *do* want a hash tag so the sum is one round trip). Writes divide by N; reads cost N reads. That is a good trade when writes vastly outnumber reads, and the reverse of the read-splitting trade. ## Operating it Splitting is targeted surgery, not a schema. Keep the set of split keys in configuration (or derive it from your hot-key detection pipeline), keep N small, document the fan-out in the write path so the next engineer does not add a fifth writer that updates only one copy, and add a metric for "copies disagreed" if correctness matters. Before reaching for it, check the cheaper options: shrink the value, add a one-second in-process cache, or route reads to replicas — all of which avoid the write-path complexity entirely.
- How does the technique change when the hot key is a counter being incremented rather than a value being read?You invert it: shard the writes instead of the reads. `INCR` goes to one of N shard keys chosen at random, so write load divides by N, and readers sum all N shards. Here you usually *want* a hash tag so the N shards share a slot and the sum is a single `MGET` or Lua call. It trades cheap writes for more expensive reads, which is correct when writes dominate.
- Why must you avoid hash tags when splitting a hot read key?Redis Cluster hashes only the substring between the first `{` and the following `}` when one is present. So `{product:123}:0` and `{product:123}:1` produce the same slot and live on the same master — the copies would still all hit one node and one command thread. Hash tags are for co-locating keys you need to operate on together, which is the opposite goal.
Instead of one noticeboard everyone crowds around, you post ten identical copies in ten corridors. Reading is instant; keeping all ten posters up to date is now ten trips.
saying these in an interview costs you the question
- Using a hash tag around the shared part of the key, which co-locates every copy and defeats the split
- Claiming splitting helps on a single non-clustered instance in the same way it does across a cluster
- Forgetting that invalidation must touch all N copies, leaving stale copies alive until TTL
- Writing all copies with an identical TTL so they expire simultaneously
- Applying splitting as a general key-naming convention instead of targeted surgery on measured hot keys