In a Redis-backed cache, what is a "hot key", and why can a single popular key become a bottleneck even though a Redis node handles hundreds of thousands of operations per second?
answer
- One key → one slot → one node → one thread
- Bandwidth wall before CPU wall
- Adding nodes moves slots, not a single key
- Head-of-line blocking hurts the whole shard
- Hot ≠ big, but hot + big is worst
basics
~20 sA hot key is one key taking a hugely disproportionate share of traffic. Because a key's name determines exactly one shard, and that shard executes commands on one thread, the load cannot be spread by adding nodes — one key saturates one CPU core and one network link.
solid answer
~50 sA hot key is a single key whose request rate is orders of magnitude above the average key — a celebrity profile, a homepage config blob, a global counter, a feature flag. It hurts for two reasons. First, placement is deterministic: in Redis Cluster the key name is hashed (CRC16 mod 16384) to one slot, and that slot lives on exactly one master. Adding nodes reshards *other* keys away; the hot one stays put. Second, that node executes commands on a single thread, so all traffic for the key queues behind one core, and every other key on that node queues behind it too — the blast radius is the whole shard, not just the key. The usual limits hit in this order: outbound network bandwidth (value size × rps), then CPU on the command thread, then client connection pools timing out. Mitigation is detection first (`redis-cli --hotkeys`, client-side sampling), then splitting the key across slots, replica reads, or a short-TTL in-process cache.
go deeper
Be able to say what a hot key is with an example (a celebrity record, a global config blob) and that all its traffic lands on one node.
Explain the key → CRC16 slot → single master → single command thread chain, and name bandwidth and CPU as the two limits.
Add diagnosis (skew across nodes, head-of-line blocking on the shard, retry amplification) and pick a mitigation with its cost, distinguishing read-hot from write-hot.
Frame it as a data-placement and traffic-distribution problem: which keys are allowed to be global, what staleness the product can buy relief with, and whether the hot entity belongs in Redis at all.
## What "hot" means Real cache traffic is never uniform. Access frequency follows a heavy-tailed (Zipf-like) distribution: a handful of keys absorb a large fraction of all requests. A **hot key** is a key on that extreme tail — think the cached record for a celebrity account, a global `config:features` blob every request reads, a leaderboard, a rate-limit counter for a shared tenant, or the cached front page. Hot is about **request rate**, not size. A separate problem is a **big key** — one key holding a huge value or a multi-million-element collection. They compound badly: a big key that is also hot multiplies bytes per second, and collection commands over a big key (`HGETALL`, `LRANGE 0 -1`, `SMEMBERS`) are O(N), so each hit costs real CPU rather than a pointer dereference. ## Why sharding does not save you Redis Cluster places data by hashing the key name: `CRC16(key) mod 16384` gives a hash slot, and each master owns a contiguous set of slots. This is deterministic and stateless — every client computes the same answer. That is what makes the cluster fast and coordination-free, and it is exactly why a hot key cannot be load-balanced. `product:42` always maps to the same slot and therefore to the same master. Resharding moves *slots*; it can move the hot key to a quieter node, but it cannot split a single key's traffic, because a key lives in exactly one slot. So the standard horizontal-scaling reflex — add nodes — buys nothing for the hot key itself. It only helps if the node was also busy with other work you can move away. ## Why one key can saturate a node Commands on a Redis node are executed by one thread (I/O threading added in 6.0 parallelizes socket reads/writes, not command execution). Practical consequences: - **Bandwidth first.** This is usually the wall you hit before CPU. A 10 KB cached JSON blob served 50,000 times a second is 4 Gbit/s of egress from one process — more than a 1 GbE link, and enough to make even a 10 GbE instance uncomfortable once replication traffic is added. Cloud instances also cap network per instance size. - **CPU next.** At small values Redis will do a few hundred thousand simple GETs per second per core. If each hit is an O(N) collection read or a Lua script, that ceiling drops by an order of magnitude. - **Head-of-line blocking.** Because one thread serves everything, the hot key's queue delays *every other key on that shard*. Your p99 for unrelated lookups degrades, which is why hot keys usually surface as "random unrelated timeouts on one node". - **Client-side amplification.** When latency rises, connection pools fill, callers time out and retry, and retries add load — a small skew turns into a cliff. ## Symptoms you will actually see One node in the cluster shows high `instantaneous_ops_per_sec` and CPU while its peers idle; `INFO stats` shows lopsided `total_net_output_bytes`; SLOWLOG may be empty (no single command is slow — there are just too many of them); client-side p99 on that shard's keys climbs while p50 stays fine. Uneven CPU across identically-sized cluster nodes is the classic tell. ## The mitigation menu 1. **Detect and quantify** before acting — `redis-cli --hotkeys` (needs an LFU eviction policy), `OBJECT FREQ` on suspects, or sampling in the cache client. Guessing which key is hot is how teams optimize the wrong thing. 2. **Split the key** into N copies with different suffixes so they hash to different slots and different masters; readers pick one at random or by client id. 3. **Read replicas** — each replica is a separate process with its own core and NIC, so N replicas multiply read capacity. Helps reads only, and reads become slightly stale. 4. **In-process (near) cache** with a very short TTL. The most effective by far for read-hot keys: a 1-second local TTL collapses any per-instance read rate to one Redis GET per second per instance. It costs staleness and per-instance divergence. 5. **Shrink the value** — hot plus large is the worst combination, so store only the fields callers need, or compress, before doing anything architectural. A hot **write** key is harder: replicas and near caches do nothing for it, since every write must reach the one master. There you either split the key (e.g. N counter shards summed on read), batch/aggregate writes in the application, or move the counter out of the critical path.
- How does a hot key differ from a big key, and can one key be both?Hot is about request rate; big is about the size of the value or the element count of the collection. A big key is dangerous because O(N) commands over it occupy the command thread for milliseconds, blocking everything else. A key that is both hot and big is the worst case: bytes per second and CPU per hit multiply, and it is usually the first thing to fix by shrinking the value.
- Why doesn't adding more nodes to Redis Cluster fix a hot key?Placement is a pure function of the key name: CRC16 of the key mod 16384 gives a hash slot, and a slot is owned by exactly one master. Resharding redistributes slots between nodes, but a single key never spans slots, so its traffic never spans nodes. You have to change the key names — split it into several keys — before the cluster can spread the load.
- How would a hot key show up in your monitoring before anyone reports an outage?Look for skew rather than absolute numbers: one cluster node with much higher CPU, ops/sec or network egress than its identically-sized peers. Client-side per-key or per-prefix latency histograms make it obvious, since p99 rises for everything on that shard while other shards are flat.
A supermarket can open more checkout lanes, but if every customer wants the one item at the back of aisle three, the extra lanes do not help — the queue forms at the shelf, not at the till.
saying these in an interview costs you the question
- Claiming Redis Cluster automatically detects and rebalances hot keys — it rebalances slots, and only when you tell it to
- Saying the fix is to add more nodes or more memory
- Assuming replicas solve hot writes; every write still goes through the single master
- Arguing that because GET is O(1) a single key cannot be a bottleneck — O(1) per operation says nothing about operations per second on one thread
- Confusing a hot key with a big key and jumping straight to `--bigkeys`