skip to content

Hot Keys

You will learn how a single celebrity key can saturate one shard while the rest of the cluster idles, and the mitigations: detection, key splitting, replicated copies, and local caches. Interviewers use hot keys to test whether you understand that sharding does not help skewed access.

part ofRedisoverview, primer and where to startread it →
on this pageshow

questions

5

In a Redis-backed cache, what is a "hot key", and why can a single popular key become a bottleneck even though a Redis node handles hundreds of thousands of operations per second?

level: middleimportance: must knowfreq 55%

answer

  1. One key → one slot → one node → one thread
  2. Bandwidth wall before CPU wall
  3. Adding nodes moves slots, not a single key
  4. Head-of-line blocking hurts the whole shard
  5. Hot ≠ big, but hot + big is worst

basics

~20 s

A hot key is one key taking a hugely disproportionate share of traffic. Because a key's name determines exactly one shard, and that shard executes commands on one thread, the load cannot be spread by adding nodes — one key saturates one CPU core and one network link.

solid answer

~50 s

A hot key is a single key whose request rate is orders of magnitude above the average key — a celebrity profile, a homepage config blob, a global counter, a feature flag. It hurts for two reasons. First, placement is deterministic: in Redis Cluster the key name is hashed (CRC16 mod 16384) to one slot, and that slot lives on exactly one master. Adding nodes reshards *other* keys away; the hot one stays put. Second, that node executes commands on a single thread, so all traffic for the key queues behind one core, and every other key on that node queues behind it too — the blast radius is the whole shard, not just the key. The usual limits hit in this order: outbound network bandwidth (value size × rps), then CPU on the command thread, then client connection pools timing out. Mitigation is detection first (`redis-cli --hotkeys`, client-side sampling), then splitting the key across slots, replica reads, or a short-TTL in-process cache.

go deeper

for a junior

Be able to say what a hot key is with an example (a celebrity record, a global config blob) and that all its traffic lands on one node.

for a middle

Explain the key → CRC16 slot → single master → single command thread chain, and name bandwidth and CPU as the two limits.

for a senior

Add diagnosis (skew across nodes, head-of-line blocking on the shard, retry amplification) and pick a mitigation with its cost, distinguishing read-hot from write-hot.

for a principal

Frame it as a data-placement and traffic-distribution problem: which keys are allowed to be global, what staleness the product can buy relief with, and whether the hot entity belongs in Redis at all.

## What "hot" means Real cache traffic is never uniform. Access frequency follows a heavy-tailed (Zipf-like) distribution: a handful of keys absorb a large fraction of all requests. A **hot key** is a key on that extreme tail — think the cached record for a celebrity account, a global `config:features` blob every request reads, a leaderboard, a rate-limit counter for a shared tenant, or the cached front page. Hot is about **request rate**, not size. A separate problem is a **big key** — one key holding a huge value or a multi-million-element collection. They compound badly: a big key that is also hot multiplies bytes per second, and collection commands over a big key (`HGETALL`, `LRANGE 0 -1`, `SMEMBERS`) are O(N), so each hit costs real CPU rather than a pointer dereference. ## Why sharding does not save you Redis Cluster places data by hashing the key name: `CRC16(key) mod 16384` gives a hash slot, and each master owns a contiguous set of slots. This is deterministic and stateless — every client computes the same answer. That is what makes the cluster fast and coordination-free, and it is exactly why a hot key cannot be load-balanced. `product:42` always maps to the same slot and therefore to the same master. Resharding moves *slots*; it can move the hot key to a quieter node, but it cannot split a single key's traffic, because a key lives in exactly one slot. So the standard horizontal-scaling reflex — add nodes — buys nothing for the hot key itself. It only helps if the node was also busy with other work you can move away. ## Why one key can saturate a node Commands on a Redis node are executed by one thread (I/O threading added in 6.0 parallelizes socket reads/writes, not command execution). Practical consequences: - **Bandwidth first.** This is usually the wall you hit before CPU. A 10 KB cached JSON blob served 50,000 times a second is 4 Gbit/s of egress from one process — more than a 1 GbE link, and enough to make even a 10 GbE instance uncomfortable once replication traffic is added. Cloud instances also cap network per instance size. - **CPU next.** At small values Redis will do a few hundred thousand simple GETs per second per core. If each hit is an O(N) collection read or a Lua script, that ceiling drops by an order of magnitude. - **Head-of-line blocking.** Because one thread serves everything, the hot key's queue delays *every other key on that shard*. Your p99 for unrelated lookups degrades, which is why hot keys usually surface as "random unrelated timeouts on one node". - **Client-side amplification.** When latency rises, connection pools fill, callers time out and retry, and retries add load — a small skew turns into a cliff. ## Symptoms you will actually see One node in the cluster shows high `instantaneous_ops_per_sec` and CPU while its peers idle; `INFO stats` shows lopsided `total_net_output_bytes`; SLOWLOG may be empty (no single command is slow — there are just too many of them); client-side p99 on that shard's keys climbs while p50 stays fine. Uneven CPU across identically-sized cluster nodes is the classic tell. ## The mitigation menu 1. **Detect and quantify** before acting — `redis-cli --hotkeys` (needs an LFU eviction policy), `OBJECT FREQ` on suspects, or sampling in the cache client. Guessing which key is hot is how teams optimize the wrong thing. 2. **Split the key** into N copies with different suffixes so they hash to different slots and different masters; readers pick one at random or by client id. 3. **Read replicas** — each replica is a separate process with its own core and NIC, so N replicas multiply read capacity. Helps reads only, and reads become slightly stale. 4. **In-process (near) cache** with a very short TTL. The most effective by far for read-hot keys: a 1-second local TTL collapses any per-instance read rate to one Redis GET per second per instance. It costs staleness and per-instance divergence. 5. **Shrink the value** — hot plus large is the worst combination, so store only the fields callers need, or compress, before doing anything architectural. A hot **write** key is harder: replicas and near caches do nothing for it, since every write must reach the one master. There you either split the key (e.g. N counter shards summed on read), batch/aggregate writes in the application, or move the counter out of the critical path.

  • How does a hot key differ from a big key, and can one key be both?
    Hot is about request rate; big is about the size of the value or the element count of the collection. A big key is dangerous because O(N) commands over it occupy the command thread for milliseconds, blocking everything else. A key that is both hot and big is the worst case: bytes per second and CPU per hit multiply, and it is usually the first thing to fix by shrinking the value.
  • Why doesn't adding more nodes to Redis Cluster fix a hot key?
    Placement is a pure function of the key name: CRC16 of the key mod 16384 gives a hash slot, and a slot is owned by exactly one master. Resharding redistributes slots between nodes, but a single key never spans slots, so its traffic never spans nodes. You have to change the key names — split it into several keys — before the cluster can spread the load.
  • How would a hot key show up in your monitoring before anyone reports an outage?
    Look for skew rather than absolute numbers: one cluster node with much higher CPU, ops/sec or network egress than its identically-sized peers. Client-side per-key or per-prefix latency histograms make it obvious, since p99 rises for everything on that shard while other shards are flat.

A supermarket can open more checkout lanes, but if every customer wants the one item at the back of aisle three, the extra lanes do not help — the queue forms at the shelf, not at the till.

saying these in an interview costs you the question

  • Claiming Redis Cluster automatically detects and rebalances hot keys — it rebalances slots, and only when you tell it to
  • Saying the fix is to add more nodes or more memory
  • Assuming replicas solve hot writes; every write still goes through the single master
  • Arguing that because GET is O(1) a single key cannot be a bottleneck — O(1) per operation says nothing about operations per second on one thread
  • Confusing a hot key with a big key and jumping straight to `--bigkeys`

context

open as a page

How would you find out which keys are receiving a disproportionate share of traffic on a production Redis instance, and what does `redis-cli --hotkeys` require in order to work at all?

level: seniorimportance: must knowfreq 45%

basics

~20 s

redis-cli --hotkeys scans the keyspace and reads each key's LFU access counter, so it only works when maxmemory-policy is an LFU policy (allkeys-lfu or volatile-lfu). Otherwise: OBJECT FREQ on suspects, brief MONITOR sampling, or counting keys in the application client.

open as a page

One Redis key is read about 50,000 times per second while being rewritten only occasionally, and it is saturating the node that owns it. You are considering adding read replicas of that Redis master so reads can be served from them. Work through the capacity arithmetic: what does each added replica actually buy you for that one key, what would change if the same key were write-hot instead of read-hot, and how does the cost per unit of relief compare with splitting the entry into several suffixed Redis keys or putting a sub-second in-process cache in front of Redis?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Each replica is a separate process with its own core and network link, so N replicas roughly multiply read capacity for that one key. Check egress first: value size times reads per second. A write-hot key gains nothing and adds master work. A near cache is far cheaper.

open as a page

Explain how splitting one heavily-read cache entry into several Redis keys with different suffixes (for example `product:123:0` through `product:123:9`) relieves load, and what that technique costs you.

level: seniorimportance: should knowfreq 35%

basics

~20 s

You store N identical copies under different key names. Because the names hash to different slots, the copies land on different cluster nodes, and each reader picks one at random — so read traffic divides by N. You pay N times the memory, an N-key write fan-out that is not atomic, and a window where copies disagree.

open as a page

You decide to put a short-lived in-process cache in front of Redis for a handful of extremely hot cache entries. How do you choose the TTL and the scope, and what failure modes does that introduce?

level: principalimportance: should knowfreq 30%

basics

~20 s

Cache only measured-hot, read-mostly entries in a bounded in-process map with a TTL of about a second. That collapses each instance's Redis traffic for the key to one GET per second, at the price of staleness up to local TTL plus Redis TTL, and instances briefly disagreeing with each other.

open as a page