skip to content

How would you find out which keys are receiving a disproportionate share of traffic on a production Redis instance, and what does `redis-cli --hotkeys` require in order to work at all?

level: seniorimportance: must knowfreq 45%

answer

  1. --hotkeys = SCAN + OBJECT FREQ
  2. Needs allkeys-lfu / volatile-lfu
  3. LFU counter: 8-bit, logarithmic, decays — rank not rate
  4. MONITOR = exact truth, expensive, seconds only
  5. Client-side prefix sampling is the durable answer

basics

~20 s

redis-cli --hotkeys scans the keyspace and reads each key's LFU access counter, so it only works when maxmemory-policy is an LFU policy (allkeys-lfu or volatile-lfu). Otherwise: OBJECT FREQ on suspects, brief MONITOR sampling, or counting keys in the application client.

solid answer

~60 s

Start by localising: `INFO stats` / `INFO commandstats` per node and CPU metrics tell you *which shard* is skewed. Then `redis-cli --hotkeys`. It works by `SCAN`ning the whole keyspace and calling `OBJECT FREQ` on each key, so it **requires an LFU eviction policy** (`allkeys-lfu` or `volatile-lfu`) — with LRU or `noeviction` it refuses to run. Note two caveats: it is a full keyspace scan (real CPU cost on a big instance, run it against a replica), and `OBJECT FREQ` returns a decayed logarithmic counter, not requests per second — it tells you relative popularity, not a rate. `MONITOR` gives exact per-key traffic but streams every command to your client and can cost a large slice of throughput; use it for a few seconds at most, or not at all on a hot node. The most reliable option at scale is **client-side sampling**: count key prefixes in your cache wrapper on 1% of calls and export a top-N metric. It measures the real thing, costs nothing on the server, and survives cluster topology changes.

go deeper

for a junior

Know that redis-cli --hotkeys exists and that it needs an LFU maxmemory-policy.

for a middle

Explain how it works (SCAN plus OBJECT FREQ) and why the LFU counter is approximate and decaying.

for a senior

Sequence a real investigation: node-level skew first, then --hotkeys on a replica, then targeted OBJECT FREQ/MEMORY USAGE, with MONITOR as a short, deliberate last resort.

for a principal

Argue for standing instrumentation — sampled per-prefix metrics in the cache client — so hot keys are visible before they cause an incident, and weigh the cost of switching eviction policy to LFU.

## The problem Redis does not keep per-key request counters. `INFO` gives you totals and per-command statistics, `SLOWLOG` gives you individually slow commands — but a hot key is usually a flood of individually fast commands, so nothing in the default telemetry points at a key name. You need a deliberate technique. ## Step 1 — localise the shard Before hunting for a key, find the node. Compare across cluster nodes: - `INFO stats` → `instantaneous_ops_per_sec`, `total_net_output_bytes` - `INFO commandstats` → calls and usec per command - host CPU Identically-sized nodes with wildly different ops/sec or egress means traffic skew. If all nodes are even, your problem is volume, not a hot key. ## Step 2 — `redis-cli --hotkeys` This built-in mode (Redis 4.0+, alongside `--bigkeys` and `--memkeys`) works like this: 1. It `SCAN`s the entire keyspace in batches (cursor-based, non-blocking, so it does not freeze the server). 2. For each key it calls `OBJECT FREQ <key>`. 3. It keeps a top-N list by counter value and prints it with a summary. **Hard requirement:** `OBJECT FREQ` is only meaningful under LFU, so `maxmemory-policy` must be `allkeys-lfu` or `volatile-lfu`. Under LRU or `noeviction`, `--hotkeys` errors out and tells you so. Switching policy is `CONFIG SET maxmemory-policy allkeys-lfu` — a live change, but it also changes what gets evicted under pressure, so treat it as a real change and revert deliberately. **What the counter means.** LFU stores an 8-bit *approximate, logarithmic* counter per object. It saturates around 255, increments probabilistically (governed by `lfu-log-factor`, so a key at 10x the traffic does not have 10x the counter), and **decays over time** (`lfu-decay-time`, in minutes). So the output ranks keys by recent-ish popularity; it is not a request rate and it cannot tell you "3,000 rps". A key hammered five minutes ago and idle since may still rank highly. **Cost.** It is an O(N) walk of the keyspace with one round trip per key (plus the network). On tens of millions of keys this takes a long time and adds meaningful load. Run it against a replica when you can — replicas apply the same writes, and their LFU counters reflect their own read traffic plus replicated writes, so read-hot detection on a replica reflects the replica's traffic. If reads are served only by the master, `--hotkeys` on the master is the honest place to look; do it during a quiet window or accept the load. ## Step 3 — targeted checks Once you have suspects, `OBJECT FREQ key` (LFU) or `OBJECT IDLETIME key` (LRU — seconds since last access, useful in the opposite direction: proving a key is *cold*) confirm them cheaply. `MEMORY USAGE key` and `OBJECT ENCODING key` tell you whether the key is also big, which changes the fix. ## Step 4 — `MONITOR`, carefully `MONITOR` streams every command the server processes to your connection. It gives exact, real-time, per-key truth — and it is expensive: the server formats and ships a line per command, which can cut throughput substantially, and if your client cannot keep up, the output buffer grows against `client-output-buffer-limit`. Rules of practice: never leave it running, never point it at an already-saturated node, pipe a few seconds into a file and aggregate offline, and prefer a replica. Something like a few seconds of capture piped through sort/uniq on the key field will identify a dominant key immediately. ## Step 5 — client-side sampling (the durable answer) The technique that actually scales in production is instrumenting your own cache client: - wrap `get`/`set` and increment a counter keyed by **key prefix** (`product:*`, `user:*`) always, and by full key on a sampled fraction (1 in 100) to bound cardinality; - export a top-N gauge, or push samples to your metrics/tracing backend; - you now have per-key rates, over time, per service, with no server cost, and it keeps working across failovers and resharding. This is also the only method that sees traffic your Redis never sees — for example requests already absorbed by a near cache. ## Redis Enterprise / managed variants Some managed offerings ship a hot-key or per-slot traffic view. Treat that as a bonus, not a substitute: the vendor-independent stack is skew metrics → `--hotkeys` → sampled client instrumentation. ## Anti-patterns - Running `KEYS *` to "look around" — O(N) *and* blocking; use `SCAN`. - Reading raw `INFO keyspace` and assuming the biggest database is the busiest one; key count is unrelated to traffic. - Concluding from a one-off `--hotkeys` run that a key is hot without checking rate over time — LFU counters decay and a batch job can leave a false trail.

  • Why is `MONITOR` dangerous on a busy production node?
    The server must format and send a line for every command it executes, on top of doing the work, which can consume a large share of throughput. If the monitoring client reads slower than the server writes, the client output buffer grows and can hit `client-output-buffer-limit`, causing disconnects. Use it for a few seconds, preferably on a replica, and aggregate the capture offline.
  • What exactly does the number returned by `OBJECT FREQ` represent?
    It is the object's LFU counter: an 8-bit approximate frequency that increments probabilistically on access — controlled by `lfu-log-factor`, so growth is logarithmic rather than linear — and decays over time according to `lfu-decay-time`. It ranks keys against each other; it is not a request count and cannot be converted to requests per second.
  • You must diagnose a suspected hot key but cannot change `maxmemory-policy` to LFU. What do you do?
    Fall back to evidence that does not need LFU: per-node skew from `INFO stats` and CPU, a few seconds of `MONITOR` captured on a replica and aggregated offline, and — best — sampled per-prefix counters added to the application's cache client. `OBJECT IDLETIME` still works under LRU but only proves coldness, not heat.

saying these in an interview costs you the question

  • Believing `--hotkeys` works under any eviction policy — it needs LFU and refuses otherwise
  • Reading `OBJECT FREQ` as a request count or a rate
  • Using `KEYS *` to enumerate the keyspace instead of `SCAN`
  • Leaving `MONITOR` running on a saturated production master
  • Confusing `--bigkeys` (value size) with `--hotkeys` (access frequency)

context