skip to content

Explain what actually goes wrong in a production Redis deployment when one key holds tens of millions of elements or a value of several hundred megabytes, and how you would remodel that data.

level: seniorimportance: must knowfreq 50%

answer

  1. One thread: O(N) key = global stall
  2. Free/DEL is O(N) too - use UNLINK
  3. Output buffer overrun -> replica full resync
  4. One key = one slot -> hot, unmovable shard
  5. Shard the key; SCAN in batches

basics

~20 s

Redis serves commands on one thread, so any whole-key operation on a giant key - reading it, deleting it, migrating it, replicating it - stalls every other client. It also unbalances shards and cannot be split. Fix by sharding the key and reading it in cursor-based chunks.

solid answer

~60 s

A big key is a latency bomb because of the single-threaded execution model. - **Command cost.** `HGETALL`, `SMEMBERS`, `LRANGE 0 -1` on ten million elements are O(N) on the serving thread; every other client waits. So is freeing it synchronously. - **Network and buffers.** A 300 MB reply must be built in a client output buffer and flushed; it can hit `client-output-buffer-limit` and get the client (or a replica) disconnected, which for a replica means a **full resynchronization** — potentially in a loop. - **Replication and persistence.** One write to a big key can push a large payload downstream, and the key is copied wholesale into snapshots and rewrites. - **Cluster.** A key lives in exactly one hash slot, so it cannot be split across shards: it skews memory, creates a hot shard, and slot migration of it is a blocking transfer that can time out. Remodel: shard by a hash of the field into `key:{id}:part:N`, iterate with `HSCAN`/`SSCAN`/`ZRANGE` limits instead of whole-key reads, cap collection sizes, and delete with `UNLINK`.

code

text · 14 lines
text
# before: everything in one key, read whole
HSET user:events:123 <field> <value>      # ends up with 20M fields
HGETALL user:events:123                   # O(N) stall for every client

# after: N physical shards, bounded reads
HSET user:events:123:{7} <field> <value>  # shard = hash(field) % 16
HSCAN user:events:123:{7} 0 COUNT 200     # cursor-based, small batches

# safe removal, and the settings that cover expiry/eviction too
UNLINK user:events:123:{7}
# redis.conf
lazyfree-lazy-expire yes
lazyfree-lazy-eviction yes
lazyfree-lazy-server-del yes

go deeper

for a junior

Say why it matters at all: Redis handles one command at a time, so any operation touching millions of elements blocks everyone else.

for a middle

List the concrete costs — O(N) reads and frees, oversized replies and buffers, snapshot copying — and the SCAN-based alternatives.

for a senior

Add the cluster and replication consequences (hot unmovable slot, buffer overrun causing full resync) and give a remodelling plan with UNLINK and lazy-free settings.

for a principal

Make it a standard: explicit size budgets in the data model, automated big-key detection, and treating a breach as a design defect rather than a hardware problem.

## Why one key can hurt an entire cluster Redis processes commands one at a time on a single thread. Latency for every client is therefore the sum of the work in front of them. A key that is large in *elements* or in *bytes* turns ordinary-looking commands into long units of work that nothing can preempt — there is no per-command time slice and no yielding halfway through. ## The failure modes **O(N) commands on the serving thread.** `SMEMBERS` on a 20-million-member set, `HGETALL` on a huge hash, `LRANGE 0 -1`, `ZRANGE key 0 -1` — each walks the whole structure. At tens of millions of elements this is tens or hundreds of milliseconds of pure stall. Every other client's p99, including unrelated GETs on tiny keys, inherits it. **Freeing is also O(N).** A synchronous `DEL` of a collection must free every element. Deleting the monster to "fix" the problem can cause the very stall you were trying to avoid; the same applies when it expires or is evicted, unless the lazy-free settings move the work off the main thread. **Output buffers and disconnects.** Replies are assembled into per-client output buffers, which count toward `used_memory` and are bounded by `client-output-buffer-limit`. A few hundred megabytes of reply can breach the limit and get the client killed; worse, replicas have their own buffer class, and a big write that overruns a replica's buffer forces a disconnect and a **full resync**, which costs a snapshot, a large transfer, and can repeat if the same write pattern recurs. **Persistence and replication amplification.** Big values are copied whole into snapshots and rewrites. A frequently-modified large key also churns pages that the persistence child shares with the parent, inflating resident memory during background saves. **Cluster-level damage.** A key always maps to exactly one hash slot and therefore one shard. Consequences: the shard holding it uses disproportionate memory, so the cluster cannot be balanced by moving slots; the traffic to that key concentrates on one node, which is the classic hot-shard; and if you do try to move that slot, the migration transfers the key in a blocking operation that can exceed the configured timeout and leave the slot in a migrating state needing manual attention. **Diagnosis is indirect.** These stalls show up as an unexplained tail-latency profile. `SLOWLOG GET` will name the offending command when it is a user command — that is your best evidence — while `LATENCY LATEST` and `LATENCY DOCTOR` catch events that are not commands. Find the keys themselves with `redis-cli --memkeys` (bytes) or `--bigkeys` (element counts), preferably against a replica. ## Remodelling **Shard the key.** Split a logical collection across N physical keys by hashing the element or field: `user:events:{123}:0` … `:15`. Reads that need one element go to one shard; reads that need everything are now N bounded commands you can pipeline, and each is short enough not to stall the loop. In Cluster mode use hash tags deliberately — `{123}` keeps the parts co-located when you need them together, and *omitting* a shared tag is what lets the parts spread across shards when you want the memory and traffic distributed. **Stop reading whole keys.** Replace `HGETALL`/`SMEMBERS` with cursor-based `HSCAN`/`SSCAN` in batches, and range reads with bounded `ZRANGE`/`LRANGE` windows or `ZRANGEBYSCORE ... LIMIT`. This converts one enormous unit of work into many small ones, which is exactly what a single-threaded server needs. **Cap growth by design.** Trim lists and streams (`LTRIM`, `XADD ... MAXLEN ~`), expire per-item keys instead of accumulating into one collection, and put an explicit size ceiling in the code path that writes. **Delete safely.** Use `UNLINK` rather than `DEL`, and enable the lazy-free settings so expiry and eviction of large values also happen off the main thread. **Guardrails.** Track the largest keys periodically (a scheduled `--memkeys` on a replica), alert when any key exceeds an agreed budget — a common working rule is to keep values under about a megabyte and collections under a few tens of thousands of elements — and treat a breach as a design defect rather than a capacity issue.

  • In Redis Cluster, why can't the cluster just spread a very large key across several shards?
    Placement is per key: the key name (or its hash tag) maps to one of the 16384 hash slots, and a slot lives on exactly one shard. Nothing splits a single key's value across nodes. That is why the only cure is application-level sharding into multiple key names, and why a big key both skews memory and blocks slot migration, since moving it is a single blocking transfer that can time out.
  • How would you detect big keys before they cause an incident?
    Run redis-cli --memkeys, and --bigkeys for element counts, on a replica on a schedule, with a sleep interval so the scan adds little load. Record the largest key per type over time and alert on a budget, for example any value over a megabyte or any collection over a few tens of thousands of elements. Correlate with SLOWLOG entries naming whole-key commands, which is usually the first symptom in production.

A single-lane checkout: one customer with a trolley of ten thousand items does not just slow themselves down, they hold up everyone in the queue, and you cannot split their trolley across two tills.

saying these in an interview costs you the question

  • Assuming a big key only affects the client that reads it
  • Reaching for DEL to clean up a huge collection on a busy primary
  • Believing a cluster will rebalance a single oversized key across shards
  • Judging key size only by element count and ignoring bytes, or vice versa
  • Blaming network or client libraries for the tail latency a whole-key read causes

context