skip to content

A single Redis hash has grown to several million fields and application code calls HGETALL on it on every request. What goes wrong, and how would you read, shrink and eventually restructure that key?

level: seniorimportance: should knowfreq 34%

answer

  1. HGETALL O(N) = server-wide stall + huge output buffer
  2. HMGET for known fields; HSCAN COUNT for iteration
  3. HSCAN: at-least-once, duplicates possible, no order
  4. DEL is O(N) too → UNLINK / lazyfree
  5. one key = one slot = one shard → shard fields into N sub-keys

basics

~20 s

HGETALL is O(fields) and runs as one command, so it stalls the server for everyone and builds a huge reply in the client output buffer. Read with HMGET for known fields or HSCAN in chunks, delete with UNLINK or batched HDEL, and split the object across several keys so no single key is that large.

solid answer

~60 s

**What breaks.** `HGETALL` is O(number of fields) and executes as a single command, so a multi-million-field hash occupies the server for the whole scan while every other client waits. The full reply is also materialised into the client output buffer — hundreds of megabytes of RAM and network per call — which can trip output-buffer limits and disconnect the client. **Read it safely.** Use `HMGET` when you know the fields you want. Use `HSCAN key cursor COUNT 500` to iterate in bounded chunks; it may return duplicates and gives no ordering, but every field present for the whole iteration is returned at least once. `HLEN` and `HSTRLEN` answer size questions without transferring data, and `redis-cli --bigkeys` / `--memkeys` finds keys like this before users do. **Shrink it.** `DEL` on a huge key blocks too — use `UNLINK` to free it in a background thread, or delete fields in batches with `HDEL`. **Restructure.** A key is one slot on one shard, so this object cannot be spread out or rebalanced. Split it into N sub-keys by hashing the field name (`obj:{id}:0…N`) and route reads to the right sub-key.

code

text · 16 lines
text
HLEN big:hash                       # (integer) 4211903 - O(1), no transfer

# only what you need
HMGET big:hash user:42 user:99

# full iteration in bounded chunks (cursor ends at 0; duplicates possible)
HSCAN big:hash 0 COUNT 500
HSCAN big:hash 3712 COUNT 500

# deletion: DEL is O(N) and blocks; UNLINK frees in a background thread
UNLINK big:hash

# find keys like this before users do
redis-cli --bigkeys
redis-cli --memkeys
MEMORY USAGE big:hash

go deeper

for a junior

Know that HGETALL reads every field, that its cost grows with the hash, and that HMGET or HSCAN should be used instead.

for a middle

Explain the three costs — server occupancy, output-buffer memory, client deserialisation — and state HSCAN's cursor guarantees and COUNT semantics accurately.

for a senior

Add detection (--bigkeys, MEMORY USAGE, SLOWLOG), safe deletion with UNLINK and lazyfree settings, batched HDEL trimming, and the single-slot cluster consequence with a field-sharding scheme.

for a principal

Treat an unbounded collection under one key as a modelling defect: decide whether it is really an index, whether entries need lifetimes, and set the size ceiling and alerting that prevent any single key from becoming a shard's limit.

## Why a huge hash is an operational problem Redis executes commands one at a time. A command's cost is therefore not just that caller's latency — it is a stall for every other client on that shard. `HGETALL` is O(N) in the number of fields, so on a hash with millions of entries a single call can occupy the server for tens or hundreds of milliseconds. Every other request queues behind it, and the symptom shows up as unexplained latency spikes on unrelated keys. Second, the reply must be **materialised**. The whole field/value set is serialised into the client's output buffer before it drains over the network. That is a large, sudden allocation; with several clients doing it concurrently, memory rises sharply, and if a client is slow to read, the connection can hit `client-output-buffer-limit` and be killed — after the server has already paid the full CPU cost. Third, the network and the client pay again: the payload crosses the wire and is then deserialised into application objects, usually to have two fields read from it. ## Reading safely - **`HMGET key f1 f2 f3`** — the first fix. If the request needs three attributes, fetch three attributes. Same single round trip, bounded cost. - **`HSCAN key cursor [MATCH pat] [COUNT n]`** — cursor iteration for when you genuinely need everything (an export, a migration, a rebuild). Each call returns a chunk and the next cursor; iteration ends when the cursor comes back `0`. The guarantees are worth stating precisely: every field present from start to end of the iteration is returned **at least once**; fields added or removed during iteration may or may not appear; **duplicates are possible**, so the consumer must be idempotent; and there is **no ordering**. `COUNT` is a hint about work per call, not an exact page size — raise it to reduce round trips, but not so far that one call becomes the very stall you were avoiding. Note also that `MATCH` filters *after* fetching, so a restrictive pattern can return empty chunks while still doing full work. - **`HLEN`** (O(1)) and **`HSTRLEN`** answer "how big is it" without transferring anything. ## Finding these keys before they hurt `redis-cli --bigkeys` samples the keyspace and reports the largest key per type; `--memkeys` reports by memory. `MEMORY USAGE key` gives a per-key estimate. `SLOWLOG GET` will already be full of `HGETALL` entries if this is happening in production. A useful habit is to alert on `HLEN` for known collection keys, since collections that grow unboundedly are almost always a modelling mistake that is cheap to catch early. ## Shrinking and deleting Deleting a huge key is itself an O(N) operation: `DEL` frees every field synchronously and stalls the server just like `HGETALL` did. Use **`UNLINK`**, which unlinks the key from the keyspace immediately and reclaims the memory in a background thread — the caller sees O(1). The `lazyfree-lazy-user-del` configuration makes `DEL` behave like `UNLINK` by default, and related `lazyfree-*` settings do the same for eviction and expiry. If you only need to trim, delete in batches: iterate with `HSCAN` and issue `HDEL` for a few hundred fields at a time, pacing the loop so the server keeps serving traffic. ## Restructuring: the real fix A key is the unit of distribution. In Redis Cluster a key belongs to exactly one hash slot on one shard, and resharding moves whole keys, so a single multi-million-field hash: - cannot be spread across shards, no matter how many you add; - concentrates all reads and writes for that object on one node; - makes that node's memory and CPU the ceiling for the whole object; - makes slot migration painful, since the key moves as one indivisible blob. The standard remedy is **sharding the hash by field name**: choose N sub-keys and place a field in `obj:{id}:<hash(field) mod N>`. Reads compute the same function, so a single-field read still costs one command; only full scans need to visit N keys. Choose N so each sub-hash stays in the low thousands of fields. Note that a hash tag like `{id}` keeps all the sub-keys of one object in the same slot — convenient for multi-key operations, but it also means they stay on the same shard, so drop the shared tag if what you need is to spread the load across nodes. Other restructurings worth considering: - **Is it really one object?** A hash with millions of fields is usually an index in disguise — "all events for a tenant", "every user's last-seen". Those want per-entity keys, or a sorted set, or a stream. - **Does it need per-field lifetime?** If entries should age out, per-field TTL (`HEXPIRE`, Redis 7.4+) or window-scoped keys with a whole-key `EXPIRE` remove the growth problem at the source. - **Does the hot path need the whole object at all?** Frequently the answer is a small summary key alongside the large one. ## Encoding, briefly Small hashes use a compact flat representation and convert to a hash table once they exceed the configured entry-count or value-length thresholds. A multi-million-field hash is long past that boundary, so its per-field overhead is the hash-table cost — one more reason the memory footprint of one giant key is worse than the same data spread over many small ones. ## The answer that lands Name all three costs of `HGETALL` (server-wide stall, reply-buffer memory, client deserialisation), give the bounded alternatives (`HMGET`, `HSCAN` with its exact guarantees), remember that deletion is equally O(N) and needs `UNLINK`, and finish on the structural point: one key is one shard, so the durable fix is to stop having a single key that large.

  • What exactly does HSCAN guarantee, and what does it not?
    It guarantees that any field present in the hash for the entire duration of the iteration is returned at least once. It does **not** guarantee that each field is returned only once — duplicates across chunks are normal, so the consumer must be idempotent — nor does it define any ordering, nor say anything about fields added or removed while the scan is running. `COUNT` is a work hint per call, not an exact page size.
  • Why is DEL on a huge hash also a problem, and what do you use instead?
    `DEL` frees every field synchronously, so it is O(N) and stalls the single command-execution thread exactly as `HGETALL` does. `UNLINK` removes the key from the keyspace immediately and reclaims memory in a background thread, so the caller sees constant time. Setting `lazyfree-lazy-user-del yes` makes `DEL` behave that way by default, and the related `lazyfree-*` options apply the same treatment to eviction and expiry.
  • In Redis Cluster, can you spread one enormous hash across shards?
    No. A key maps to a single hash slot and resharding relocates whole keys, so the entire hash lives on one shard regardless of cluster size. Spreading the load requires splitting the data into multiple keys — typically N sub-hashes chosen by hashing the field name — and deliberately *not* giving them a common hash tag, since a shared tag would pin them all back into the same slot.

Asking for the whole hash is like making one clerk read out a million-line ledger while everyone else in the queue waits: reading only the lines you need, or a page at a time, keeps the queue moving.

saying these in an interview costs you the question

  • Assuming HGETALL is cheap because Redis is in memory
  • Believing a slow HGETALL only affects the caller rather than every client on that shard
  • Thinking Redis Cluster will spread a single large key across shards
  • Using DEL on a multi-million-field key and expecting it to be instant
  • Treating HSCAN as returning each field exactly once, or in a stable order

context