skip to content

A Redis instance holds 50 million tiny key-value pairs and uses far more memory than the sum of the values suggests. How would you investigate, and what restructuring would let the encoding layer cut that footprint?

level: seniorimportance: should knowfreq 36%

answer

  1. top-level key tax ≈ 50-100 bytes, payload-independent
  2. MEMORY USAGE / OBJECT ENCODING / --memkeys on a replica
  3. bucket ids into hashes under hash-max-listpack-entries
  4. one oversized field kills the whole bucket's encoding
  5. lost: per-record TTL, eviction granularity, slot spread

basics

~20 s

Each top-level key carries its own object, SDS and dict-entry overhead — often 50-100 bytes regardless of value size. Group the records into hashes bucketed so each stays under hash-max-listpack-entries/value, so many records share one compact listpack. Measure with MEMORY USAGE, OBJECT ENCODING and MEMORY DOCTOR.

solid answer

~60 s

**Investigate first.** `INFO memory` for `used_memory` versus `used_memory_rss` and the fragmentation ratio; `MEMORY USAGE <key>` on samples to see the real per-key cost including overhead; `OBJECT ENCODING` to check which layout each value uses; `redis-cli --memkeys`/`--bigkeys` (against a replica) for the distribution; `MEMORY DOCTOR` for obvious pathologies. **The likely cause.** A top-level key is not free. It costs a dict entry in the keyspace, a robj header, an SDS header for the key name and another for the value, plus allocator rounding — commonly 50–100 bytes before any payload. With 50 M keys holding 8-byte values, overhead is the dataset. **The restructuring.** Move the records inside hashes, bucketed by a deterministic function of the id (`user:{id/1000}` → field `id`). Each bucket holds ~1000 records under one top-level key, and if you size buckets to stay under `hash-max-listpack-entries` (raise it to, say, 512–1000 and keep values short), the bucket is one compact listpack with near-zero per-field overhead. Reductions of several times are typical. **The costs**: no per-record TTL, field lookups become a scan of the listpack, and multi-key atomicity changes shape.

code

text · 16 lines
text
# per-key cost of the flat layout
> SET user:12345 "a"
> MEMORY USAGE user:12345
(integer) 88          # ~88 bytes for a 1-byte value

# bucketed layout, kept under the listpack threshold
> CONFIG SET hash-max-listpack-entries 1000
> CONFIG SET hash-max-listpack-value 64
> HSET user:12 345 "a"
> OBJECT ENCODING user:12
"listpack"
> MEMORY USAGE user:12       # amortised across ~1000 fields

# instance-level view
> INFO memory | grep -E 'used_memory:|used_memory_rss:|fragmentation'
> MEMORY DOCTOR

go deeper

for a junior

Know that each Redis key has fixed overhead and that grouping small records into hashes saves memory.

for a middle

Name the measurement commands and the bucketing pattern, and state the listpack thresholds it depends on.

for a senior

Run the full investigation — fragmentation versus dataset, per-key cost, encoding checks on a replica — and articulate the CPU, TTL and eviction costs of bucketing.

for a principal

Sequence the levers by cost and blast radius, decide whether the data model change is justified against simply provisioning more memory, and set the bucket-size and threshold policy the team will follow.

## Step 1 — establish where the memory actually goes Before restructuring anything, get numbers: - `INFO memory` — `used_memory` (what Redis accounts for), `used_memory_rss` (what the OS gave the process), `mem_fragmentation_ratio`, and `used_memory_dataset` versus overhead. A high fragmentation ratio points at the allocator, not at your key design, and has a different fix (`activedefrag`, restart, jemalloc tuning). - `DBSIZE` — key count. Divide `used_memory_dataset` by it for the average cost per key; compare that with the average payload you expect. A 10× gap is the signature of per-key overhead. - `MEMORY USAGE <key>` on a sample — the actual bytes for that key including its overhead. Running it on a key with an 8-byte value and seeing ~90 bytes back makes the problem concrete. - `OBJECT ENCODING <key>` — confirms whether values are in compact or general encodings. - `redis-cli --memkeys` (or `--bigkeys`) against a **replica**, since it scans the keyspace and you do not want that load on a master. - `MEMORY DOCTOR` and `MEMORY STATS` for a summary including per-database overhead and the size of client output buffers, replication backlog and AOF buffers — occasionally the "missing" memory is there rather than in the dataset. ## Step 2 — understand the per-key tax Every top-level key in Redis carries fixed costs regardless of how small its value is: - an entry in the global keyspace dict (key pointer, value pointer, next pointer), - a `robj` header for the value (type, encoding, LRU/LFU field, refcount), - an SDS header for the key name and, for string values, another for the value, - allocator size-class rounding on each of those allocations, - plus, if the key has a TTL, an entry in the expires dict as well. Realistically that lands in the 50–100 byte range per key. It is invisible in a benchmark with a thousand keys and completely dominant at fifty million. ## Step 3 — the restructuring that works The fix is to stop having 50 million top-level keys. Group records into hashes so that the fixed per-key tax is paid once per *bucket* instead of once per record, and the records inside enjoy the listpack encoding's near-zero per-entry overhead: ``` # before: 50 000 000 keys SET user:12345 "..." # after: 50 000 keys, 1000 fields each HSET user:12 345 "..." # bucket = id / 1000, field = id ``` For this to pay off, each bucket must stay in the `listpack` encoding, which means: - bucket size comfortably below `hash-max-listpack-entries` — raising it from 128 to something like 512 or 1000 is the usual move, chosen together with the bucket function; - every field name and value below `hash-max-listpack-value` (default 64 bytes) — one oversized value converts the whole bucket to `hashtable` and you lose the saving for all 1000 records in it. Verify empirically rather than trusting the theory: load a representative slice both ways and compare `MEMORY USAGE` totals with `OBJECT ENCODING` confirming `listpack`. ## Step 4 — weigh what you give up 1. **No per-record TTL.** Redis expires whole keys, not hash fields — hash-field TTLs exist only on Redis 7.4+ (`HEXPIRE`), and using them defeats the compact encoding anyway. If each record needs its own expiry, bucketing is usually the wrong answer. 2. **CPU per access.** Reading a field from a 1000-entry listpack is a linear scan of a contiguous blob, and writes may rewrite the blob. Both run on Redis's single command thread, so an over-large bucket converts a memory win into a latency problem. This is exactly the trade the default of 128 encodes; moving it is a deliberate choice to be measured. 3. **Cluster placement.** All fields of a bucket live in one slot on one master, so bucket size influences load distribution and a hot bucket becomes a hot slot that resharding cannot split. 4. **Eviction granularity.** With `maxmemory` set, eviction removes whole keys — a bucket, meaning a thousand records at once. For a cache that changes the miss pattern significantly. ## Step 5 — the other levers, in order of cheapness - **Shorten key names.** With 50 M keys, trimming a prefix from `application:user:profile:` to `u:` saves gigabytes and costs nothing but readability. - **Shrink values.** Store an integer rather than its decimal string (Redis encodes it as `int`), drop JSON field names in favour of a hash or a compact binary encoding, keep strings under 44 bytes so they use `embstr`. - **Check TTL coverage.** Keys without expiry in a cache are a leak; `maxmemory-policy` plus TTLs may be the actual fix. - **Only then** consider bucketing, which changes the data model and therefore the application. The order matters: name and value shrinking are local changes, while bucketing is an architectural one whose downsides (TTL, hot slots, scan cost) must be genuinely acceptable.

  • What breaks if each of those records needs its own TTL?
    Bucketing largely stops working. Redis expires whole keys, so a bucket of a thousand records has a single lifetime; per-field expiry only exists from Redis 7.4 via HEXPIRE, and adding field TTLs adds per-field metadata that undermines the compact encoding you were bucketing for. If per-record expiry is a real requirement, keep separate top-level keys and attack memory another way — shorter key names, smaller values, tighter maxmemory policy, or simply more RAM.
  • You raise hash-max-listpack-entries to 1000 and memory drops, but p99 latency rises. What is happening?
    Listpack access is a linear scan of a contiguous blob and writes may reallocate and memmove it, so a 1000-entry bucket costs roughly eight times the per-operation work of a 128-entry one — all on Redis's single command-execution thread, where it delays every other client. The default of 128 is calibrated to keep that work in the microsecond range. The remedy is to find the bucket size where the memory saving is still most of the win but the scan cost is acceptable, measuring both, rather than maximising bucket size.
  • Before restructuring the data model, what cheaper memory levers would you try?
    Shorten key names, which with tens of millions of keys can save gigabytes for a trivial change; store integers as integers so Redis uses the int encoding rather than a string; keep strings at or under 44 bytes so they use embstr's single allocation; strip redundant JSON field names by moving to a hash or a compact encoding; and verify TTL coverage and maxmemory-policy, since a cache with unexpiring keys is really a leak. These are local changes with no impact on data model, atomicity or eviction behaviour.

Fifty million individually addressed envelopes cost fifty million stamps. Putting a thousand letters into each of fifty thousand parcels pays postage once per parcel — but now you cannot recall a single letter, and someone has to open the parcel to find it.

saying these in an interview costs you the question

  • Assuming memory usage equals the sum of value sizes
  • Bucketing into hashes without checking the resulting OBJECT ENCODING
  • Raising hash-max-listpack-entries arbitrarily with no latency measurement
  • Ignoring that bucketing removes per-record expiry
  • Running --bigkeys or --memkeys against a busy master
  • Blaming fragmentation without checking used_memory versus used_memory_rss

context