A Redis instance is using far more memory than expected. Which built-in commands and redis-cli modes would you use to find out where the memory went, and what does each of them actually measure?
answer
- INFO memory -> STATS -> DOCTOR -> USAGE
- used_memory vs rss vs peak
- bytes-per-key from MEMORY STATS
- --bigkeys = count, --memkeys = bytes
- Never KEYS *; scan a replica
basics
~20 sINFO memory for the totals, MEMORY STATS for the breakdown and bytes-per-key, MEMORY DOCTOR for hints, MEMORY USAGE <key> for one key's bytes, and redis-cli --bigkeys (largest by element count) or --memkeys (largest by bytes) to find offenders. Never KEYS *.
solid answer
~50 sWork top-down. 1. **`INFO memory`** — the totals: `used_memory` (what the allocator has handed out), `used_memory_rss` (what the OS sees), `used_memory_peak`, `maxmemory`, `maxmemory_policy`, `mem_fragmentation_ratio`. This tells you whether the problem is data, overhead, or fragmentation. 2. **`MEMORY STATS`** — the breakdown: dataset bytes versus overhead, `keys.count`, `keys.bytes-per-key`, replication backlog, client output buffers, AOF buffer. Overhead the dataset does not explain shows up here. 3. **`MEMORY DOCTOR`** — heuristic commentary; useful as a sanity check, not as evidence. 4. **`MEMORY USAGE key [SAMPLES n]`** — bytes for one key including its value; for collections it *samples* nested elements (default 5), so it is an estimate unless you pass `SAMPLES 0`. 5. **`redis-cli --bigkeys`** — SCAN-based, reports the largest key per type by **element count**; **`--memkeys`** ranks by **actual bytes** using MEMORY USAGE. Run the scanning modes against a replica when you can, and never use `KEYS *` on a live instance — it blocks the single thread.
code
text · 15 linesredis-cli INFO memory
# used_memory_human:9.83G used_memory_rss_human:12.1G
# used_memory_peak_human:14.2G maxmemory_human:12.00G
# mem_fragmentation_ratio:1.23
redis-cli MEMORY STATS # dataset.bytes, overhead.total,
# keys.count, keys.bytes-per-key,
# replication.backlog, clients.normal
redis-cli MEMORY DOCTOR
# find offenders without blocking the server (prefer a replica)
redis-cli --memkeys -i 0.01
redis-cli --bigkeys --pattern 'session:*'
redis-cli MEMORY USAGE cart:9182 SAMPLES 0 # exact, but O(N): use with carego deeper
Name the tools and what each measures: INFO memory for totals, MEMORY USAGE for one key, --bigkeys/--memkeys to find offenders, and never KEYS *.
Show the top-down order and the count-versus-bytes distinction, plus that MEMORY USAGE samples nested elements by default.
Add the safety and interpretation layer: scan a replica, throttle with -i, read MEMORY STATS overhead buckets, and know that buffers and backlog count toward maxmemory.
Turn it into a standing practice: baselined bytes-per-key and dataset-versus-overhead ratios, offline RDB analysis for deep audits, and thresholds that trigger remodelling rather than more RAM.
## Start with the totals, not with keys The first question is not "which key is big" but "is this even the dataset?". `INFO memory` answers that: - `used_memory` — bytes the allocator has handed to Redis. This is the number `maxmemory` compares against, and it includes far more than your values: client output buffers, the replication backlog, the AOF buffer, script caches. - `used_memory_rss` — resident bytes as the operating system sees them. This is what a container limit and the OOM killer care about. - `used_memory_peak` — the high-water mark; a huge gap between peak and current usually means a transient burst (a giant pipeline, a slow consumer, a rewrite) rather than steady-state data. - `mem_fragmentation_ratio` — RSS divided by used_memory, plus the finer `allocator_frag_ratio` and `rss_overhead_ratio`. If `used_memory` itself is modest and RSS is huge, the story is fragmentation or process overhead, and chasing individual keys is wasted effort. ## Break down the total `MEMORY STATS` returns a structured report: `dataset.bytes` and `dataset.percentage`, `overhead.total`, `keys.count`, `keys.bytes-per-key`, `replication.backlog`, `clients.slaves`, `clients.normal`, `aof.buffer`, plus per-database overhead. Two very common findings live here rather than in your data: a replication backlog sized generously, and client output buffers inflated by a slow consumer or a subscriber that cannot keep up. `keys.bytes-per-key` is the number to remember for capacity work — it folds in the dictionary entry, the key string, the object header and the expires entry, which is why tiny values still cost on the order of a hundred bytes each. `MEMORY DOCTOR` runs heuristics over these numbers and prints prose advice ("fragmentation is high", "peak is much larger than current"). Treat it as a checklist prompt. ## Then find the offenders `MEMORY USAGE <key>` reports the bytes for one key and its value, including the key itself and its administrative overhead. For aggregate types it estimates by sampling nested elements — five by default — so for a hash with wildly varying field sizes the estimate can be off; `SAMPLES 0` makes it exact at the price of walking everything, which is an O(N) command on the single thread and must not be aimed at a huge key on a busy primary. To sweep the whole keyspace, `redis-cli` has two modes built on SCAN, so they use cursor-based iteration in small batches instead of blocking: - `--bigkeys` reports, per type, the key with the most elements, using STRLEN/LLEN/SCARD/HLEN/ZCARD/XLEN. Crucially this ranks by **cardinality, not bytes**: a 2-million-field hash of tiny flags outranks a single 200 MB string in its output, and only the latter may be your actual problem. - `--memkeys` calls MEMORY USAGE per key and ranks by **bytes**, which is what you usually want, at higher cost per key. Both accept `--pattern` to restrict the sweep, and `-i <seconds>` to sleep between batches so you do not add meaningful load. Even so, prefer running them on a replica. ## What not to do `KEYS *` walks the entire keyspace in one command on the single thread and will stall every other client for as long as it takes — seconds on a large instance. The same goes for `DEBUG OBJECT` loops, `HGETALL` on unknown-size hashes, and `MEMORY USAGE ... SAMPLES 0` on a suspected monster key. For deep offline analysis, take an RDB snapshot and run an offline analyzer against the file instead, which costs the live instance nothing beyond the snapshot itself. ## Putting it together A workable order: `INFO memory` → is it data or overhead? → `MEMORY STATS` → which bucket? → if dataset, `--memkeys` (on a replica) → confirm the top offenders with `MEMORY USAGE` → decide whether the fix is remodelling a key, adding TTLs, changing the eviction policy, or reclaiming fragmentation.
- Why can --bigkeys point at a key that is not actually your memory problem?It ranks by number of elements, obtained from cheap cardinality commands, not by bytes. A set holding ten million short integers may report as the biggest key while consuming less memory than one string holding a few hundred megabytes of serialized JSON. Use --memkeys, or confirm candidates with MEMORY USAGE, when bytes are the question.
- MEMORY STATS shows a large overhead that the dataset does not explain. What are the usual causes?Most often the replication backlog sized large, client output buffers held by a slow consumer or a Pub/Sub subscriber that cannot keep up, and the AOF buffer during a rewrite. All of these count toward used_memory and therefore toward maxmemory, so they can trigger eviction even when your data has not grown.
saying these in an interview costs you the question
- Running KEYS * on a production instance to inventory memory
- Reading --bigkeys output as a ranking by bytes
- Assuming used_memory covers only stored values
- Using MEMORY USAGE with SAMPLES 0 on a suspected huge key on the primary
- Treating MEMORY DOCTOR output as authoritative measurement