skip to content

INFO memory on one Redis instance reports mem_fragmentation_ratio of 1.9, and on another it reports 0.7. Interpret each of those numbers and say what you would do about them.

level: seniorimportance: should knowfreq 42%

answer

  1. ratio = RSS / used_memory
  2. 1.0-1.5 normal; ~1.9 = waste or fork overhead
  3. allocator_frag_ratio vs rss_overhead_ratio
  4. <1 means swapping - emergency
  5. activedefrag (jemalloc) or restart

basics

~20 s

The ratio is resident memory divided by allocated memory. 1.9 means the process holds nearly twice what Redis uses - real fragmentation or transient overhead; consider activedefrag or a restart. Below 1, like 0.7, means part of the process is swapped to disk, which is a latency emergency.

solid answer

~50 s

`mem_fragmentation_ratio = used_memory_rss / used_memory`. **1.9** — the OS holds nearly double what the allocator handed out. Two causes: genuine allocator fragmentation (typical after mass deletes or workloads whose value sizes churn across size classes, since freed memory returns to jemalloc, not to the OS), or non-data resident overhead (a snapshot fork's copied pages, a recent peak, large buffers). Split them with `allocator_frag_ratio` versus `rss_overhead_ratio`. If it is allocator fragmentation and jemalloc is the allocator, enable `activedefrag yes` and tune `active-defrag-ignore-bytes`, `active-defrag-threshold-lower/upper` and the CPU bounds; defrag costs CPU on the main thread. A failover-and-restart also fixes it instantly. **0.7** — resident is *smaller* than allocated, meaning pages have been swapped out. Redis latency collapses when the single thread touches a swapped page. Fix the machine: remove or shrink swap pressure, lower `maxmemory`, add RAM, and check for other processes on the host. On tiny datasets the ratio is meaningless — baseline overhead dominates.

code

text · 15 lines
text
redis-cli INFO memory | grep -E 'frag|rss|used_memory:'
#  used_memory:10737418240
#  used_memory_rss:20401094656
#  mem_fragmentation_ratio:1.90
#  allocator_frag_ratio:1.62     <- real allocator waste
#  allocator_frag_bytes:4.1G
#  rss_overhead_ratio:1.14       <- fork/buffers, defrag won't help this part

# redis.conf (jemalloc only)
activedefrag yes
active-defrag-ignore-bytes 100mb
active-defrag-threshold-lower 10
active-defrag-threshold-upper 100
active-defrag-cycle-min 1
active-defrag-cycle-max 25

go deeper

for a junior

Recall the formula and the two directions: well above 1 means overhead or fragmentation, below 1 means swapping and is bad.

for a middle

Explain why freed memory stays with the allocator and name activedefrag plus restart as the two remedies.

for a senior

Distinguish allocator fragmentation from RSS overhead using the finer ratios, weigh defrag's CPU cost, and treat a sub-1 ratio as a host-level incident.

for a principal

Set policy: headroom in maxmemory sizing so forks and fragmentation cannot push into swap, when to defrag versus rotate a node, and what to alert on.

## What the ratio is made of `INFO memory` reports two independent measurements. `used_memory` is what the memory allocator (jemalloc by default) says it has handed to Redis. `used_memory_rss` is the resident set size the operating system attributes to the process. The ratio is simply RSS divided by used_memory, and it answers: how much physical memory is this process holding per byte Redis thinks it is using? A value in the neighbourhood of 1.0 to 1.5 is unremarkable. There is always some gap: allocators round requests up into size classes, keep per-thread caches, and hold pages they are not currently handing out. ## Reading 1.9 This says roughly 90% overhead. There are two distinct stories. **Allocator fragmentation.** Memory freed by Redis goes back to jemalloc, not to the kernel. If a workload allocated millions of 200-byte values and then deleted most of them, the pages remain resident with live objects scattered across them; no page can be returned because each still holds something. Workloads that churn value sizes across allocator size classes — growing hashes, lists that swing in length, mass expiry followed by different-shaped data — produce this reliably. **Non-data resident overhead.** A snapshot fork causes the kernel to copy pages the parent modifies while the child writes, so RSS climbs during and after persistence work. Big client output buffers, a recent peak now freed, or a large replication backlog also raise RSS relative to the dataset. Redis reports finer figures to separate these: `allocator_frag_ratio` and `allocator_frag_bytes` isolate the allocator's own waste, while `rss_overhead_ratio` captures what is resident beyond the allocator's accounting. If the allocator ratio is near 1.0 while the overall ratio is 1.9, defragmentation will do nothing — the memory is elsewhere. **Remedies, in order of intrusiveness.** Confirm it is allocator fragmentation. Enable active defragmentation (`activedefrag yes`, jemalloc-only), which incrementally relocates objects to free whole pages; gate it with `active-defrag-ignore-bytes` (don't bother below an absolute waste) and `active-defrag-threshold-lower`/`upper` (percentages at which it starts and goes full-throttle), and bound its CPU with the min/max effort settings, because defrag runs on the thread that serves commands. Watch tail latency while it works. If fragmentation is the result of a one-off event rather than a steady pattern, the cheapest fix is a controlled failover to a replica and a restart of the fragmented node, which returns everything to the OS. ## Reading 0.7 A ratio below 1 means the OS has less of the process resident than the allocator has allocated — the missing pages are on disk, swapped out. For a database that expects every access to be a memory access on one thread, this is severe: a single swapped page touched by a command turns a microsecond operation into a millisecond-plus disk read, and because everything is serialized on one thread, every other client waits behind it. Latency graphs go from flat to spiky and throughput craters. The fix is environmental, not a Redis tunable: find out why the host is under memory pressure. Common causes are `maxmemory` set too close to (or above) physical RAM, other processes co-located on the box, a container limit lower than the workload, or a snapshot fork transiently doubling demand. Lower `maxmemory` to leave real headroom, move noisy neighbours off, add RAM, or reduce the dataset. Confirm from the OS side too — the process's swap usage — rather than trusting the ratio alone. ## Caveats worth stating On a nearly empty instance the ratio is noise: baseline process memory swamps a small dataset, so 3.0 on a 20 MB dataset means nothing. Read the ratio together with absolute numbers (`allocator_frag_bytes` in particular) and only act when the wasted bytes matter at your scale. And remember the direction of the arithmetic: memory freed inside Redis lowers `used_memory` immediately, which *raises* the ratio even though nothing got worse — a mass delete or a wave of expiry routinely produces an alarming-looking ratio for a while.

  • Right after deleting several million keys the fragmentation ratio jumped from 1.1 to 2.3. Did something break?
    No. Deleting lowers used_memory immediately, but the freed memory returns to the allocator rather than to the kernel, so resident size barely moves and the ratio rises arithmetically. If the workload will refill that space, the pages get reused and the ratio settles; if not, active defragmentation or a restart is what actually returns memory to the OS.
  • Why is active defragmentation not simply left on everywhere?
    It relocates live objects on the same thread that serves commands, so it consumes CPU and adds tail latency exactly when the instance is already stressed. It also only helps allocator fragmentation, and only with jemalloc. It is gated behind absolute-bytes and percentage thresholds so it engages when the waste is genuinely worth paying for.

saying these in an interview costs you the question

  • Reading a ratio below 1 as unusually good memory efficiency
  • Turning on active defragmentation when the overhead is fork or buffer related
  • Treating a high ratio on a nearly empty instance as a real problem
  • Expecting freed Redis memory to return to the operating system automatically
  • Assuming a restart is the only possible remedy for fragmentation

context