skip to content

A Redis process holding about 8 GB of data momentarily grows toward 16 GB of resident memory whenever a background snapshot runs, and the host occasionally OOM-kills it. Explain the mechanism behind that growth and what you would change to stop the kills.

level: seniorimportance: must knowfreq 55%

answer

  1. fork copies page tables, then copy-on-write
  2. child only reads; parent's writes cause copies
  3. worst case ~2x resident memory
  4. THP turns 4KB copies into 2MB copies
  5. rdb_last_cow_size and latest_fork_usec

basics

~20 s

The snapshot child is a fork() sharing memory copy-on-write. Every page the parent writes during the save gets duplicated, so a write-heavy instance can approach double its size. Fix with memory headroom (maxmemory well under RAM), disabled transparent huge pages, vm.overcommit_memory=1, and snapshotting from a replica.

solid answer

~60 s

`BGSAVE` forks. The child does not copy 8 GB up front — the kernel marks the pages read-only and shares them, copying a page only when either side writes to it. The child never writes, so every copy is caused by the **parent still serving writes**. At the worst extreme, if the parent touches every page before the child finishes, you pay a full second copy: ~2× resident memory. What drives it: write rate, spread of writes across the keyspace, and how long the save takes. Mitigations, roughly in order: - **Headroom.** Set `maxmemory` to about half of physical RAM on a write-heavy box, so the worst case fits. - **Disable transparent huge pages.** With THP, copy-on-write duplicates 2 MB pages instead of 4 KB ones — massive amplification plus latency. Redis logs a warning about this at startup. - **`vm.overcommit_memory = 1`**, or `fork()` itself fails and the save errors out (which, with `stop-writes-on-bgsave-error yes`, starts rejecting writes). - **Move snapshots to a replica** and keep the primary on `save ""`, so the fork cost lives where latency and memory are cheap.

code

text · 7 lines
text
$ redis-cli INFO persistence | grep -E 'rdb_bgsave_in_progress|rdb_last_cow_size|rdb_last_bgsave_status'
rdb_bgsave_in_progress:1
rdb_last_cow_size:3187671040        # ~3 GB duplicated by the previous save
rdb_last_bgsave_status:ok

$ redis-cli INFO stats | grep latest_fork_usec
latest_fork_usec:412000              # 412 ms main-thread stall at fork

go deeper

for a junior

Know that a background snapshot forks a child, memory is shared until written, and writes during the save cause extra memory use.

for a middle

Explain copy-on-write page by page, why the worst case approaches double, and name the fork pause as a separate cost from the memory growth.

for a senior

Diagnose with rdb_last_cow_size, latest_fork_usec, and RSS versus used_memory; prescribe THP off, overcommit 1, maxmemory headroom, and snapshotting from a replica.

for a principal

Turn it into a capacity and topology decision: instance sizing versus shard count, where backup forks are allowed to run, and an explicit memory budget that accounts for worst-case copy-on-write rather than steady state.

## What `fork()` actually does When Redis runs a background save it calls `fork()`. The child is a separate process with a logically complete copy of the parent's address space — but the kernel does not copy 8 GB of data. It copies the **page tables**, marks every shared page read-only in both processes, and returns. That is copy-on-write (COW): the duplicate is created lazily, one page at a time, the first time either process writes to it. Two distinct costs follow from this, and candidates routinely conflate them. **Cost one: the fork pause.** Copying page tables is work proportional to how much memory the process maps. On bare metal it is roughly on the order of ten milliseconds per gigabyte; on some virtualized instances it is several times worse. During it the main thread is stopped, so it shows up as a periodic latency spike exactly aligned with save points. `INFO stats` reports `latest_fork_usec` — the single most useful number when someone says "Redis stalls every five minutes". **Cost two: copy-on-write memory growth.** This is the 8 GB → 16 GB symptom. The child only reads (it is serializing a frozen view), so every page copy is triggered by the **parent**, which is still accepting writes. Each write to a shared page forces the kernel to allocate a private copy for the parent. The upper bound is one full duplicate of the dataset — hence "up to 2×" — reached when the parent touches every page before the child finishes. ## What makes it worse than theory - **Write rate and spread.** A million writes concentrated on a few hot keys copy few pages. The same volume scattered across the keyspace copies many. Random-access write patterns are the pathological case. - **Save duration.** The longer the child runs (big dataset, slow disk, saturated I/O), the more of the keyspace the parent has time to dirty. Slow disks therefore cost memory, not just time. - **Transparent huge pages (THP).** With THP enabled, the kernel backs memory with 2 MB pages, so a single 8-byte write copies 2 MB instead of 4 KB — up to a 512× amplification of the copy, plus much worse latency. Redis prints an explicit startup warning recommending `never` for `/sys/kernel/mm/transparent_hugepage/enabled`. This is the single most common cause of a snapshot spike that looks far larger than the write rate should justify. - **Allocator behavior.** Fragmentation means logically small updates can touch pages across many arenas; `mem_fragmentation_ratio` in `INFO memory` is worth reading alongside the spike. - **Expiration and eviction.** These are writes too. Active expiry cycles and eviction under `maxmemory` dirty pages during the save just like client traffic does. ## Why the process gets killed If the box has 16 GB and Redis normally sits at 8 GB, everything looks fine until a save runs and the copies push the resident set toward the limit. Then either the kernel OOM killer picks Redis (it is the biggest process), or `fork()` itself fails with `ENOMEM` and the save errors — and with the default `stop-writes-on-bgsave-error yes`, the surviving server starts rejecting writes with `MISCONF`. Both outcomes are self-inflicted capacity problems. ## The fixes, in the order to apply them 1. **`vm.overcommit_memory = 1`.** Linux's default heuristic overcommit can refuse a fork that would nominally need another 8 GB even though COW means it will not actually use it. Redis warns about this at startup. Set it in `/etc/sysctl.conf`. Note this makes forks succeed; it does not create memory, so it must be paired with real headroom. 2. **Disable THP** (`never`, and remove any boot-time `always`). This alone often collapses the spike. 3. **Set `maxmemory` with headroom.** On a write-heavy instance the classic guidance is to keep the dataset around half of physical RAM so the worst-case duplicate fits. If writes are light and localized, you can run hotter — measure rather than guess. 4. **Reduce snapshot frequency** (`save` points) so the exposure window per hour shrinks — accepting a larger data-loss window in exchange. 5. **Move snapshotting off the primary.** Attach a replica, set `save ""` on the primary, and take snapshots and backups from the replica. The primary then only forks for AOF rewrites and replica syncs. 6. **Shard.** Two 4 GB instances fork faster and spike less than one 8 GB instance, and the spikes are not simultaneous. ## Things that do *not* help - Adding CPU cores: the fork pause is a page-table copy on one thread, not a parallelizable workload. - Switching to the append-only file to "avoid the fork": AOF **rewrites** fork too, with the same copy-on-write dynamics. So does a replica full-sync, diskless or not. If forking is the problem, you must reason about all three sources. ## What to look at while diagnosing `INFO persistence` for `rdb_bgsave_in_progress`, `rdb_last_bgsave_status`, `rdb_last_cow_size` (how much memory the last child's copy-on-write actually consumed — the direct measurement of this phenomenon); `INFO stats` for `latest_fork_usec`; `INFO memory` for `used_memory_rss` against `used_memory`; and the host's own OOM/kernel log to confirm who killed what.

  • Would switching from snapshots to the append-only file avoid this memory spike?
    No, not by itself. An AOF rewrite also forks a child and has exactly the same copy-on-write dynamics; so does a replica full-synchronization. You would trade a frequent save-point fork for a less frequent rewrite fork, which reduces how often you pay it but does not remove the mechanism. If forks are the problem, reduce dataset size per instance or move the work to a replica.
  • How do you measure how much memory a given background save actually duplicated?
    `INFO persistence` exposes `rdb_last_cow_size` — the copy-on-write memory consumed by the last RDB child (with `aof_last_cow_size` for the rewrite child). Compare it against `used_memory` to see what fraction of the dataset was dirtied during the save, and watch `used_memory_rss` during the window to see the real resident peak the host must absorb.
  • Why does Redis warn at startup when `vm.overcommit_memory` is 0?
    Because Linux's default heuristic can refuse a `fork()` whose nominal memory requirement (a full copy of the address space) exceeds available memory plus swap, even though copy-on-write means the child will never actually use that much. The refused fork makes background saves and AOF rewrites fail, which — with `stop-writes-on-bgsave-error yes` — can escalate into the server rejecting writes.

Handing a photographer a snapshot of the warehouse: nothing is duplicated until you move a crate, and each crate you move has to be cloned so the photo still shows where it was. Move every crate and you have cloned the whole warehouse.

saying these in an interview costs you the question

  • Saying fork() immediately copies the whole dataset — it copies page tables and defers data copies
  • Blaming the child process for the extra memory; the parent's writes cause the copies
  • Claiming the spike is proportional to dataset size alone, ignoring write rate and save duration
  • Leaving transparent huge pages enabled, or not knowing they amplify copy-on-write
  • Proposing the append-only file as a way to avoid forking

context