skip to content

Redis can be configured to append every write command to an append-only file (AOF). What does the BGREWRITEAOF command do, and why does an append-only log need rewriting at all?

level: middleimportance: must knowfreq 58%

answer

  1. Log of commands, grows with write volume not data size
  2. Rewrite is generated from memory, never from the old file
  3. RDB preamble base by default since 4.0
  4. Restart = replay, so file size is downtime
  5. One child at a time: rewrite may be 'scheduled'

basics

~20 s

The AOF records every write, so it grows with write volume and repeats history. BGREWRITEAOF forks a child that writes a new minimal file reproducing the current in-memory dataset, then atomically swaps it in: smaller file, faster restart.

solid answer

~50 s

The AOF is a log of the write commands Redis executed, serialized in RESP. It grows with **write traffic**, not with dataset size: a counter incremented a million times is a million `INCR` entries for one key, and keys later deleted still occupy log space. That costs disk and, more importantly, restart time, because loading an AOF means replaying it. `BGREWRITEAOF` regenerates the file from **memory**, not from the old log. Redis forks a child; the child walks the current dataset and emits the shortest representation that reproduces it (an RDB image of the dataset when `aof-use-rdb-preamble yes`, the default since 4.0). The parent keeps serving; writes arriving during the rewrite go to a rewrite buffer (before 7.0) or to a fresh incremental file (7.x). On completion Redis swaps atomically: rename in older versions, manifest update in 7.x. Rewrites also start automatically via `auto-aof-rewrite-percentage`/`auto-aof-rewrite-min-size`, and when you enable AOF at runtime with `CONFIG SET appendonly yes`.

code

text · 10 lines
text
# 1000 writes to one key produce ~1000 AOF entries
redis-cli EVAL "for i=1,1000 do redis.call('INCR', KEYS[1]) end return redis.call('GET', KEYS[1])" 1 hits

redis-cli INFO persistence | grep -E 'aof_base_size|aof_current_size|aof_rewrite_in_progress'

# Regenerate the file from the in-memory dataset
redis-cli BGREWRITEAOF
# Background append only file rewriting started

redis-cli INFO persistence | grep -E 'aof_last_bgrewrite_status|aof_base_size|aof_current_size'

go deeper

for a junior

Know that AOF logs write commands, grows over time, and BGREWRITEAOF shrinks it in the background without stopping the server.

for a middle

Explain that the file tracks write volume, that the rewrite is generated from memory by a forked child, and that restart time is the real motivation.

for a senior

Add the swap mechanics, the RDB-preamble base, single-child scheduling, and the INFO fields you would check after a failed rewrite.

for a principal

Frame it as an RTO and capacity decision: rewrite frequency trades fork and copy-on-write cost against file size and restart time, and on large instances persistence often belongs on a replica.

## What the AOF actually contains With `appendonly yes`, Redis appends every command that **changed** the dataset to a file, serialized in the same RESP protocol clients speak. Reads are never logged. On startup the server replays that file command by command and rebuilds the dataset. The AOF is therefore a *command log*: it stores what happened, not what the data currently looks like. That distinction is the whole reason rewriting exists. ## Why an append-only log must be rewritten Because the file is history, its size tracks **write volume over time**, not the size of the data: - A single counter hit a million times produces roughly a million `INCR` entries, for one key holding one integer. - A list used as a queue that receives and drains ten million items leaves twenty million entries behind, for a list that is currently empty. - Keys written and later deleted, or expired, still occupy their original entries; the deletion adds more. Two costs follow. The obvious one is disk: an AOF can be many multiples of the dataset. The one that actually hurts in production is **restart time**. Loading an AOF is not reading a data file, it is re-executing every command in it, single-threaded. A 40 GB AOF can take many minutes to load, and that time is downtime after a crash or a deploy. ## What BGREWRITEAOF does `BGREWRITEAOF` produces a new AOF that reproduces the dataset with the minimum amount of data, and it does so **from the in-memory dataset**, never by reading or compacting the old file. Mechanically: 1. Redis forks a child process. The child sees a frozen, consistent snapshot of memory (courtesy of copy-on-write) at the instant of the fork. 2. The child serializes that snapshot. With `aof-use-rdb-preamble yes` (default since 4.0) it writes it in RDB binary format, which is far more compact and much faster to load than commands; with the option off it writes plain commands such as one `RPUSH` per list, one `HSET` per hash, batched to bounded sizes. 3. Meanwhile the parent keeps serving clients. Writes that arrive during the rewrite must not be lost: before 7.0 they were accumulated in an in-memory *AOF rewrite buffer* and appended to the child's output at the end; from 7.0 they are written straight into a new incremental file that the manifest will reference. 4. When the child finishes successfully, Redis swaps the new content in atomically: a `rename()` over the old file in older versions, an atomic manifest update in 7.x. The old files are then removed. The `BG` prefix means background: the command returns immediately and only the `fork()` itself blocks the server. ## Consequences of "from memory, not from the file" - The result is **minimal by construction**. There is no partial compaction, no incremental merge of the old log. - A corrupt existing AOF is *not* repaired by a rewrite; the corrupt file is simply discarded once the new one is in place. Do not think of the rewrite as a repair tool. - Anything not in memory cannot be in the new file. Keys that expired before the fork are gone, which is exactly what you want, but it also means the rewrite reflects the dataset at fork time plus everything appended afterwards. ## What starts a rewrite - **Manually**: `BGREWRITEAOF`. - **Automatically**: when the file has grown by `auto-aof-rewrite-percentage` over the size recorded after the previous rewrite, and is at least `auto-aof-rewrite-min-size`. - **On enabling AOF at runtime**: `CONFIG SET appendonly yes` triggers a rewrite immediately, because that is how the file gets seeded from the current dataset. Only one persistence child may run at a time. If a `BGSAVE` is in flight, the rewrite is *scheduled* rather than started; `INFO persistence` shows `aof_rewrite_scheduled:1` until it actually begins. ## Observability `INFO persistence` is the interface: `aof_enabled`, `aof_rewrite_in_progress`, `aof_rewrite_scheduled`, `aof_last_bgrewrite_status`, `aof_last_write_status`, `aof_base_size`, `aof_current_size`, and `latest_fork_usec` for the fork stall. A rewrite that fails (typically disk full) leaves `aof_last_bgrewrite_status:err` and Redis retries later; the old file stays authoritative, so you keep durability but keep growing. ## The cost side A rewrite is not free: it forks (a latency spike proportional to how much memory the process maps), it consumes copy-on-write memory for the duration proportional to the write rate, and it writes the whole dataset to disk, competing with the main AOF's own writes. That is why the trigger thresholds are tunable and why very large instances often move persistence to a replica.

  • Does a rewrite read and compact the existing append-only file?
    No. The child serializes the in-memory dataset and produces a brand-new file; the old one is discarded on success. This is why a rewrite cannot repair a corrupted AOF, and why the output is minimal rather than partially compacted. It also means the new file reflects the fork-time snapshot plus everything appended afterwards.
  • What happens if you issue BGREWRITEAOF while a BGSAVE child is already running?
    Redis does not fork a second persistence child. The rewrite is marked as scheduled and starts once the current child exits; `INFO persistence` reports `aof_rewrite_scheduled:1` in the meantime. The command still returns success immediately, so scripts that assume the rewrite started right away can be misled.
  • Why does turning on AOF with CONFIG SET appendonly yes trigger a rewrite?
    Because there is no existing log describing the current dataset, and Redis cannot invent history it did not record. A rewrite is exactly the operation that produces a file matching what is in memory, so enabling AOF at runtime seeds the file that way and normal appending continues from there.

The AOF is a diary of every move in a chess game; a rewrite throws the diary away and photographs the current board, which is all you need to keep playing.

saying these in an interview costs you the question

  • Saying the AOF grows in proportion to dataset size rather than to write volume
  • Claiming BGREWRITEAOF compacts or repairs the existing file
  • Thinking the rewrite blocks the server for its whole duration rather than only for the fork
  • Assuming the rewritten file is always plain commands, missing the RDB preamble default
  • Believing a rewrite improves durability; it only changes file size and load time

context