skip to content

Redis runs an append-only-file (AOF) rewrite in a forked child process, the same way a background RDB snapshot (BGSAVE) does. Setting aside the generic cost of fork() and copy-on-write, what is different about an AOF-rewrite child in production, and which Redis INFO fields and LATENCY monitor events would you watch to observe one?

level: seniorimportance: must knowfreq 48%

answer

  1. latest_fork_usec = syscall only; aof_current_rewrite_time_sec = child life
  2. scheduled ≠ started: aof_rewrite_scheduled:1
  3. aof_last_cow_size usually > rdb_last_cow_size (wider window)
  4. no-appendfsync-on-rewrite = silent fsync holiday
  5. LATENCY LATEST: fork vs aof-write-active-child

basics

~20 s

An AOF-rewrite child lives far longer than a BGSAVE on the same data, and the parent keeps writing AOF to the same disk throughout. Watch latest_fork_usec, aof_rewrite_in_progress, aof_rewrite_scheduled, aof_last_bgrewrite_status, aof_last_cow_size, and LATENCY LATEST's fork event.

solid answer

~50 s

The fork syscall itself is identical to a BGSAVE's — page-table copy and copy-on-write cost belong to the RDB fork/BGSAVE discussion, not here. Four things are rewrite-specific: 1. The child usually lives much longer than a BGSAVE over the same data, so the copy-on-write window is wider — `aof_last_cow_size` typically exceeds `rdb_last_cow_size`. 2. The parent is writing AOF to the same device for the child's whole lifetime, so I/O contention lasts the entire rewrite, not just the fork instant. 3. `no-appendfsync-on-rewrite yes` silently suspends the parent's fsyncs for the whole rewrite, turning a one-second loss window into a minutes-long one, with no metric announcing it. 4. `BGREWRITEAOF` issued while another persistence child runs is only *scheduled* (`aof_rewrite_scheduled:1`), so it looks like nothing happened. Observe via `INFO persistence` (`aof_rewrite_in_progress`, `aof_rewrite_scheduled`, `aof_current_rewrite_time_sec`, `aof_last_rewrite_time_sec`, `aof_last_bgrewrite_status`, `aof_rewrites_consecutive_failures`, `aof_last_cow_size`), `INFO stats` → `latest_fork_usec`, and `LATENCY LATEST` (`fork`, `aof-write-active-child`).

code

text · 13 lines
text
# The fork syscall stall (microseconds) - NOT the rewrite duration
redis-cli INFO stats | grep -E 'latest_fork_usec|total_forks'

# Rewrite child state, duration, outcome and copy-on-write cost
redis-cli INFO persistence | grep -E \
  'aof_rewrite_in_progress|aof_rewrite_scheduled|aof_current_rewrite_time_sec|aof_last_rewrite_time_sec|aof_last_bgrewrite_status|aof_rewrites_consecutive_failures|aof_last_cow_size|rdb_last_cow_size|current_cow_size|aof_base_size|aof_current_size'

# "Nothing happened" after BGREWRITEAOF: queued behind another child
redis-cli BGREWRITEAOF
# Background append only file rewriting scheduled   -> aof_rewrite_scheduled:1

# The silent durability trade - only visible in the config
redis-cli CONFIG GET no-appendfsync-on-rewrite appendfsync

go deeper

for a junior

Know that a rewrite happens in a forked child and that INFO persistence tells you whether one is running (aof_rewrite_in_progress) and whether the last one succeeded (aof_last_bgrewrite_status).

for a middle

Be able to name latest_fork_usec as the fork stall versus aof_last_rewrite_time_sec as the child's lifetime, and explain that BGREWRITEAOF may only be scheduled when another persistence child is active.

for a senior

Diagnose from the metrics: separate the fork spike from whole-window disk contention using LATENCY's fork and aof-write-active-child events, compare aof_last_cow_size with rdb_last_cow_size, and call out no-appendfsync-on-rewrite as a silent durability change.

for a principal

Frame it as an observability and risk contract: which of these fields are actually scraped and alerted on, what the fsync-suspension window means for the stated RPO, and whether persistence should fork on this node at all rather than on a replica.

## Scope: what is *not* the subject here Every `fork()` in Redis costs a synchronous page-table copy that stalls command execution, and every forked child creates a copy-on-write window during which pages the parent writes get physically duplicated. That mechanism is generic to background persistence and is covered by the RDB fork / copy-on-write and `SAVE` vs `BGSAVE` material — take the derivation from there. This question is about what makes the *AOF-rewrite* child a different operational animal, and how you read its state off a running server. ## Four rewrite-specific effects a snapshot fork does not have **1. The child lives much longer, so the copy-on-write window is wider.** A `BGSAVE` child writes one compact binary RDB and exits. An AOF-rewrite child writes a full base file and, on older configurations that serialize the base as plain commands rather than an RDB preamble, that output is several times bulkier and slower to produce. Longer child life means more parent writes land inside the window, so peak copy-on-write memory for a rewrite is normally *higher* than for a snapshot of the same dataset. Redis reports both separately: `aof_last_cow_size` for the last rewrite, `rdb_last_cow_size` for the last snapshot. Compare them; if the AOF number is several times the RDB number, that is the wider window showing up, not a different mechanism. **2. The parent is writing to the same device the whole time.** During a `BGSAVE` with AOF disabled, only the child touches the disk. During a rewrite, the parent keeps appending live writes (in Redis 7.0+ to an incremental file alongside the new base; the multi-part AOF layout is its own topic) while the child streams the base out. Two writers, one device. The signature is a latency profile that is *elevated for the whole rewrite duration* rather than one spike at the start — and that duration is directly visible as `aof_current_rewrite_time_sec` while it runs and `aof_last_rewrite_time_sec` afterwards. **3. `no-appendfsync-on-rewrite yes` silently widens the durability window.** With this set, the parent stops issuing `fsync` on the AOF for as long as *any* persistence child is running. With `appendfsync everysec` you believe your worst-case loss is one second; during a rewrite it is however long the rewrite takes — minutes on a large instance. Nothing logs this and no INFO field says "fsyncs suspended"; the only way to know is to read `CONFIG GET no-appendfsync-on-rewrite` and multiply by `aof_last_rewrite_time_sec`. It is a real latency-for-durability trade, but it must be a stated one. **4. A rewrite requested while another child runs is scheduled, not started.** `BGREWRITEAOF` while an AOF rewrite is already in progress returns an error. While a *different* persistence child is active — an RDB save, a replica full-sync fork, a module fork — or inside `MULTI`, Redis replies `Background append only file rewriting scheduled`, sets `aof_rewrite_scheduled:1`, and starts the rewrite when the other child exits. Automatic size-based triggers are deferred the same way. So "I ran BGREWRITEAOF and nothing happened" is usually `aof_rewrite_scheduled:1` plus something that keeps occupying the fork slot; if that something recurs constantly (frequent `save` points, repeated replica resyncs), the AOF grows unbounded while the rewrite never gets a turn. ## The fields, and what each one actually tells you - `INFO stats` → `latest_fork_usec`: duration of the *fork syscall only*, in microseconds. This is the instantaneous full-server stall, not the rewrite duration. `total_forks` tells you how often you pay it. - `aof_rewrite_in_progress`: 1 while a child is running. - `aof_rewrite_scheduled`: 1 when a rewrite is queued behind another child. Stuck at 1 = starvation. - `aof_current_rewrite_time_sec` / `aof_last_rewrite_time_sec`: the child's lifetime — the width of the copy-on-write and contention window. - `aof_last_bgrewrite_status`: `ok` or `err`. `err` means the AOF is still growing; Redis backs off before retrying, and `aof_rewrites_consecutive_failures` counts the streak. - `aof_last_cow_size` (and `rdb_last_cow_size` for comparison), plus the live Redis 7 progress fields `current_cow_size`, `current_cow_peak`, `current_fork_perc`. - `aof_base_size` vs `aof_current_size`: the growth ratio the auto-trigger reasons about. - `aof_rewrite_buffer_length` existed before 7.0 as the parent-side diff buffer; multi-part AOF removed it, so do not look for it on 7.x. ## LATENCY monitor Set `latency-monitor-threshold` to a few milliseconds and Redis records named latency events. `LATENCY LATEST` returns each event with its last and max value; the `fork` event isolates the stall so you can prove or disprove that the fork caused a spike, and `LATENCY HISTORY fork` gives the series. The AOF write-path events matter just as much: `aof-write-active-child` fires when a parent AOF write blocked *while a child was running* — direct evidence of rewrite-induced disk contention rather than fork cost — alongside `aof-write-pending-fsync`, `aof-fsync-always` and `aof-rename`. ## Reading it as a routine A single spike whose width matches `latest_fork_usec` at the moment `aof_rewrite_in_progress` flips to 1 is the fork. Degradation spanning `aof_current_rewrite_time_sec` with `aof-write-active-child` events is contention. Rising `current_cow_size` through the rewrite is copy-on-write. `aof_rewrite_scheduled:1` with a growing `aof_current_size` is starvation. `aof_last_bgrewrite_status:err` is a rewrite that never happened at all.

  • aof_rewrite_scheduled has read 1 for the last hour and the AOF file keeps growing. What is happening, and how do you confirm it?
    A rewrite is queued behind some other persistence child that keeps re-occupying the fork slot, so it never starts. Check rdb_bgsave_in_progress and the save points, sync_full/sync_partial_ok in INFO stats for repeated replica full syncs, and the log for repeated "Background saving started" lines. Also check aof_last_bgrewrite_status and aof_rewrites_consecutive_failures — a failing rewrite is retried only after a back-off delay, which looks similar from the outside.
  • During a rewrite you see latency degraded for the whole child lifetime rather than one spike at the start. How do you tell fork cost from disk contention?
    Fork cost is a single stall at the moment aof_rewrite_in_progress flips to 1, and its width should match latest_fork_usec and the LATENCY monitor's fork event. Degradation spanning aof_current_rewrite_time_sec is the parent and child competing for the same device; aof-write-active-child and aof-write-pending-fsync events plus device utilisation in iostat confirm it. Copy-on-write pressure is a third pattern — rising current_cow_size and RSS through the window.
  • Which INFO field shows the memory the last rewrite child actually cost, and why is it usually larger than the snapshot equivalent?
    aof_last_cow_size, against rdb_last_cow_size for the last background save. It is normally larger because the rewrite child lives much longer than a BGSAVE over the same dataset, so more parent writes fall inside the copy-on-write window and dirty more pages. The underlying copy-on-write mechanism is identical; only the window width differs.

saying these in an interview costs you the question

  • Reading latest_fork_usec as the duration of the whole rewrite — it is the fork syscall alone; the child's lifetime is aof_current_rewrite_time_sec / aof_last_rewrite_time_sec.
  • Assuming BGREWRITEAOF always starts a rewrite immediately, and never checking aof_rewrite_scheduled when nothing appears to happen.
  • Recommending no-appendfsync-on-rewrite yes as a free latency win, without saying it suspends the parent's fsyncs for the entire rewrite and stretches the loss window from one second to minutes.
  • Treating aof_rewrite_in_progress:0 as proof the last rewrite succeeded, instead of reading aof_last_bgrewrite_status and aof_rewrites_consecutive_failures.
  • Claiming an AOF rewrite costs exactly what a BGSAVE costs "because it's the same fork", ignoring the longer child life and the parent writing to the same device throughout.

context