skip to content

On Redis 7.x the append-only file is a directory (default `appendonlydir`) containing a `.manifest` file, a base file, and one or more `.incr.aof` files. As the operator: what is the unit you must copy or restore, how does the server decide what to load at startup, where do you point `redis-check-aof`, and what breaks if a script copies or truncates one file inside that directory on its own?

level: seniorimportance: should knowfreq 24%

answer

  1. directory is the unit, manifest is the authority
  2. base.rdb = preamble on, base.aof = off
  3. load order: base then incr, in manifest order
  4. redis-check-aof takes the manifest path
  5. only last-incr tail truncation survives

basics

~20 s

The copy unit is the whole directory — manifest plus the base file plus every incr file it names, as one consistent set. Startup reads the manifest and loads base first, then incr files in manifest order. Point redis-check-aof at the manifest. A single file copied or truncated in isolation gives an unloadable or silently short dataset.

solid answer

~60 s

**The backup unit is the directory, not a file.** The manifest is an index; it names one base file and the incr files that follow it. A copy is only restorable if the manifest and every member it names travel together. **Base extension tells you the format.** With `aof-use-rdb-preamble yes` (the default) the base is `*.base.rdb`, a binary snapshot; with it `no` the base is `*.base.aof`, a command log. Either way the incr files are RESP commands. **Startup:** Redis reads the manifest, loads the base with the matching loader, then replays each incr file in manifest order. Missing or extra files are not discovered by scanning the directory — the manifest is the authority. **Verification:** on 7.x `redis-check-aof` takes the manifest path, not a member. **Partial handling breaks it:** a manifest naming a file you didn't copy aborts the load; a truncated base or a truncated middle incr aborts too; only a truncated tail of the *last* incr is tolerated, and only with `aof-load-truncated yes` — which silently loses the trailing writes. Pre-7.0 the same base+tail lived concatenated in one `appendonly.aof`, which is why old `cp appendonly.aof` scripts must be rewritten.

code

text · 8 lines
text
$ ls appendonlydir/
appendonly.aof.1.base.rdb
appendonly.aof.1.incr.aof
appendonly.aof.manifest

$ cat appendonlydir/appendonly.aof.manifest
file appendonly.aof.1.base.rdb seq 1 type b
file appendonly.aof.1.incr.aof seq 1 type i

go deeper

for a junior

Recall that on Redis 7.x the AOF is a directory, that the manifest lists which files belong to it, and that you copy the whole directory rather than any single file.

for a middle

Explain the roles — manifest indexes, base holds the dataset at last rewrite, incr holds the tail — and that .base.rdb vs .base.aof reflects aof-use-rdb-preamble. State that startup loads base then incr in manifest order.

for a senior

Own the operational failure modes: manifest/member mismatch aborts startup, only the last incr's tail is tolerantly truncated (and that loses writes silently), redis-check-aof takes the manifest, and hand-deleting a member breaks the next restart. Mention that a concurrent rewrite can make a naive cp -r inconsistent.

for a principal

Frame it as a backup-consistency problem: the directory is a multi-file artifact with no atomic read, so either take a point-in-time filesystem snapshot, or use a BGSAVE RDB as the backup artifact and keep the AOF for crash recovery. Set restore verification (redis-check-aof on the manifest) and log-based alerting on truncated loads as policy rather than leaving it to whoever runs the restore.

## What is actually on disk On Redis 7.0 and later the append-only file is a directory placed next to `dir` and named by `appenddirname` (default `appendonlydir`). It holds three kinds of member: - **The manifest** (`appendonly.aof.manifest`) — a small text index. Each line names one member with a sequence number and a type: `b` for base, `i` for incr, `h` for history (superseded members awaiting deletion). - **The base file** — the dataset as of the last rewrite. Its extension is a direct readout of configuration: `*.base.rdb` when `aof-use-rdb-preamble yes` (the default), meaning the body is a binary snapshot; `*.base.aof` when the preamble is disabled, meaning the body is a plain RESP command log that rebuilds the dataset by replay. - **One or more incr files** (`*.incr.aof`) — the RESP command tail written since that base, appended under the configured `appendfsync` policy. Before 7.0 those same two sections lived concatenated inside a single `appendonly.aof` (snapshot head, command tail); the only thing that changed is the packaging and the fact that ordering is now explicit in a manifest rather than implicit in byte offsets. ## The copy/backup unit **The unit is the whole directory as one consistent set: manifest + base + every incr file the manifest names.** There is no valid partial copy. Two rules follow: 1. *Never copy or restore a single member.* Grabbing just the `.base.rdb` looks tempting — it is an RDB image — but it is the dataset as of the last rewrite, missing everything in the incr tail. Grabbing just an incr file gives you a command fragment with no starting state. 2. *The manifest and the members must agree.* If the manifest names a file that is absent from your copy, the server refuses to load. If you copy files but not the manifest, there is nothing to tell the server what the current set is. A subtlety worth stating in an interview: the directory is live. If a `BGREWRITEAOF` (or an auto-rewrite triggered by `auto-aof-rewrite-percentage`) completes while your `cp -r` is in flight, the old base and incr files can be deleted out from under you and the manifest swapped, leaving your copy internally inconsistent. Robust practice is a point-in-time filesystem or volume snapshot of the directory, or copying the manifest first and then each file it names and re-verifying afterwards; incr files only grow, so bytes appended after you read the manifest are harmless, but deletions are not. Many shops sidestep the problem for backup purposes entirely by taking a `BGSAVE` RDB, which is a single self-consistent artifact. ## How startup uses it Redis does **not** scan the directory and guess. It opens the manifest, loads the base with the loader matching its extension (RDB loader for `.base.rdb`, command replay for `.base.aof`), then replays each incr file in manifest order. Members present on disk but absent from the manifest are ignored as garbage; members named but missing are a fatal error. ## Verifying and repairing on 7.x `redis-check-aof` became manifest-aware: you point it at the manifest, not at a member. ``` redis-check-aof appendonlydir/appendonly.aof.manifest redis-check-aof --fix appendonlydir/appendonly.aof.manifest ``` It walks the whole set. `--fix` truncates the trailing partial command of the last incr file — which means it repairs exactly the one damage class Redis considers survivable, and it does so by discarding writes. ## What breaks with partial handling - **Missing member.** Manifest names a file you did not copy → load aborts at startup; the server exits with an AOF error rather than coming up with partial data. - **Truncated base.** A short `.base.rdb` fails its checksum, or a short `.base.aof` ends mid-command → hard failure. `aof-load-truncated` does not rescue the base. - **Truncated middle incr.** A hole in the middle of the sequence is a hard failure too; tolerance applies only to the tail of the final file. - **Truncated tail of the last incr.** With `aof-load-truncated yes` (default) Redis logs a warning, truncates to the last complete command and starts. This is the dangerous one: the server comes up *looking healthy* while the trailing writes are gone. Alert on the log line rather than trusting a clean start. - **Hand-deleting an "old-looking" incr file** to reclaim disk while the manifest still lists it turns a healthy server into one that cannot restart. The supported way to collapse the set is `BGREWRITEAOF`. - **Legacy scripts.** Anything doing `cp appendonly.aof` from the 6.x era silently backs up nothing on 7.x, because that path no longer exists. ## The line to remember The manifest is the authority, the directory is the artifact, and any operation that treats one member as standalone — copy, restore, truncate, delete — is a bug.

  • Your restore comes up cleanly but is missing the last few seconds of writes. What most likely happened, and where would you confirm it?
    Almost certainly the tail of the final incr file was truncated — an incomplete copy, a torn write, or `redis-check-aof --fix` having trimmed it. With the default `aof-load-truncated yes` Redis truncates to the last complete command, logs a warning, and starts normally, so the server looks healthy. Confirm in the startup log, which explicitly reports the truncated AOF; set `aof-load-truncated no` if you would rather fail loudly than start with silent data loss.
  • A disk-space alert fires and an engineer wants to delete the oldest-looking file in `appendonlydir`. What do you tell them?
    Nothing in that directory is safe to delete by hand while the manifest names it — removing a member makes the AOF unloadable on the next restart, even though the running server is unaffected until then. The supported way to shrink the set is `BGREWRITEAOF`, which produces a fresh base, starts a new incr file, and atomically swaps in a manifest that no longer references the old members, after which Redis removes them itself. Note that a rewrite transiently needs room for both the old and new base.
  • Why can the base file's extension differ between two servers, and does it change how you restore them?
    The extension reflects `aof-use-rdb-preamble`: `yes` (the default) produces `*.base.rdb`, a binary snapshot; `no` produces `*.base.aof`, a RESP command log replayed to rebuild the dataset. Restore procedure is identical either way — copy the directory intact and let the manifest drive loading — but the `.base.aof` variant reloads noticeably slower because it replays commands instead of loading a snapshot image.

saying these in an interview costs you the question

  • Saying you can back up or restore just the `.base.rdb`, since "it's an RDB anyway" — it omits the entire incr tail
  • Believing Redis scans the directory to find members instead of reading the manifest
  • Pointing `redis-check-aof` at an individual `.incr.aof` on a 7.x server instead of at the manifest
  • Assuming a clean startup proves nothing was lost, when `aof-load-truncated yes` silently trims the last incr file
  • Treating hand-deleting or renaming a member as a safe way to reclaim disk instead of running `BGREWRITEAOF`

context