What does the Redis append-only file (AOF) actually record, and what does turning on 'appendonly yes' give you that keeping only periodic point-in-time snapshots does not?
answer
- log of write commands, RESP format
- appended AFTER execution
- SPOP→SREM, EXPIRE→PEXPIREAT (deterministic replay)
- grows with write volume, not dataset size
- recovery = re-execute the log
basics
~20 sThe AOF is a log of every write command that changed the dataset, appended after the command runs; recovery replays the commands. A snapshot only stores state as of the moment it was taken, so a crash loses everything since it — the AOF narrows that loss to a second or less.
solid answer
~50 sWith `appendonly yes`, every command that modifies the dataset is appended to a log file in the same RESP wire format clients use. Restart recovery means re-executing that log from the top, which rebuilds the dataset. Commands are logged **after** they execute, and in an *effect-normalised* form: non-deterministic commands are rewritten into deterministic ones, so `SPOP` becomes an `SREM` of the member actually popped and `EXPIRE` becomes `PEXPIREAT` with an absolute millisecond timestamp. Without that, a replay would produce a different dataset than the original run. The difference from snapshots is loss granularity. A snapshot captures the whole dataset at one instant, taken every few minutes, so a crash discards every write since the last one. The AOF is appended continuously, so with the default `appendfsync everysec` you lose roughly the last second. The costs are a bigger file, continuous disk writes, and slower startup — replay is re-executing every command.
code
text · 15 lines# redis.conf
appendonly yes
appenddirname "appendonlydir" # Redis 7.0+
# Client sends: SET user:1 ada
# AOF incremental file contains (RESP):
*3\r\n$3\r\nSET\r\n$6\r\nuser:1\r\n$3\r\nada\r\n
# Client sends: EXPIRE user:1 60
# AOF stores an ABSOLUTE deadline so replay is deterministic:
*3\r\n$9\r\nPEXPIREAT\r\n$6\r\nuser:1\r\n$13\r\n1786000000000\r\n
# Client sends: SPOP myset -> returned "c"
# AOF stores the concrete effect:
*3\r\n$4\r\nSREM\r\n$5\r\nmyset\r\n$1\r\nc\r\ngo deeper
Say what it is: a log of every write command, replayed on startup, giving a much smaller loss window than periodic snapshots.
Add that entries are RESP-encoded and appended after execution, and that non-deterministic commands are normalised (SPOP→SREM, EXPIRE→PEXPIREAT) so replay reproduces the same dataset.
Discuss the cost profile — growth proportional to write volume, continuous I/O, slower startup than a binary snapshot — and operational traps like CONFIG SET without CONFIG REWRITE, plus the 7.0 multi-part layout.
Frame it as choosing a recovery-point objective and paying for it in I/O and restart time, and connect the same effect-normalisation to the replication stream so that durability and replica consistency are reasoned about together.
## What Redis is trying to solve Redis serves data from memory. If the process dies, memory is gone. Persistence exists so a restarted instance can rebuild the dataset. Redis offers two mechanisms with quite different shapes: a **snapshot** of the whole dataset (RDB), and an **operation log** of everything that changed it (AOF). This leaf is about the log. ## What the AOF contains When `appendonly yes` is set, every command that **modifies** the dataset is appended to the append-only log. Read commands are never logged — replaying a `GET` would change nothing. The entries are written in the **same RESP protocol encoding clients use**, which is why an AOF is human-readable: `SET user:1 ada` appears as a small array of bulk strings. Recovery is conceptually just "connect a fake client and feed it the file": Redis starts empty and re-executes the whole log. ## Logged after execution, not before This is a genuinely important detail and a favourite follow-up. Redis appends a command to the AOF **after** it has been executed successfully against the in-memory dataset, not before. Consequences: - A command that fails (wrong type, syntax error) never reaches the log, so a replay never re-runs a failure. - The log describes effects that definitely happened in memory — but a crash between execution and the durable write means an acknowledged effect can be missing from the file. That is precisely what the fsync policy governs. ## Effect normalisation: turning non-deterministic commands into deterministic ones Some commands do not produce the same result twice. `SPOP` removes a random member. `INCRBYFLOAT` depends on the current value. `EXPIRE key 60` means "sixty seconds from *now*", and "now" during a replay hours later is a different instant. So Redis rewrites such commands into a deterministic equivalent before appending them: - `SPOP myset` → `SREM myset <the member that was actually popped>` - `EXPIRE key 60` / `SETEX` → `PEXPIREAT key <absolute unix ms>` - `INCRBYFLOAT k 1.1` → `SET k <resulting value>` - Commands with random or time-dependent behaviour inside scripts are handled by propagating their effects rather than the script call. Without this, replay would silently diverge from the original dataset — a different random member removed, a key that should have expired still alive. The same normalisation feeds the replication stream to replicas. ## AOF versus snapshots: the real difference is the loss window A snapshot is a complete copy of the dataset at one instant, taken on a schedule (say every 5 minutes or every N changes). If the process dies four minutes after the last snapshot, four minutes of writes are gone — no matter how fast the disk is, because nothing was written for them. The AOF is appended continuously, so the exposure is the time between a write happening and that write reaching disk durably. With the default `appendfsync everysec`, that is about one second. With `appendfsync always`, the fsync happens before the client is told the command succeeded, so an acknowledged write survives a crash of that process. A useful way to phrase the distinction: a snapshot answers "what did the dataset look like at time T", the AOF answers "what happened, in order". The log's granularity is per command; the snapshot's is per schedule. ## What the AOF costs - **Size.** The file grows with the *volume of writes*, not with the size of the dataset. A single counter incremented a million times produces a million log entries for one key. (Redis solves that with a periodic rewrite that compacts the log into the minimal set of commands needed to recreate the current dataset — its own topic.) - **Continuous I/O.** Every event-loop cycle writes buffered commands to the file, and fsyncs happen per policy. - **Startup time.** Loading an AOF means re-executing every command through the normal command path, which is slower than loading a compact binary snapshot of the same dataset. ## Where it lives Since Redis 7.0 the AOF is not a single file but a **multi-part set** — a base file plus incremental files, tracked by a manifest, all inside the directory named by `appenddirname` (default `appendonlydir`) under `dir`. Before 7.0 it was one file named by `appendfilename`. The concept is unchanged; only the on-disk layout differs. ## Enabling and disabling it `appendonly yes` in the config file, or at runtime `CONFIG SET appendonly yes` (which triggers an initial rewrite to create the base). `CONFIG SET` is not persistent unless you also `CONFIG REWRITE` the config file — a classic operational surprise where durability silently disappears at the next restart. ## The interview-ready summary "The AOF logs every write command in RESP format, appended after execution and normalised so replay is deterministic; recovery re-executes the log. Compared with periodic snapshots it shrinks the crash-loss window from minutes to about a second, at the cost of a larger file, continuous writes, and slower startup."
- Why does Redis rewrite SPOP as SREM in the AOF instead of logging SPOP itself?SPOP removes a random member, so replaying the literal command would very likely remove a different member than the original execution did, and the restored dataset would diverge from the one that crashed. Redis therefore logs the concrete effect — an SREM of the member actually removed. The same normalisation applies to time-dependent commands such as EXPIRE, which becomes PEXPIREAT with an absolute timestamp.
- Does the AOF grow with the size of the dataset or with something else?With the volume of write traffic. One key that is incremented a million times generates a million log entries even though the dataset holds a single value, whereas a large dataset written once produces a small log. That is why the file must be periodically compacted into a minimal set of commands that recreates the current state, otherwise it grows without bound.
- You ran CONFIG SET appendonly yes during an incident. Is the instance now durably configured?Only until it restarts. CONFIG SET changes the running server, not the config file, so the next restart comes back with appendonly no and no AOF at all. You must also run CONFIG REWRITE (or edit the config file) to make the change survive. This is a common way instances silently lose persistence weeks after an incident.
A snapshot is a photograph of the room every five minutes; the AOF is a running list of every object moved. The photo restores the room as of the photo; the list can rebuild it move by move.
saying these in an interview costs you the question
- "The AOF stores the data / the values" — it stores the write commands, and the data is a replay of them.
- "Commands are written to the AOF before they execute, like a write-ahead log" — Redis appends after execution.
- "EXPIRE is logged verbatim" — it is normalised to PEXPIREAT with an absolute timestamp so replay is deterministic.
- "AOF file size tracks dataset size" — it tracks write volume; a tiny dataset can produce a huge log.
- "Turning on AOF makes writes durable immediately" — durability depends entirely on the fsync policy.