Redis exposes three values for the appendfsync setting: always, everysec and no. What does each one do, and how much data does each risk losing when the process or the machine crashes?
answer
- write() = survives Redis crash; fsync() = survives power loss
- always: fsync before reply, ~0 loss, slowest
- everysec: bg thread, ~1s (worst ~2s), default
- no: kernel decides, ~30s on Linux, fastest
- no-appendfsync-on-rewrite turns everysec into no
basics
~20 salways fsyncs before replying, so an acknowledged write survives a crash, at a big throughput cost. everysec (the default) fsyncs once per second in a background thread, risking about a second of writes. no never fsyncs — the OS flushes when it likes, so a machine crash can lose tens of seconds.
solid answer
~60 sThe AOF write path has two stages: Redis buffers commands in memory and calls `write()` once per event-loop cycle to hand them to the **OS page cache**, then `fsync()` forces the page cache to the physical device. `appendfsync` controls only the second stage. - **always** — fsync on every event-loop cycle, before replies go back to clients. A write that was acknowledged is on disk, so a crash of the Redis process *or* the machine loses nothing on that node. Cost: a disk sync in the hot path, cutting write throughput by an order of magnitude on non-optimal storage. - **everysec** (default) — a background thread fsyncs once per second. A **Redis process** crash loses only data still in the in-memory buffer; a **machine/power** failure loses up to ~1 second (worst case ~2, when a slow fsync forces Redis to delay the write). - **no** — Redis never calls fsync; the kernel flushes on its own schedule (typically up to ~30s on Linux). A process crash still loses nothing beyond the buffer, but a power loss can discard tens of seconds. The key distinction: `write()` survives a Redis crash, `fsync()` survives a machine crash.
code
text · 19 lines# redis.conf
appendonly yes
appendfsync everysec # always | everysec | no
no-appendfsync-on-rewrite no # yes => no fsync at all during a rewrite
# Inspect at runtime
> CONFIG GET appendfsync
1) "appendfsync"
2) "everysec"
> INFO persistence
aof_enabled:1
aof_last_write_status:ok
aof_pending_fsync:0
aof_delayed_fsync:17 # fsync still running when a write was due
# Runtime change is NOT persistent by itself:
> CONFIG SET appendfsync always
> CONFIG REWRITE # <- required to survive restartgo deeper
Name the three values and their rough loss windows: always ≈ nothing, everysec ≈ one second (the default), no ≈ whatever the OS decides.
Explain the two-stage path — buffer, write() to page cache, fsync() to device — and use it to distinguish process-crash loss from machine-crash loss for each policy.
Add the everysec two-second worst case and aof_delayed_fsync, the throughput cost of always on network storage, and the way no-appendfsync-on-rewrite silently downgrades durability during a rewrite.
Turn it into an RPO decision: pick the policy from the tolerable loss window and storage characteristics, state explicitly that node-level fsync does not give cluster-level durability under asynchronous replication, and decide where the real guarantee lives (quorum, WAIT, or an upstream system of record).
## The two-stage write path Understanding the policies requires understanding what happens to a logged command: 1. **Redis buffer.** After a write command executes, its normalised form is appended to an in-memory AOF buffer inside the Redis process. 2. **`write()` to the OS.** Once per event-loop cycle (before the server goes back to waiting for I/O), Redis flushes that buffer with a `write()` syscall. The bytes are now in the **kernel page cache**. They are no longer lost if the *Redis process* dies — the kernel owns them and will write them out — but they are not on the physical device yet. 3. **`fsync()` to the device.** Only `fsync()` tells the kernel to push those pages to durable storage and waits for the device to confirm. `appendfsync` governs step 3 exclusively. This is why "process crash" and "machine crash" are two different questions and why a candidate who conflates them gets the loss windows wrong. ## appendfsync always Redis performs the `fsync()` in the same event-loop cycle as the `write()`, **before** replies for those commands are sent to the clients. So when your client receives `+OK`, the command is durable on that node: neither `kill -9` nor a power cut loses it. The cost is that every batch of writes pays the storage device's sync latency, in the path that also serves every other client. On a consumer SSD an fsync is often a fraction of a millisecond; on network-attached storage it can be several milliseconds. Throughput commonly drops by an order of magnitude compared with `everysec`. Two nuances worth stating: the fsync covers everything in the cycle, so under heavy pipelining many commands amortise one sync; and "durable" means durable *on that machine* — it says nothing about replicas, which are updated asynchronously. ## appendfsync everysec — the default Redis calls `write()` every cycle as usual, but delegates `fsync()` to a **background thread** that runs roughly once per second. Command processing never waits for the device in the normal case. Loss windows: - **Redis process crashes** (bug, OOM kill, `kill -9`): only whatever is still sitting in the in-memory AOF buffer is lost — typically sub-millisecond of writes, because the buffer is flushed with `write()` every cycle. The kernel still holds and will persist everything already written. - **Machine crash or power loss**: up to about one second of writes, since that is the sync interval. The true worst case is closer to **two seconds**: if a previous fsync is still in flight when the next cycle wants to write, Redis postpones the `write()` rather than blocking on a busy disk, and only after ~2 seconds does it write anyway. This is the default because it is the knee of the curve: near-`no` throughput with a bounded, small, well-understood loss window. ## appendfsync no Redis performs the `write()` and never calls `fsync()` at all (barring rewrite bookkeeping). Flushing is entirely the kernel's decision, driven by its dirty-page writeback settings — on Linux, commonly up to around 30 seconds. - **Redis process crash**: still only the in-memory buffer is lost, exactly as with `everysec`. This surprises people; the OS is the safety net. - **Machine crash**: potentially tens of seconds of writes. Throughput is the highest of the three, and latency is the smoothest, because nothing in Redis ever waits on the device. ## Choosing - `everysec` for almost everything. The default is the default for good reasons. - `always` when a single lost acknowledged write is genuinely unacceptable *and* the storage is fast enough — and even then, be explicit that this protects one node, not the cluster, because replication is asynchronous. - `no` when Redis holds regenerable data (a cache, derived state) and you want maximum throughput; the honest version of that choice is often "persistence is a warm-start optimisation, not a durability guarantee". ## Things that silently change the policy - `no-appendfsync-on-rewrite yes` suspends fsyncs entirely while a background rewrite is running, to avoid two processes hammering the same disk. During that window, an `everysec` instance effectively behaves like `no`. It is a deliberate latency-versus-durability trade you should know you have made. - `CONFIG SET appendfsync ...` applies immediately but is lost on restart unless you also `CONFIG REWRITE`. - Virtualised or network-attached storage may lie about, batch, or simply be slow at fsync; the policy is only as good as the device honouring it. ## Observability `INFO persistence` exposes `aof_last_write_status`, `aof_last_bgrewrite_status`, `aof_pending_fsync` and `aof_delayed_fsync` — the last one counts how often an fsync was still running when Redis wanted to write, and a growing value is direct evidence that the disk cannot keep up with the chosen policy. The latency monitor tracks `aof-write`, `aof-fsync-always` and related events. ## The one-line answer "`write()` protects against a Redis crash, `fsync()` protects against a machine crash. `always` fsyncs before the reply — zero acknowledged loss, big throughput cost. `everysec` fsyncs in a background thread — about a second at risk, up to two in the worst case. `no` leaves it to the kernel — tens of seconds at risk, fastest."
- With appendfsync no, how much do you lose if the Redis process is killed with kill -9 but the machine keeps running?Almost nothing — only whatever is still in the in-memory AOF buffer, which Redis flushes with write() every event-loop cycle. Once write() returns, the bytes live in the kernel page cache, and the kernel survives the process and will eventually write them out. The tens-of-seconds exposure of appendfsync no applies to a machine crash or power loss, not to a process crash.
- Does appendfsync always guarantee that an acknowledged write is never lost by the system as a whole?No — it guarantees the write is durable on that one node before the client is told OK. Redis replication is asynchronous, so if that node dies and a replica is promoted, writes the master had persisted but not yet shipped are gone from the new master's view. Node-level durability and cluster-level durability are separate problems, and WAIT or a quorum design addresses the second.
- Why is the worst-case loss under everysec sometimes described as two seconds rather than one?The background fsync can take longer than a second on a slow or busy disk. If a previous fsync is still in flight when the main thread is ready to write more data, Redis postpones the write rather than blocking command processing on the device — but only for up to about two seconds, after which it writes anyway. That postponement stretches the unsynced window beyond the nominal one second, and the aof_delayed_fsync counter in INFO persistence records each occurrence.
write() is dropping a letter into the office mailroom; fsync() is watching the van leave. Losing your desk (the process) after the mailroom has it is fine; losing the whole building (the machine) is not.
saying these in an interview costs you the question
- "With appendfsync no you lose 30 seconds if Redis crashes" — a process crash only loses the in-memory buffer; the 30s figure is for a machine/power failure.
- "everysec means Redis blocks for the fsync once a second" — the fsync runs in a background thread.
- "appendfsync always makes the write durable across the cluster" — replication is asynchronous; it is node-level durability only.
- "appendfsync controls how often Redis writes to the file" — write() happens every event-loop cycle regardless; the setting controls fsync only.
- "CONFIG SET appendfsync always is enough" — without CONFIG REWRITE the setting is lost at the next restart.