A team documents that their Redis instance can lose at most one second of writes because it runs with 'appendfsync everysec'. Under what circumstances is that claim wrong, and how would you detect it on a running instance?
answer
- everysec = 1s only if fsync finishes in <1s
- write deferred while fsync in flight, up to ~2s
- aof_delayed_fsync / aof_pending_fsync in INFO persistence
- no-appendfsync-on-rewrite yes → no fsync during rewrite
- stall sits outside command execution
basics
~20 sIf the background fsync is still running when the next write is due, Redis postpones the write for up to about two seconds rather than blocking, so the unsynced window can reach ~2s. And with no-appendfsync-on-rewrite yes, fsyncs stop entirely during a rewrite. Check aof_delayed_fsync in INFO persistence.
solid answer
~50 sTwo things break the one-second claim. **Slow disks.** With `everysec` the fsync runs in a background thread. Before each `write()`, Redis checks whether a previous fsync is still in flight; if it is, Redis **defers the write** rather than blocking command processing on a busy device. It will defer for up to about two seconds, then write anyway (now potentially blocking). So on storage that cannot complete an fsync within a second, the unsynced window stretches toward two seconds — and the deferral itself shows up as a latency spike. **Rewrites.** `no-appendfsync-on-rewrite yes` suppresses fsyncs completely while a background AOF rewrite is running, to keep two processes from hammering the same disk. For the duration, the instance behaves like `appendfsync no` — potentially tens of seconds at risk. Detection: `INFO persistence` → `aof_delayed_fsync` counts deferrals; a rising counter means the disk is behind. `aof_pending_fsync` shows in-flight work, and the latency monitor reports `aof-write` / `aof-fsync-always` events.
code
text · 18 lines> INFO persistence
aof_enabled:1
aof_rewrite_in_progress:0
aof_last_write_status:ok
aof_pending_fsync:1
aof_delayed_fsync:2841 # <-- writes postponed; window > 1s
> CONFIG GET no-appendfsync-on-rewrite
1) "no-appendfsync-on-rewrite"
2) "yes" # <-- durability suspended during rewrites
# Attribute the stalls
> CONFIG SET latency-monitor-threshold 100
> LATENCY LATEST
1) 1) "aof-write-pending-fsync"
2) (integer) 1786000123
3) (integer) 412 # last event, ms
4) (integer) 980 # max event, msgo deeper
It is enough to know that everysec is approximate: on a slow disk the window can be larger, and there is a counter in INFO persistence that shows it.
Explain the deferral mechanism — Redis will not block on write() while an fsync is in flight, so it postpones for up to about two seconds — and name aof_delayed_fsync.
Diagnose end to end: correlate aof_delayed_fsync and latency-monitor aof-* events with device utilisation, notice that the slow-command log stays empty, and identify no-appendfsync-on-rewrite as a silent durability downgrade.
Refuse to state a durability number that the storage stack cannot back: verify fsync honesty, budget I/O for the rewrite child, and if the recovery-point objective is real, move the guarantee to replica acknowledgement or an upstream system of record rather than a single node's fsync.
## Where the "one second" number comes from With `appendfsync everysec`, Redis writes buffered AOF data to the OS with `write()` on each event-loop cycle, and a **background thread** issues `fsync()` roughly once per second. The nominal exposure to a machine crash is therefore the data written since the last completed fsync — about one second. That is the number everybody quotes, and under healthy conditions it is right. ## Failure mode 1: the disk cannot keep up An fsync is not instantaneous. On network-attached or heavily contended storage it can take hundreds of milliseconds, occasionally seconds. Redis has to decide what to do when it is time to flush the buffer with `write()` but the previous `fsync()` has not finished. Blocking on `write()` would be catastrophic: the call can stall behind the in-flight sync, and because commands are processed serially, every client would stall with it. So Redis takes the other option: it **postpones the write**, keeps the data in its own buffer, and carries on serving. It records the moment it started postponing. If the fsync still has not finished after roughly **two seconds**, Redis gives up waiting and performs the write anyway — at which point it may genuinely block. Two consequences follow, and both contradict the naive claim: - **The durability window widens.** Data that has not even reached the page cache certainly has not been fsynced, so the crash exposure grows to roughly two seconds rather than one. - **Latency spikes appear from nowhere.** A command completes in microseconds but the reply is delayed because the event-loop cycle stalled on a write that finally went through. Note that this stall happens *outside* command execution, so the slow-command log does not attribute it to any command — a classic "Redis is slow but nothing shows up as slow" investigation. ## Failure mode 2: rewrites suppress fsync entirely A background AOF rewrite is a second process writing a fresh copy of the dataset to the same disk. On spinning disks and on contended cloud volumes, that competition can make the parent's fsyncs pathologically slow. The escape valve is `no-appendfsync-on-rewrite`: - `no` (default): keep fsyncing per policy during the rewrite. Durability holds; latency may suffer. - `yes`: **do not fsync at all** while a rewrite (or a snapshot child) is running. Latency stays smooth; for that entire window the instance is effectively running `appendfsync no`, with a loss exposure of however long the kernel takes to flush — tens of seconds. Many "tuned for latency" configurations set this to `yes` and then describe the instance as "one second durable", which is only true when no rewrite is in progress. If rewrites happen often — high write volume, aggressive growth thresholds — the unprotected fraction of the day can be substantial. ## Failure mode 3: the device lies fsync's guarantee is only as strong as the storage stack honouring it. Virtual disks with volatile write-back caches, some container storage layers, and misconfigured RAID controllers can acknowledge an fsync before data is on stable media. Redis has no way to detect this; it is a platform property you verify out of band. Say so explicitly when someone asks for a hard durability number. ## Detecting the problem on a live instance `INFO persistence` is the primary source: - **`aof_delayed_fsync`** — a cumulative counter of how many times Redis postponed a write because an fsync was still running. Zero is healthy. A number that grows during traffic peaks is direct proof the disk is behind the policy. - **`aof_pending_fsync`** — fsyncs currently outstanding. - **`aof_last_write_status`** — `err` means writes to the AOF are failing outright, which is an incident in itself (Redis will start rejecting writes). - **`aof_rewrite_in_progress`** — combine with `no-appendfsync-on-rewrite yes` to know when durability is suspended. The latency monitor (`CONFIG SET latency-monitor-threshold <ms>`, then `LATENCY LATEST` / `LATENCY HISTORY`) reports dedicated events such as `aof-write`, `aof-write-pending-fsync`, `aof-write-active-child` and `aof-fsync-always`, which distinguish "slow because a child is running" from "slow because a sync is pending". At the OS level, elevated `%util` and await times on the AOF's device, or a shared volume's IOPS credits being exhausted, corroborate the diagnosis. ## Remediation 1. **Give the AOF faster or dedicated storage** — local NVMe rather than a shared network volume, or at least separate the AOF device from the snapshot device. 2. **Reduce rewrite pressure** so the competing child runs less often. 3. **Decide the trade honestly**: keep `no-appendfsync-on-rewrite no` and accept latency during rewrites, or set it to `yes` and document that durability is suspended for that window. 4. **Move the guarantee elsewhere** if the number really matters — an upstream durable system of record, or acknowledgement from replicas — rather than asking a single node's fsync to carry it. ## The answer in one breath "Everysec means one second only when the disk can finish an fsync in under a second. Redis postpones writes while a sync is in flight, up to about two seconds, so the real window is up to ~2s — and with `no-appendfsync-on-rewrite yes` there is no fsync at all during a rewrite. `aof_delayed_fsync` in INFO persistence is the tell."
- Why does Redis postpone the write instead of just calling write() and letting it block?Because write() on a file with an fsync in flight can stall for as long as the sync takes, and commands are processed serially, so that stall becomes latency for every connected client. Postponing keeps the data in Redis's own buffer and lets the server keep serving traffic. The compromise is a bounded postponement of about two seconds, after which correctness of the durability promise wins and Redis writes anyway.
- The instance shows occasional 500 ms client latency but the slow-command log is empty. How does that fit?The slow-command log measures time spent executing a command, whereas an AOF write or fsync stall happens in the event-loop cycle around command execution. So the server really was blocked, but no individual command owns the time. The latency monitor's aof-write and aof-write-pending-fsync events, together with a rising aof_delayed_fsync, are what attribute the stall correctly.
saying these in an interview costs you the question
- "everysec always means exactly one second, by design" — it means one second only when fsync completes in under a second.
- "aof_delayed_fsync counts lost writes" — it counts postponed writes, an indicator of disk pressure, not data loss.
- "no-appendfsync-on-rewrite yes is a pure latency win" — it suspends durability entirely for the rewrite window.
- "If nothing shows in the slow-command log, Redis wasn't blocked" — AOF stalls occur outside command execution.
- "Using a cloud volume with high IOPS makes fsync latency a non-issue" — burst credits, network hops and volatile caches all still bite.