After switching a Redis instance to 'appendfsync always', p99 latency for all clients jumps and occasionally the whole instance appears to freeze for tens of milliseconds. How do you confirm the AOF sync path is the cause, and what are your options?
answer
- always = fsync in the serving path → global latency
- slow-command log clean; stall is outside execution
- LATENCY LATEST → aof-fsync-always / aof-write
- INFO persistence: pending/delayed fsync, write status
- fixes: NVMe, everysec, split dataset, replica acks
basics
~20 sWith always, an fsync runs in the serving path each event-loop cycle, so every client pays the device's sync latency. Confirm with the latency monitor's aof-fsync-always and aof-write events plus INFO persistence, and device I/O stats. Options: faster local storage, back to everysec, or move the guarantee to replicas.
solid answer
~50 s**Why it happens.** `appendfsync always` performs the fsync inside the event-loop cycle, before replies are sent. Commands are processed serially, so the device's sync latency is added to everyone's response time, not just the writer's. On network-attached storage an fsync of a few milliseconds turns into a few milliseconds of instance-wide stall, repeatedly. **Confirming it.** Enable the latency monitor (`CONFIG SET latency-monitor-threshold 20`) and read `LATENCY LATEST` / `LATENCY HISTORY`; the `aof-fsync-always` and `aof-write` events attribute the stalls directly. Cross-check `INFO persistence` (`aof_last_write_status`, `aof_pending_fsync`, `aof_delayed_fsync`) and `INFO commandstats`. Note the slow-command log will look innocent: the stall is in the event-loop cycle, not inside any command's execution. Confirm at the OS level with device await/utilisation, and remember `latency-tracking`/`LATENCY HISTOGRAM` reports per-command percentiles. **Options.** Move the AOF to fast local NVMe; drop to `everysec` and accept ~1s exposure; keep `always` only on a node whose throughput is low enough; or obtain durability from replica acknowledgement instead of a single node's disk.
code
text · 18 lines> CONFIG SET latency-monitor-threshold 20 # record events >= 20ms
> LATENCY LATEST
1) 1) "aof-fsync-always"
2) (integer) 1786001200 # last event timestamp
3) (integer) 34 # last event duration, ms
4) (integer) 91 # max duration seen, ms
> LATENCY HISTORY aof-fsync-always
> LATENCY DOCTOR # prose diagnosis
> INFO persistence
aof_last_write_status:ok
aof_pending_fsync:0
aof_delayed_fsync:0 # 'always' does not defer; it blocks
# The trade, stated explicitly:
> CONFIG SET appendfsync everysec # fsync leaves the serving path; ~1s at risk
> CONFIG REWRITE # make it survive restartgo deeper
Know the shape of the answer: always makes Redis wait for the disk before replying, and because Redis handles one command at a time, everybody waits.
Add the evidence path — latency-monitor events for the AOF, INFO persistence fields — and the obvious remedy of returning to everysec with its one-second exposure.
Run the full diagnosis: rule out command execution, attribute with LATENCY LATEST/HISTORY, corroborate with device metrics and burst-credit behaviour, then present options with their durability consequences, including pipelining as amortisation.
Decide where the durability guarantee actually lives: split datasets by requirement, use replica acknowledgement or an upstream system of record for the strict subset, and treat a single node's fsync policy as one input to an availability-and-recovery budget rather than the guarantee itself.
## The mechanism you are diagnosing Under `appendfsync always`, each event-loop cycle that produced writes performs the sequence: `write()` the AOF buffer to the OS, then `fsync()` the file, and only then send the replies for those commands. Because Redis executes commands one at a time, the fsync is not "the writer's problem" — it occupies the same execution path that every other client's command must pass through. A 3 ms fsync repeated hundreds of times a second is hundreds of milliseconds per second of the instance doing nothing but waiting on the device. This is why the symptom is *global*: read-only clients that never wrote anything see their p99 rise, which is often what confuses the initial investigation. ## Step 1 — establish that the stall is not command execution The first surprising observation is that the slow-command log is usually clean. That log records time spent *executing* a command; the AOF write and fsync happen in the surrounding event-loop cycle. So "latency is terrible but nothing is slow" is itself a strong clue pointing at the I/O path (or at another out-of-band stall). ## Step 2 — attribute the stall with the latency monitor Redis has a built-in latency monitor with named events. Enable it and look: ``` CONFIG SET latency-monitor-threshold 20 # ms LATENCY LATEST LATENCY HISTORY aof-fsync-always LATENCY DOCTOR ``` The events that matter here are `aof-fsync-always` (time spent in the synchronous fsync), `aof-write` (the write syscall itself), `aof-write-pending-fsync` and `aof-write-active-child` (write slowed because a sync or a background child was in progress), and `aof-rewrite-diff-write`. `LATENCY LATEST` gives the last and maximum durations per event, which is usually enough to close the question in one command. `LATENCY DOCTOR` will even name the likely cause in prose. ## Step 3 — corroborate with INFO and the OS `INFO persistence` supplies `aof_last_write_status` (an `err` here means writes are failing, at which point Redis starts refusing writes — a different and worse incident), `aof_pending_fsync`, and `aof_delayed_fsync`. `INFO stats` gives `instantaneous_ops_per_sec` to correlate stalls with load; `latency-tracking` plus `LATENCY HISTOGRAM` provides per-command percentile histograms so you can show that *all* commands, not one family, degraded at once. At the OS level, look at the AOF device's average service time and utilisation, and at whether the volume is a shared/network device subject to throttling or burst-credit exhaustion. Instances that behaved fine for days and then degraded are a classic credit-exhaustion signature. ## Step 4 — the options, with their trade-offs 1. **Faster storage.** Local NVMe rather than a network volume is the single most effective change. If the AOF and the snapshot output share a device, separate them so a background child is not competing with the sync path. 2. **Move to `everysec`.** This takes the fsync off the serving path entirely (a background thread does it) at the cost of about one second of exposure on a machine crash. For the overwhelming majority of workloads this is the right answer, and it is the default for exactly this reason. 3. **Keep `always` but shrink what it must protect.** If only a small subset of the data genuinely needs per-write durability, put that subset on its own low-throughput instance running `always`, and let the high-throughput instance run `everysec`. Durability policy becomes a per-dataset decision rather than a per-process one. 4. **Relocate the guarantee.** Node-level fsync protects one machine; asynchronous replication means a failover can still lose acknowledged writes. If the real requirement is "this write must not vanish", replica acknowledgement (for instance, requiring propagation to N replicas before treating the write as committed) or an upstream durable system of record addresses it more honestly than any fsync setting — and often lets you go back to `everysec`. 5. **Amortise with pipelining.** One fsync covers everything written in that cycle, so a client that pipelines batches of writes pays far fewer syncs per write than one issuing them singly. Where the client can batch, throughput under `always` improves substantially without changing durability. ## What not to do Do not set `no-appendfsync-on-rewrite yes` as a fix for this. It will smooth the graph during rewrites, but it does so by suspending durability for the whole rewrite window — which is exactly the property you switched to `always` to obtain. Choosing latency is legitimate; choosing it silently while still claiming per-write durability is not. ## How to present it in an interview Walk the chain: the mechanism (fsync in the serving path, serial command processing, so global impact), the evidence (latency monitor `aof-fsync-always`, INFO persistence, device stats, and the notable absence of slow commands), and the decision (faster disk, `everysec`, split the dataset, or move the guarantee to replicas) with the durability consequence of each stated out loud.
- Why do read-only clients also see higher latency when appendfsync always is enabled?Because Redis processes commands serially, so the fsync performed for other clients' writes occupies the same execution path a read must traverse. A read that takes microseconds can still wait behind a multi-millisecond sync. The impact is a property of the shared execution path, not of the individual command, which is why the whole instance degrades together.
- The team insists no acknowledged write may ever be lost. Is appendfsync always sufficient?It is necessary at the node level but not sufficient for the system. Replication is asynchronous, so if the master dies after fsyncing a write but before shipping it, a promoted replica will not have it, and the acknowledged write is gone. Meeting the requirement means acknowledging writes only after they reach a quorum of replicas, or keeping an upstream durable system of record, with node-level fsync as one component.
- Would client-side pipelining help throughput under appendfsync always?Yes, substantially. A single fsync covers everything written during that event-loop cycle, so a batch of pipelined writes amortises one sync across many commands instead of paying one per round trip. It does not weaken durability, since replies are still sent only after the sync. The limit is that batching adds queueing latency for the individual command.
saying these in an interview costs you the question
- "Only write commands get slower with appendfsync always" — serial execution makes it an instance-wide latency tax.
- "Nothing in the slow-command log, so Redis wasn't blocked" — the stall is in the event-loop cycle, outside command execution.
- "Set no-appendfsync-on-rewrite yes to fix always latency" — that silently suspends the durability you enabled.
- "appendfsync always means no data can ever be lost" — asynchronous replication can still lose acknowledged writes on failover.
- "Provisioned-IOPS network storage removes fsync latency" — network round trips, throttling and burst credits all still apply.