skip to content

A client sends a write to Redis and receives an OK reply. What has actually been guaranteed at that instant, and what would you have to add for "this write is safe" to hold end-to-end — through the Redis process, the operating system page cache, the storage device, and other nodes?

level: seniorimportance: must knowfreq 50%

answer

  1. OK = applied in master memory, nothing more
  2. Ladder: process → page cache → device → other node
  3. Device cache can ack an fsync that never landed
  4. Async replication: failover loses acked writes even with always
  5. WAIT / WAITAOF = post-hoc confirmation, no rollback

basics

~20 s

An OK from Redis means only that the command was applied in the master's memory. Safety is a ladder: process memory, OS page cache, the device (whose own cache can lie), another node. The WAIT and WAITAOF commands confirm replication or fsync after the fact, not synchronously.

solid answer

~50 s

**An OK is a statement about memory, not about disk or about the cluster.** Redis applies the command to the in-memory dataset and replies; only `appendfsync always` inserts an fsync before that reply, and even then the promise stops at one machine. Think in four failure domains: Redis process memory (gone on any crash), the OS page cache (survives `kill -9`, not a power cut or kernel panic), the storage device (survives a power cut only if the device is honest — a volatile disk or RAID-controller cache can acknowledge an fsync that never reached flash), and another node (the only rung that survives losing the machine). Replication is asynchronous, so a failover discards acknowledged writes regardless of `appendfsync`. `WAIT numreplicas timeout` and `WAITAOF numlocal numreplicas timeout` (Redis 7.0+) let the client confirm afterwards that replicas received, or fsynced, the writes it already issued — post-hoc, with no rollback, and each costs a round trip or an fsync on the write path.

code

text · 10 lines
text
# Write, then ask (after the fact) how many replicas took it
redis-cli SET order:1042 paid
# OK            <- master memory only
redis-cli WAIT 2 1000
# (integer) 1   <- only ONE replica acked within 1000 ms; nothing rolled back

# Redis 7.0+: confirm fsync locally AND on replicas
redis-cli WAITAOF 1 1 1000
# 1) (integer) 1   local AOF fsynced
# 2) (integer) 0   no replica fsynced in time -> your call what to do now

go deeper

for a junior

Recall the core fact: Redis answers from memory, so an OK does not mean the data is on disk, and a crash can lose recent writes. Knowing that persistence must be configured and that a replica exists for machine-level failure is enough.

for a middle

Separate the failure domains cleanly — process crash versus power loss versus losing the machine — and explain which rung (page cache, fsync, replica) covers each. Know that replication is asynchronous and that WAIT exists.

for a senior

Lead with what the acknowledgement does and does not promise, walk the ladder including dishonest device caches, and explain the failover hole precisely: promoted replica, discarded tail, old master resyncing. Describe WAIT/WAITAOF as post-hoc confirmation with a per-write latency cost and say how you would apply them selectively.

for a principal

Frame it as choosing an RPO per class of data and paying for it deliberately: which writes get a confirmation round trip, which stop at the page cache, and which belong in a store with real synchronous commit rather than in Redis at all. Address correlated failure of the acknowledging nodes, the availability cost of min-replicas gates, and how the choice is enforced and measured rather than merely documented.

## What the OK reply actually promises When Redis returns `+OK`, one thing has certainly happened: the command was executed against the in-memory dataset on the node that answered, and the effect is immediately visible to every other client reading that node. Nothing else is implied. With AOF enabled, the command has additionally been appended to an in-memory buffer that will be handed to the kernel at the end of the event-loop iteration. Only the strictest fsync policy makes Redis force that buffer to stable storage *before* the reply leaves. The exact loss window that each `appendfsync` policy and each RDB save point produces is a separate subject, owned by the persistence-configuration material; the point here is that visibility and durability are different properties, and the reply confirms only the first. There is also a rung below Redis: the client. If the connection breaks after the server executed but before the reply arrived, the client cannot tell whether the write happened. "Safe end-to-end" therefore starts with idempotent writes or a retry key, not with disk settings. ## The failure-domain ladder Each rung survives a strictly larger class of failure, and each costs more. **1. Redis process memory.** Survives nothing. `kill -9`, a segfault, the Linux OOM killer, or a container restart takes it. This is where an ordinary acknowledged write lives. **2. The OS page cache.** Reached when Redis calls `write()` on the AOF. The kernel now owns those bytes, so the data survives the Redis process dying — the kernel flushes later on its own schedule. It does not survive the kernel dying: a power cut, a panic, a hypervisor loss, or a cloud instance being terminated destroys page cache and process memory alike. **3. The storage device.** Reached when `fsync()` returns. This is the rung people believe is absolute, and it is the one that lies most often. A consumer SSD, a desktop disk with write-back caching enabled, or a RAID controller without a battery/flash-backed cache can acknowledge an fsync while the data still sits in the device's own volatile DRAM; a power cut then loses writes the operating system believes are durable. Enterprise drives with power-loss protection, or controllers with BBU/FBWC, are what make an fsync mean what you think it means. Network-attached cloud volumes add their own semantics and, importantly, their own latency — an fsync there can cost milliseconds, which is why `appendfsync always` collapses throughput on such storage. **4. Another node.** The only rung that survives losing the machine outright: destroyed disk, terminated instance, rack fire, or an availability zone going dark. No local fsync setting reaches it. The practical skill is naming which failure class you are defending against and stopping at the rung that covers it — a session cache legitimately stops at rung 1. ## The failover hole Redis replication is asynchronous by design: the master applies the command, replies to the client, and *then* streams it to replicas. So if the master dies and Sentinel or Cluster promotes a replica, everything the master had acknowledged but not yet shipped is gone — even if it was fsynced locally, because the promoted replica never received it. Worse, when the old master comes back it becomes a replica of the new one and resynchronises, discarding its own divergent tail. This is why "we set `appendfsync always`" is not an answer to "can we lose an acknowledged write": with automatic failover, the answer stays yes. `min-replicas-to-write` and `min-replicas-max-lag` reduce the exposure by making the master *refuse* writes when too few replicas are sufficiently caught up. That is an availability-for-durability trade, not synchronous commit: the writes it does accept are still acknowledged before replication. ## WAIT and WAITAOF: confirmation, not commit `WAIT numreplicas timeout` blocks the calling client until at least `numreplicas` replicas have acknowledged all write commands previously sent on that connection, and returns how many actually did. `WAITAOF numlocal numreplicas timeout`, added in Redis 7.0, blocks until those writes have been fsynced to the AOF on the local node and/or on that many replicas, returning both counts; it requires AOF to be enabled on the nodes being counted. Their shared limitation is that they are **post-hoc**. The write was applied and became readable before you called them. If the command returns a count lower than you asked for, nothing rolls back — you have an applied, visible, possibly non-durable write and must decide what to do (retry, alert, or fail the request while knowing the data may still be there). They are also not proof against correlated failure: if the acknowledging nodes share a rack, a power feed, or an availability zone, they can all disappear together. The cost is real. Waiting for a replica adds at least one replication round trip to the write path; `WAITAOF` additionally waits on a real fsync, which under a once-per-second policy can mean waiting for the next fsync tick. Applying either to every write turns Redis into a much slower system, so the usual pattern is to apply it selectively — to the small set of writes that carry value, or once at a batch or transaction boundary rather than per command. ## How to say it in an interview Start with "an OK is a memory-level acknowledgement", walk the four domains, then deliver the punchline: asynchronous replication means a failover can lose acknowledged writes no matter how aggressively you fsync, and `WAIT`/`WAITAOF` buy you after-the-fact confirmation with a latency bill, not a synchronous commit. Finish by naming the failure class you would actually design for.

  • Why does a Sentinel failover still lose acknowledged writes even when the master ran with the strictest appendfsync setting?
    Because fsync and replication are independent paths. The strictest policy guarantees the write reached the master's own disk before the reply, but the master streams to replicas asynchronously afterwards. When the master dies, the promoted replica never received the tail, so it starts serving without it; when the old master returns it resynchronises from the new one and discards its own divergent data. Local durability protects against that node restarting, not against that node being replaced.
  • WAIT 2 1000 returns 1. What do you do?
    Recognise that the write is already applied and readable — there is no rollback available. You decide at the application level: retry the same idempotent write, fail the user-visible operation while flagging that the data may persist anyway, or degrade to a slower durable store for that record. Operationally a short count means replicas are lagging or down, so it should also raise an alert; treating WAIT's return value as advisory and ignoring it defeats the point of calling it.
  • When is stopping at the lowest rung — no persistence at all — the right engineering answer?
    When the dataset is derivable from a system of record and the cost of loss is a cache miss rather than a wrong answer: session lookaside caches, rate-limit counters, computed leaderboards, denormalised read models. The honest caveat is the secondary effect: an empty Redis after a restart sends full traffic to the origin, so you need request coalescing, staged warming, or origin capacity for the cold-start herd. Saying this confidently is stronger than bolting on persistence you never validated.

Handing a letter to the office mailroom protects it if you personally vanish, but not if the building burns; only a copy already in another city survives that. WAIT is phoning the other city afterwards to ask whether the copy arrived — useful, but the original letter has already been read by then.

saying these in an interview costs you the question

  • Saying an OK reply means the data is on disk — it means the command was applied in memory
  • Believing appendfsync always makes acknowledged writes failover-safe; replication is asynchronous either way
  • Treating fsync as absolute and ignoring volatile device/controller caches that acknowledge without persisting
  • Describing WAIT or WAITAOF as synchronous replication or two-phase commit — they confirm after the write is already applied and visible, and never roll back
  • Ignoring WAIT's return value, or calling it on every write and then being surprised by the throughput collapse

context