When Redis documentation calls a MULTI/EXEC transaction atomic, exactly which guarantees does that cover — with respect to other clients, to a crash, and to replicas?
answer
- atomic = grouped and uninterrupted, not revertible
- single thread → no interleaving during EXEC
- durability = AOF/RDB settings, not EXEC
- replication async → committed EXEC can vanish on failover; WAIT narrows only
- big EXEC = latency spike for every client
basics
~20 sIt covers isolation only: EXEC runs the queued commands back to back on the single execution thread, so no other client sees a half-applied state. It does not cover durability — survival of a crash depends on your AOF/RDB settings — and replication is asynchronous, so a committed EXEC can be lost in a failover.
solid answer
~60 s"Atomic" here means **all commands run together, uninterrupted**. Redis executes commands on one thread and dispatches EXEC as a single unit, so no other client's command interleaves; every observer sees the state before the transaction or after it. That is genuine isolation, and it is the guarantee EXEC exists to provide. It stops there. **No rollback**: a command that fails at runtime does not undo its neighbours. **No durability of its own**: EXEC returning does not mean the data survives a crash — that depends entirely on persistence settings; with the default `appendfsync everysec` you can lose up to a second of writes including a committed transaction, and with RDB-only persistence you can lose much more. **No replication guarantee**: replication is asynchronous, so a primary can acknowledge EXEC and fail over before a replica received it, losing the whole transaction. `WAIT` lets you block until N replicas acknowledge, which narrows but does not close the window. Also note the cost side: EXEC occupies the execution thread for the whole batch, so a huge transaction is a latency spike for everyone.
code
text · 7 linesMULTI
INCRBY account:42 -100
INCRBY account:99 100
EXEC # applied in memory, replies immediately
WAIT 1 100 # block up to 100ms for 1 replica to ack; returns the count that did
# a return of 0 means no replica has it yet - a failover now would lose the transactiongo deeper
Say that atomic means the queued commands run together with nothing from other clients in between, and that it does not mean the data is safely on disk.
Attribute the isolation to single-threaded execution, and separate it clearly from rollback (absent) and durability (a persistence setting).
Cover all three axes with the specifics: no undo log, appendfsync windows, asynchronous replication and failover loss, WAIT's limits, and the latency cost of a large EXEC.
Reason about whether Redis is the right system of record given async replication and failover loss, how min-replicas settings convert silent loss into visible failure, and how transaction size interacts with tail latency.
## The guarantee that is real: isolation Redis executes commands from all clients on a single thread. When EXEC arrives, the server takes the queued commands and runs them one after another as a single unit of work; it does not return to the event loop to serve another client in the middle. The consequences are strong and easy to state: - No other client's command can observe the keyspace between two commands of the transaction. - No other client's write can land in the middle of the transaction and be overwritten by half of it. - Every external observer sees either the pre-transaction state or the post-transaction state. That is the entire value proposition of MULTI/EXEC over sending the same commands one at a time or in a plain pipeline. Pipelining reduces round trips but allows interleaving; EXEC does not. It is worth being precise that this isolation covers only the execution window. Anything you read *before* MULTI can be stale by the time EXEC runs, unless you also WATCH the keys involved. ## The guarantee that does not exist: atomicity in the rollback sense In a relational database, "atomic" carries the promise that a failed statement can undo the whole unit. Redis does not keep an undo log, so if a queued command fails at execution time, the commands around it still take effect and stay. The only all-or-nothing case is a command rejected at queue time, which makes EXEC abort before running anything. Using the word "atomic" for both properties is the source of most confusion here; when answering, separate them explicitly: **grouped and uninterrupted, yes; revertible, no.** ## Durability is a separate axis entirely EXEC returning its reply array means the commands were applied to the in-memory dataset. It says nothing about disk. - With **no persistence**, a crash loses everything since the last restart, transaction or not. - With **RDB snapshots only**, you lose everything written since the last snapshot — potentially minutes. - With **AOF and `appendfsync everysec`** (the common default), the transaction is written into the AOF buffer and fsynced at most a second later; a power loss in that window loses it even though EXEC succeeded. - With **AOF and `appendfsync always`**, the write is fsynced before the reply, at a substantial throughput cost. One useful property: the transaction is propagated to the AOF as a single `MULTI ... EXEC` block, so on reload the batch is replayed as a unit rather than partially. If a crash truncates the AOF mid-block, `redis-check-aof` trims the incomplete tail, so you lose the whole transaction rather than half of it. The block-level integrity is preserved; the survival of the block is not guaranteed. ## Replication is asynchronous A primary applies the transaction, replies to the client, and propagates it to replicas afterwards — again wrapped in a MULTI/EXEC block so a replica never applies a partial batch. But there is a window in which the primary has acknowledged and no replica has the data. If the primary dies in that window and a replica is promoted, the committed transaction is gone from the surviving dataset, and a client that read its own write can observe it vanish. `WAIT numreplicas timeout` blocks until the given number of replicas have acknowledged the writes so far, which shrinks the window, but it is not a consensus commit: it can time out, and it cannot rule out a failover to a replica that lagged. `min-replicas-to-write` / `min-replicas-max-lag` make the primary refuse writes when too few replicas are keeping up, converting silent loss into a visible error. Neither turns Redis into a strongly consistent store; if the data cannot tolerate loss on failover, Redis is the wrong system of record for it. ## The cost side of running as one unit The same single-threaded execution that provides isolation makes a large transaction a latency event: while EXEC runs, no other client is served. A transaction containing thousands of commands, or one containing an O(N) command over a big collection, produces a stall visible to everyone and shows up in SLOWLOG under EXEC. Keep transactions small and bounded, and prefer many small units over one giant batch unless the isolation genuinely requires the batch. Memory matters too: the queue is held on the connection until EXEC, so an enormous queued batch consumes memory on the server before it does anything useful. ## Answering the question well A strong answer states the three axes separately: **isolation** — yes, guaranteed by the single execution thread; **rollback** — no, by design, with only the queue-time abort as an all-or-nothing case; **durability and replication** — orthogonal, governed by persistence configuration and by asynchronous replication, so EXEC succeeding is not a promise the data survives a crash or a failover.
- EXEC returned successfully and then the primary crashed a moment later. Can the transaction be lost?Yes, on two independent paths. If persistence is AOF with the default everysec fsync, the block may not have reached disk yet, so a restart from that AOF loses it; RDB-only persistence has a far larger window. And because replication is asynchronous, a replica promoted immediately after the crash may never have received the block. EXEC's reply means "applied in memory", nothing more.
- Does wrapping commands in MULTI/EXEC make them faster than sending them individually?Only incidentally. The gain comes from sending them as one batch, which is pipelining, and you would get the same round-trip saving by pipelining without MULTI. What the transaction adds is isolation — no interleaving — not throughput. A very large EXEC is in fact worse for the system overall, because it occupies the single execution thread for its whole duration.
- How does a replica or an AOF file avoid ever holding half a transaction?Redis propagates the transaction wrapped in a MULTI/EXEC block, so both the AOF and the replication stream carry it as a unit and a replica applies all of it or none. If a crash leaves a truncated block at the end of the AOF, redis-check-aof trims the incomplete tail, which discards the whole partial transaction rather than replaying part of it.
saying these in an interview costs you the question
- Saying a Redis transaction is ACID, or that EXEC guarantees the data is on disk
- Believing MULTI/EXEC is primarily a performance optimisation rather than an isolation mechanism
- Assuming WAIT makes replication synchronous or provides a consensus commit
- Thinking values read before MULTI are automatically protected during EXEC without WATCH
- Ignoring that a large EXEC blocks every other client while it runs