Redis will not undo a batch of commands that was only half applied, and a client can lose its connection at any point — including after sending EXEC. How do you design a multi-command write so that retrying it cannot corrupt state?
answer
- at-least-once, effects idempotent
- absolute value beats delta; EXPIREAT beats EXPIRE
- SET op:<id> NX EX = idempotency key
- payload before pointer; pointer before delete
- TTL on orphan-prone keys
basics
~20 sAssume at-least-once. Prefer idempotent commands (SET, HSET, ZADD) over accumulating ones (INCR, LPUSH, SADD of unique members is fine), order writes so an interrupted batch leaves a benign state, guard non-idempotent effects with a dedupe key set via SET token NX EX, and give orphan-prone keys a TTL.
solid answer
~60 sStart from the honest failure model: the client can die before EXEC (nothing ran), after EXEC (everything ran), or with the reply lost (unknown). Since Redis will not roll back and cannot tell you which happened, the batch must be safe to run twice. Four levers: 1. **Choose idempotent commands.** `SET k v` and `HSET h f v` converge; `INCR`, `LPUSH` and `RPUSH` double on retry. If you need a counter that tolerates retries, key the increment by an operation id and deduplicate. 2. **Deduplicate explicitly.** `SET op:<uuid> 1 NX EX 600` as the first command; if it returns nil the work already happened, so skip. This turns any batch into an effectively-once one within the TTL window. 3. **Order for benign intermediates.** Write the payload key before adding its id to an index or sorted set, so a crash leaves an unreferenced object rather than an index entry pointing at nothing. 4. **Let garbage expire.** Give orphan-prone keys a TTL so a partial batch cleans itself instead of needing a repair job. Where the logic is read-modify-write, a Lua script removes the round trips — but it still gives no rollback, so it must be idempotent too.
code
text · 9 lines# 1. claim: nil means this operation already ran
SET op:9f3c1 1 NX EX 3600
# 2. payload first (safe orphan if we die here), with a TTL
HSET item:42 title "Hello" body "..."
EXPIREAT item:42 1735689600
# 3. pointer last — never advertises a missing payload
ZADD feed 1735600000 42go deeper
Know the two buckets: SET/HSET are safe to repeat, INCR/LPUSH are not, and a retry must not double an effect.
Explain the retry window around EXEC and show the dedupe-key pattern with SET NX EX plus the absolute-value-over-delta rule.
Lead with the failure model, then the ordering rule (payload before pointer), TTLs for orphans, and where a Lua script does and does not help.
Argue the posture — Redis gives isolation, not atomic recovery, so correctness is an application property — and decide which invariants get write-time enforcement versus detection and reconciliation.
## The failure model you must design for Three things can happen around `EXEC`: - The connection dies **before** EXEC reaches the server: nothing ran; the queued commands die with the connection. - EXEC reached the server and **ran**, but the reply never got back (timeout, network drop, proxy reset): everything is applied and you do not know it. - EXEC ran and one command errored: partially applied, permanently, and the error is buried in the array reply. No Redis feature closes the middle gap. There is no transaction id you can query, no "did this commit?" call. So the only viable posture is **at-least-once delivery with idempotent effects** — exactly the discipline you would apply to a retried HTTP POST. ## Lever 1 — prefer commands that converge Sort your writes into two buckets. *Idempotent (safe to repeat):* `SET`, `HSET`, `SETEX`, `ZADD` with a computed score, `SADD` of a fixed member, `DEL`, `EXPIREAT` with an absolute timestamp. *Accumulating (doubles on retry):* `INCR`/`INCRBY`, `LPUSH`/`RPUSH`, `ZINCRBY`, `APPEND`, `EXPIRE` with a relative TTL (each retry pushes the deadline out), `BITCOUNT`-style counters. Whenever a design lets you compute an absolute target value instead of a delta, take it: `SET balance 500` retried is still 500; `INCRBY balance 100` retried is 200 too much. `EXPIREAT` over `EXPIRE` is the same idea for deadlines. ## Lever 2 — an explicit dedupe token When the effect is genuinely accumulating — a credit, an append to a stream of events, a push onto a work list — attach a caller-generated operation id and claim it first: If `SET op:<id> 1 NX EX 3600` returns nil, the operation already ran; return the cached outcome or simply stop. If it returns OK, you own the operation and proceed. This is the Redis-primitive version of an idempotency key. Two caveats to state out loud in an interview: the guarantee only holds for the token's TTL, and there is a window between claiming the token and completing the writes in which a crash leaves the token claimed but the work unfinished — which is why the token should be claimed *inside* the same atomic unit (a Lua script) when the effect must not be lost, or given a short TTL when losing it is preferable to duplicating it. ## Lever 3 — order so the intermediate state is benign Since a batch may stop at any command boundary only when it never ran at all — but a *client-side sequence* of batches certainly can stop midway — decide which dangling state you can live with and order to produce that one. The usual rule: **create the thing before you advertise it**. Write `item:42` (the payload hash) first, then `ZADD feed 1699… 42` (the pointer). An interrupted sequence leaves an item nobody references, which is invisible garbage. The reverse order leaves a feed entry whose payload lookup returns nil — a user-visible error and a source of null-handling bugs everywhere downstream. Same rule for deletions, mirrored: remove the pointer first, then the payload. ## Lever 4 — make garbage self-cleaning Give orphan-prone keys a TTL at creation. An unreferenced `item:42` that expires in an hour needs no repair job; the same key without a TTL is a permanent leak that shows up months later as unexplained memory growth. Where TTLs are unacceptable, plan a reconciliation sweep (`SCAN`, never `KEYS`) that compares index membership against payload existence and repairs one side. ## Where Lua fits Moving read-modify-write logic into a script collapses several round trips into one server-side unit, which removes the class of failures that happen *between* your commands and lets you make the decision and the write inseparable. What it does not remove is the retry problem: if the client loses the reply it still does not know whether the script ran, and a script that errors after writing leaves those writes behind. So a script must be written to be idempotent for the same reasons a batch must be. ## Checklist to say out loud - Can this batch run twice with the same end state? If not, which command breaks it? - Is there an operation id, and where is it claimed relative to the effects? - If the process dies between two commands, is the resulting state invisible garbage or a broken reference? - Does every key a partial run can create have a TTL or a reconciliation owner? - Am I checking every element of the EXEC reply, or trusting that no exception means success?
- The client times out waiting for EXEC's reply. How do you find out whether the batch was applied?Redis cannot tell you — there is no transaction handle to query. You either read back a value the batch would have written and infer from it, or, better, you designed the batch so the question does not matter: claim an operation id up front and make every effect idempotent, so retrying is correct regardless of what actually happened.
- You must increment a counter exactly once per request. How do you do it with Redis primitives?Attach a request id and make the claim and the increment inseparable — a small Lua script that does `SET op:<id> 1 NX EX <ttl>` and only calls `INCR` when the claim succeeded. Doing the two as separate round trips leaves a window where the claim exists but the increment never happened, so the count silently under-counts on crash.
- When is a reconciliation sweep the right answer instead of stronger write-time guarantees?When the invariant spans keys that no single atomic unit can cover — different hash slots in a cluster, or Redis plus an external database — and when a brief inconsistency is tolerable. A periodic SCAN-based sweep that repairs dangling pointers is far cheaper than contorting the write path, provided the anomaly is detectable and the window is acceptable to the product.
Treat every batch like a retried payment request: you cannot un-charge the card, so you attach an idempotency token and make the second attempt a no-op.
saying these in an interview costs you the question
- Assuming a timeout means the batch did not run
- Using INCR or LPUSH in a batch that the client retries on failure
- Setting a relative EXPIRE inside a retryable batch, so each retry extends the deadline
- Adding the index entry before the payload it points at
- Claiming a Lua script makes the operation exactly-once by itself
- Planning to 'clean up manually' instead of TTLs or a reconciliation job