skip to content

You need a read-modify-write over Redis data that must not lose concurrent updates. Compare doing it with WATCH-based optimistic concurrency against doing it in a single server-side Lua script, and say when you would pick each.

level: seniorimportance: must knowfreq 44%

answer

  1. WATCH = detect conflict, retry; Lua = remove the window
  2. 3 RTT per attempt vs 1 RTT flat
  3. retries collapse throughput as p rises
  4. script blocks the thread — keep it microseconds
  5. KEYS[] + hash tags for cluster routing; NOSCRIPT fallback

basics

~20 s

WATCH validates client-side logic and retries on conflict — flexible, but costs multiple round trips per attempt and degrades as contention rises. A Lua script does read, decide and write in one uninterrupted server-side pass — one round trip, no retries — but must be short, same-slot, deployed, and still offers no rollback.

solid answer

~60 s

**WATCH** keeps the decision in your application: watch the keys, read, compute in whatever language you like, queue the writes, and let EXEC return nil if anything moved. Cost: about three round trips per attempt, all of them repeated on conflict. It shines when the logic is complex, uses application state or libraries, or the conflict rate is genuinely low. **Lua** moves the decision into the server. The script executes atomically — nothing interleaves — so there is no conflict window at all and therefore no retry loop and one round trip. It shines for short, hot, high-contention operations: conditional debit, rate-limit token bucket, compare-and-delete for a lock. The script's costs are real: it occupies the execution thread, so a slow script raises latency for every client; in cluster mode all keys must be in one hash slot and should be passed as `KEYS` so the client can route; you take on script deployment (`SCRIPT LOAD`/`EVALSHA` with a NOSCRIPT fallback, or Redis Functions in 7.0+); and errors mid-script leave earlier writes applied, exactly as with EXEC. Rule of thumb: low contention plus complex logic → WATCH; high contention plus small logic → script.

code

lua · 8 lines
lua
-- KEYS[1] = balance key, ARGV[1] = amount to debit
local bal = tonumber(redis.call('GET', KEYS[1]) or '0')
local amt = tonumber(ARGV[1])
if bal < amt then
  return -1                       -- rejected, nothing written
end
redis.call('SET', KEYS[1], bal - amt)
return bal - amt

go deeper

for a junior

State the core difference: WATCH retries when someone else changed the data; a script does read and write together on the server so there is nothing to retry.

for a middle

Quantify the round trips, note that scripts run atomically and must be short, and mention EVALSHA plus the NOSCRIPT fallback.

for a senior

Drive the choice from the conflict rate and script duration, cover cluster KEYS routing, script deployment and the fact that neither gives rollback.

for a principal

Frame it as where computation belongs relative to the data, weigh the operational surface of shipping scripts against retry-storm risk, and set the team rule (atomic command first, script for hot small logic, WATCH for complex cold logic).

## The same problem, two shapes Both mechanisms solve lost updates for a read-modify-write. They differ in *where the decision runs*. **WATCH** is optimistic validation: you tell the server which keys your decision depends on, do the thinking on the client, and the server refuses your write if those keys moved. Concurrency is unrestricted; conflict is detected and paid for after the fact with a retry. **Lua** removes the window instead of detecting it. A script runs as one uninterrupted unit on the server, so nothing can change your inputs between read and write. There is no conflict to detect and no retry to write. ## Round-trip and contention economics A WATCH attempt is roughly `WATCH` → read → `MULTI`+queue+`EXEC`: three network exchanges (some foldable with pipelining). A script is one `EVALSHA`. Now add contention. With conflict probability *p*, the expected number of WATCH attempts is about 1/(1−p). At p=0.1 that is barely noticeable. At p=0.7 you are doing three attempts on average, nine-ish round trips, and — worse — every failed attempt still consumed a full read of the data and a server-side transaction setup, so throughput collapses precisely when the system is busiest. This is the classic optimistic-concurrency failure mode: it converts contention into wasted work rather than into waiting. A script's cost is flat in contention, because clients simply queue for the execution thread. Its cost scales with *script duration* instead: while it runs, no other command runs, so a 5 ms script caps the instance at 200 ops/s of anything. Keep scripts to the microsecond scale, avoid loops over unbounded collections, and never call anything of unpredictable size inside one. ## Expressiveness WATCH lets your decision use anything the client can do: business rules, a library, another service's response, a locale-aware comparison. Lua gives you a small sandbox — no network, no time-dependent nondeterminism you should rely on, no libraries beyond what Redis embeds. If the decision cannot be expressed compactly in Lua, WATCH is not a fallback, it is the right tool. Conversely, if the decision *can* be expressed in five lines, a script is almost always better: it is shorter, has no retry code to get wrong (re-WATCH bugs are common), and behaves predictably under load. ## Cluster and key routing Both need all involved keys in one hash slot, because both execute on one node. For scripts the discipline is stricter and healthier: pass every key through `KEYS[]` so the client can compute the slot and route correctly, and use hash tags to co-locate. Keys constructed inside the script body are invisible to routing and are the standard cause of "works on standalone, breaks in cluster". ## Operational surface WATCH needs no deployment: it is just commands. A script needs a lifecycle — `SCRIPT LOAD` returning a SHA, `EVALSHA` in the hot path, a `NOSCRIPT` error handler that re-loads after a restart or failover, and a story for keeping the script text in version control. Redis 7.0 Functions formalise this with libraries registered on the server that survive restarts and replicate, at the cost of a registration step. Scripts also complicate incident response: a script that has already written cannot be terminated by `SCRIPT KILL` (killing it would leave a torn state the server refuses to accept), so the only escape from a runaway writing script is a hard `SHUTDOWN NOSAVE` with the data loss that implies. That is a strong argument for keeping scripts provably short. ## What neither gives you Neither provides rollback. A script that errors after two writes leaves those two writes applied and returns an error; an EXEC that runs with a stale-but-unwatched assumption applies fully. Both therefore need the same idempotency discipline for client retries after a lost reply. ## And the third option Before choosing either, check whether a single existing command already does the job atomically: `INCRBY`, `SET … NX`, `SET … XX GET`, `HSETNX`, `ZADD … GT`, `SETRANGE`, `GETDEL`, `SMOVE`. A surprising share of hand-rolled CAS loops and scripts reimplement one of these with more code and worse latency. ## The answer to give "Single atomic command if one exists. If not: short, hot, contended logic goes into a Lua script or a Function, because it removes the conflict window and costs one round trip. Complex or client-dependent logic with low conflict rates stays in a WATCH loop with bounded, jittered retries. Either way the operation must be idempotent, because neither mechanism rolls back and neither tells me whether a lost reply meant applied or not."

  • At roughly what conflict rate does a WATCH loop stop being a good idea?
    There is no hard threshold, but the expected attempts are about 1/(1−p), so the wasted work grows sharply past a conflict probability of a few tens of percent. Practically: if your metrics show more than one retry per successful operation on average, the loop is spending more time re-reading than working and the logic should move server-side.
  • What can go wrong with EVALSHA in production and how do you handle it?
    The script cache is not persistent — after a restart, a failover to a replica that never had the script, or a SCRIPT FLUSH, EVALSHA returns NOSCRIPT. The client must catch that error, re-send the body with EVAL or SCRIPT LOAD, and retry. Redis 7 Functions avoid the issue by persisting and replicating registered libraries.
  • A Lua script has entered an infinite loop after performing a write. What are your options?
    SCRIPT KILL refuses to terminate a script that has already written, because stopping it would leave partially applied state the server will not accept. The remaining option is SHUTDOWN NOSAVE, which drops the writes since the last save point. That severity is exactly why scripts must be bounded and reviewed for unbounded loops before they ship.

WATCH is submitting a form and being told someone edited the record first; a script is handing the clerk your instructions and letting them do the whole edit at the counter while nobody else is served.

saying these in an interview costs you the question

  • Claiming a Lua script gives rollback or transactional undo
  • Recommending WATCH for a hot contended key without discussing retry cost
  • Writing long or unbounded-loop scripts and ignoring that they block every other client
  • Building key names inside the script instead of passing them via KEYS in cluster mode
  • Using EVALSHA with no NOSCRIPT fallback path
  • Reaching for either mechanism when a single atomic command like INCRBY or SET NX already exists

context