skip to content

A Lua script sent to Redis with EVAL is described as executing atomically. What exactly does that guarantee, what does it cost the rest of the server, and what can an operator do about a script that will not finish?

level: middleimportance: must knowfreq 54%

answer

  1. One unit of work on the single command thread
  2. No interleaving — but no rollback either
  3. busy-reply-threshold (was lua-time-limit, 5s) → BUSY replies
  4. SCRIPT KILL only if nothing was written
  5. Wrote + stuck ⇒ SHUTDOWN NOSAVE, data loss

basics

~20 s

The script runs to completion with no other client command interleaved, so read-modify-write logic is safe. The cost: it occupies the single command-processing thread, so every other client waits. After busy-reply-threshold Redis replies BUSY and accepts SCRIPT KILL — but only if the script has not written.

solid answer

~50 s

Redis processes commands on one thread, and a script is dispatched as a single unit of work. So from any other client's perspective the script's effects appear all at once: no command from anyone else runs between two `redis.call`s. That is what makes check-then-act logic — compare a value, then conditionally write — safe without locking. The price is that the script **is** the server for its duration. A loop over a million elements is a million-element stall for every other client, visible as a latency spike, timeouts and, in a Sentinel/Cluster setup, potentially a spurious failover. Redis guards this with `busy-reply-threshold` (formerly `lua-time-limit`, default 5000 ms). After it elapses, the server keeps running the script but starts answering other clients with `BUSY`, accepting only `SCRIPT KILL`/`FUNCTION KILL` and `SHUTDOWN NOSAVE`. `SCRIPT KILL` works **only if the script has performed no write** — there is no rollback, so a partially-written script cannot be aborted; then `SHUTDOWN NOSAVE` is the only exit. Keep scripts short and bounded.

code

text · 14 lines
text
# client A
> EVAL "while true do end" 0
(blocked)

# client B, after busy-reply-threshold
> GET foo
(error) BUSY Redis is busy running a script. You can only call SCRIPT KILL or SHUTDOWN NOSAVE.

> SCRIPT KILL
OK                      # script performed no write

# but if it had written first:
> SCRIPT KILL
(error) UNKILLABLE Sorry the script already executed write commands against the dataset.

go deeper

for a junior

Say that the script runs as one indivisible unit so no other command interleaves, and that a slow script makes everyone else wait.

for a middle

Add the mechanics: single command-processing thread, busy-reply-threshold producing BUSY replies, SCRIPT KILL, and the absence of rollback.

for a senior

Diagnose and prevent: SLOWLOG and latency correlation, unkillable-after-write and the SHUTDOWN NOSAVE consequence, read-only invocation to stay killable, bounded batch sizes, and the failover risk from long stalls.

for a principal

Reason about the bargain — buying atomicity with head-of-line blocking — and decide where that trade stops paying, e.g. moving large computations out of the data tier or sharding hot logic.

## What atomic means here Redis executes client commands on a single command-processing thread. When a script arrives via `EVAL`/`EVALSHA`, the server runs the whole script as one such unit of work: the Lua interpreter runs, its `redis.call` invocations execute inline against the data, and only when the script returns does Redis move on to the next client's command. So the guarantee is **isolation by serialization**: no other client's command is interleaved between any two operations of the script. If your script does ```lua local v = tonumber(redis.call('GET', KEYS[1]) or '0') if v < tonumber(ARGV[1]) then return redis.call('INCR', KEYS[1]) end return v ``` nobody can change `KEYS[1]` between the `GET` and the `INCR`. This is why Lua is the standard answer for conditional, multi-step, multi-key updates in Redis — a capped counter, a token bucket, a compare-and-delete lock release. What atomicity does **not** give you is rollback. If the fifth command in a script fails or the script raises a Lua error after four successful writes, those four writes stand. Atomicity means "nothing interleaves", not "all or nothing". ## The cost: the script owns the server Because it is one unit of work on the only command-processing thread, the script's runtime is dead time for every other connection. The relevant discipline is the same as for any command: know the complexity of what you call and bound it. - `redis.call('KEYS', '*')` inside a script is O(N) over the whole keyspace — a stall proportional to the database. - Looping `for i = 1, 1000000` over `LPOP` is a million commands' worth of work in one blocking unit. - Fetching a big collection into Lua (`LRANGE key 0 -1` on a huge list) costs both time and a large temporary allocation. Symptoms are latency spikes across all clients, client-side timeouts, and — importantly — health-check timeouts. A script that runs longer than a Sentinel's `down-after-milliseconds` or a Cluster's node timeout can cause the node to be judged failed and trigger an unnecessary failover. ## The BUSY state Redis does not preempt scripts, but it does not leave you blind either. The `busy-reply-threshold` configuration (Redis 7.0 name; earlier `lua-time-limit`, default 5000 ms) controls when the server, while still running the script, begins responding to other clients with: ``` (error) BUSY Redis is busy running a script. You can only call SCRIPT KILL or SHUTDOWN NOSAVE. ``` This is a signal, not a timeout — the script is not aborted. In that state the server accepts only `SCRIPT KILL` (or `FUNCTION KILL` for Functions) and `SHUTDOWN NOSAVE`. ## Killing — and when you cannot `SCRIPT KILL` terminates the running script and returns an error to its caller. It is refused with `UNKILLABLE` if the script has **already executed a write command**. The reason is the absence of rollback plus replication integrity: the writes performed so far may already be queued for propagation to replicas and the AOF, and aborting mid-way would leave the primary in a state that cannot be described coherently to them. When a writing script is stuck, the only remedy is `SHUTDOWN NOSAVE` — kill the server without writing an RDB — losing everything since the last save point (or relying on the AOF, if enabled, up to its last fsync). This is a genuine data-loss event, which is why "bound your scripts" is not pedantry. Two mitigations follow directly: 1. **Declare read-only scripts as such.** With the `no-writes` shebang flag (7.0) or `EVAL_RO`/`EVALSHA_RO`, the script cannot write, and therefore is always killable. 2. **Do writes last.** A script that computes first and writes at the end has a large killable window and a small unkillable one. ## Practical rules - Treat a script as a single command whose complexity is the sum of its parts; aim for microseconds to low milliseconds. - Never iterate an unbounded collection inside a script; pass a bounded batch size in `ARGV` and let the caller loop. - Avoid `KEYS`, unbounded `LRANGE`/`ZRANGE`/`SMEMBERS`, and unbounded `while` loops. - Watch `SLOWLOG` — script executions appear there — and `INFO commandstats` / latency percentiles to catch regressions. - If the work is genuinely large, do it outside Redis or chunk it, rather than buying atomicity at the price of a server-wide stall. ## Why the guarantee is worth the constraint The reason this bargain is acceptable is that the alternative — a lock protocol between clients — costs more round trips, introduces failure modes when a client dies holding a lock, and still cannot make a multi-key update indivisible. A short script buys real atomicity for the cost of the microseconds it runs. The failure mode only appears when the script stops being short.

  • If a script fails halfway through, do the writes it already performed get undone?
    No. Redis has no rollback for scripts: commands already executed have taken effect and will be propagated. Atomicity here means no other client's command interleaved, not that the script is all-or-nothing. If you need to avoid a half-applied state, validate everything up front and perform the writes only after all checks pass.
  • Why does Redis refuse SCRIPT KILL once the script has written?
    Because there is no way to undo the writes already applied, and those effects may already be on their way to replicas and the AOF. Aborting would leave the primary in a state that cannot be described consistently to its followers. The only exit is SHUTDOWN NOSAVE, which discards data since the last save point, so the real defence is keeping scripts bounded and doing writes last.
  • How would you notice that a script is hurting production before someone gets a BUSY error?
    Script executions show up in SLOWLOG, so a rising count of slow EVAL entries is the earliest signal. Latency percentiles across all commands rise together during a stall, because everything queues behind the script, and INFO commandstats shows the per-call cost of EVAL/EVALSHA. Correlating a p99 spike affecting every command type with EVAL entries in SLOWLOG is the usual diagnosis.

It is a single-lane bridge with no passing places: your car is guaranteed an uninterrupted crossing, and everyone behind you waits exactly as long as you take.

saying these in an interview costs you the question

  • Saying a script is transactional and rolls back on error
  • Assuming Redis preempts or times out a long script automatically
  • Believing SCRIPT KILL always works
  • Thinking BUSY means the script was aborted
  • Putting KEYS * or an unbounded loop inside a script because 'it's atomic anyway'

context