skip to content

What extra constraints does Redis Cluster impose on Lua scripts executed with EVAL and on transactions opened with MULTI, compared with running the same code against a single Redis instance?

level: seniorimportance: must knowfreq 46%

answer

  1. script + transaction = one node, one slot
  2. every key in KEYS, none built from ARGV
  3. undeclared or non-local key = error
  4. WATCH only works within the slot
  5. scripts/functions are per-node state, load on every master

basics

~20 s

Both run entirely on one node. Every key a script or transaction touches must hash to one slot, all script keys must be declared in KEYS (Redis refuses undeclared or non-local key access), and a cross-slot command inside MULTI aborts the transaction.

solid answer

~50 s

A script and a `MULTI` block are both executed by a single node against local data, so cluster mode adds two rules. **Declare and co-locate keys.** `EVAL script numkeys key... arg...` must list every key it touches in `KEYS`, and all of them must hash to one slot - otherwise the call is rejected before execution. That is also how the client knows where to route the script. Redis actively enforces locality at runtime: touching a key that was not declared, or that belongs to another slot, raises an error rather than reading someone else's shard. Building key names inside the script from `ARGV` is therefore a bug in cluster mode. The same holds for `FCALL` with Redis Functions (7.0+). **Transactions are single-node too.** Commands queued after `MULTI` must all target the same slot; a queued command for another slot fails and the transaction is discarded at `EXEC`. `WATCH` keys must be in that slot as well. So cluster-safe scripts operate on one hash-tagged group.

code

text · 6 lines
text
# OK: both keys share the hash tag {acct:42} -> one slot
> EVAL "redis.call('DECRBY', KEYS[1], ARGV[1]); return redis.call('RPUSH', KEYS[2], ARGV[1])" 2 "{acct:42}:balance" "{acct:42}:ledger" 10

# Rejected: declared keys span slots
> EVAL "return 1" 2 acct:42:balance acct:99:ledger
(error) CROSSSLOT Keys in request don't hash to the same slot

go deeper

for a junior

Recall that scripts and transactions run on one node, so all their keys must share a slot and be listed in KEYS.

for a middle

Explain why KEYS is load-bearing for routing and enforcement, and that a cross-slot command aborts a MULTI block.

for a senior

Cover runtime enforcement of non-local key access, connection pinning for transactions, per-node script/function distribution, and retry on topology change.

for a principal

Position the single-slot rule as the atomicity boundary of the system: decide which invariants deserve co-location and design explicit application-level protocols for everything that crosses shards.

## Why the constraint exists On a standalone Redis, a Lua script or a `MULTI/EXEC` block is atomic because the server executes it as one unit on its single command-execution thread. In cluster mode there is no distributed execution engine: whatever runs, runs on one node, seeing only the slots that node owns. To keep the atomicity guarantee meaningful, Redis requires that the whole unit fit inside one slot. ## Rules for EVAL / EVALSHA / FCALL **1. Declare every key in KEYS.** The `numkeys` argument splits key names from arguments: `EVAL script 2 k1 k2 arg1`. On a standalone server, `KEYS` is a convention you can cheat on; in cluster mode it is load-bearing. The client uses `KEYS[1]` to compute the slot and route the call, and the server uses the declared set to verify locality. **2. All declared keys must hash to one slot.** If they do not, the call fails immediately - the script never runs. In practice cluster-safe scripts operate on a hash-tagged family: `{order:77}:items`, `{order:77}:total`. **3. No undeclared or non-local keys at runtime.** Modern Redis blocks a script that accesses a key it did not declare, and blocks access to a key that does not belong to the node's slots, with errors along the lines of *"Script attempted to access keys that do not hash to the same slot"* or *"attempted to access a non local key in a cluster node"*. The tempting pattern `redis.call('GET', 'prefix:' .. ARGV[1])` therefore breaks under cluster even though it works on a single instance - and it is unroutable in principle, since the client cannot know the slot in advance. **4. Consequences for design.** Anything that needs to read one shard and write another cannot be a script. Split it into two operations in the application, and reason explicitly about the failure between them (idempotent retries, a reconciliation pass, or an outbox) instead of pretending you still have one atomic step. **5. Functions (Redis 7.0+).** `FUNCTION LOAD` registers a library on a node; `FCALL fn numkeys keys...` obeys the same key-declaration and single-slot rules. Note a distribution wrinkle: libraries and cached scripts are per-node state. `SCRIPT LOAD` must reach every master (clients that use `EVALSHA` handle `NOSCRIPT` by re-sending the source), and functions must be loaded on each shard - `redis-cli --cluster call` is the usual tool. Replicas receive them via replication, but a newly added shard does not have them. ## Rules for MULTI / EXEC / WATCH A transaction is queued on one connection to one node and then executed there. Cluster adds: - **Single slot for the whole block.** The first command's key fixes the node; a later command whose key belongs to another slot is answered with an error at queue time, and the transaction is aborted so `EXEC` fails rather than partially applying. - **WATCH must be local too.** Optimistic concurrency with `WATCH` only observes keys in the same slot on the same node - you cannot watch a key on another shard. - **Connection affinity.** A cluster client normally routes per command; for a transaction it must pin the whole block to one connection to one node. Some clients require you to be explicit about this, and a naive wrapper that routes each queued command independently will break. - **Topology changes mid-transaction.** If the slot moves while a transaction is open, the block fails and must be retried against the new owner; treat `MULTI` in cluster mode as retryable. ## Practical patterns - **Hash-tag the unit of atomicity.** Decide up front which keys must be updated together and give them a shared, narrow tag: `{acct:42}:balance`, `{acct:42}:ledger`. That group defines both your atomicity boundary and your co-location cost. - **Keep scripts short.** Cluster does not change the fact that a script monopolises the node's execution thread for its duration; a long script is a latency incident for every key in that shard. - **Prefer single-key structures.** A hash, a sorted set, or a stream that holds the whole unit avoids the problem entirely - one key is always one slot. - **Do not fake cross-shard atomicity.** Two scripts on two shards give you no joint guarantee. If you need cross-shard consistency, that is an application-level protocol (idempotency keys, compensations), not something Redis will provide. ## Quick self-check If you can answer "which single slot does this script/transaction operate on, and how did the client know that before sending it?" then it is cluster-safe. If the answer depends on data read at runtime, it is not.

  • A script works on a standalone Redis but fails in cluster mode with an error about non-local keys. What is the most likely cause?
    The script constructs key names at runtime instead of receiving them through KEYS. Redis Cluster verifies that a script only touches declared keys belonging to the node's slots, and the client cannot route a call whose keys it cannot see. The fix is to pass every key as a KEYS argument and hash-tag them into one slot.
  • Can you use MULTI/EXEC to update two keys on different shards atomically?
    No. The transaction is queued and executed on one node, so a command for another slot is rejected and the block is discarded. Cross-shard atomicity does not exist in Redis Cluster; you either co-locate the keys with a hash tag or design an application-level protocol with idempotent, retryable steps.
  • Why can a newly added shard return NOSCRIPT for a script that has been running for months?
    Script and function definitions are per-node state propagated to replicas by replication, not gossiped cluster-wide. A freshly added master has an empty script cache, so EVALSHA fails with NOSCRIPT until the client falls back to EVAL with the source or the operator loads it on every master.

saying these in an interview costs you the question

  • Building key names inside a Lua script from ARGV and expecting it to work in cluster mode
  • Believing MULTI/EXEC gives cross-shard atomicity
  • Assuming SCRIPT LOAD or FUNCTION LOAD propagates to every master in the cluster
  • Thinking KEYS versus ARGV is only a stylistic convention
  • Expecting a client that routes each command independently to handle a transaction correctly

context