You're implementing a rate limiter shared across many stateless application server instances, backed by Redis, using a simple GET-then-SET pattern: read the current counter, check it against the limit, and if under, increment and write it back. Under concurrent load, why does this let clients exceed their limit, and how do you fix it?
answer
- check-then-act race
- INCR is atomic, GET+SET isn't
- Lua script = atomic multi-step
- EXPIRE right after first INCR
- fail-open vs fail-closed on Redis outage
basics
~20 sTwo requests can both read the counter before either writes it back, so both see 'under limit' and both proceed — a race condition called a check-then-act bug. The fix is to make the check-and-increment a single atomic operation, e.g. Redis INCR or a Lua script.
solid answer
~50 sA read-check-write sequence done as three separate Redis round trips is not atomic: if two requests from the same client arrive concurrently, both can execute GET and see the counter at, say, 99 (under a 100 limit) before either has issued the write, so both proceed and increment to 101 — the limiter has been bypassed by a classic time-of-check-to-time-of-use race. The fix is to collapse check-and-increment into a single atomic Redis operation. The simplest is INCR (atomic by design) combined with EXPIRE set only on the first increment (INCR returns 1), then compare the returned value against the limit after incrementing rather than before. For anything more complex — token bucket refill math, sliding window counters — use a Lua script executed via EVAL, since Redis guarantees a script runs to completion without interleaving with other clients' commands, giving a correct multi-step check-and-update as one atomic unit.
go deeper
Should recognize, at a high level, that reading then writing in two separate steps can go wrong when done concurrently.
Should name INCR as the atomic fix for a simple counter and understand why GET+SET isn't safe under concurrency.
Expected to know when a plain INCR is insufficient (token bucket, sliding window counter) and reach for a Lua script for multi-step atomicity, plus handle the EXPIRE race correctly.
Expected to additionally weigh Redis as a shared dependency's latency and single-point-of-failure implications, and make an explicit fail-open/fail-closed decision for when it's unreachable.
## The bug The bug in a GET-then-SET rate limiter is a textbook **check-then-act race condition**, made worse by the fact that Redis itself is safe per-command but the client-side sequence of commands is not. Concretely: 1. client issues `GET counter:userA` and gets 99; 2. the application compares 99 < 100 and decides to allow; 3. the application then issues `INCR` (or `SET 100`) to record the request. If two requests from userA arrive on two different application server instances (or even two threads on one instance) close enough in time, both can complete their GET before either completes its write-back. Both see 99, both conclude "under limit, allow," and both then write, leaving the true count at 101 having allowed 101 requests through a 100-request limit — and under higher concurrency the overshoot can be far larger than 1. This is invisible in single-instance, low-concurrency testing and only shows up under real production load, which is exactly why it is a dangerous bug to ship. ## Why it happens The root problem is that the check (read + compare) and the act (write) are **not a single indivisible operation** from Redis's point of view, even though each individual GET or SET command is atomic. The fix is architectural: make the entire check-and-increment sequence atomic as one Redis operation, so no other client can observe or interleave with an in-progress increment. ## The simple fix — one atomic command The simplest correct pattern for a basic fixed-window counter uses Redis's `INCR` command, which atomically increments a key and returns the new value in one round trip — there is no separate read step to race on. The application does: 1. `result = INCR(key)`; 2. if `result == 1`, that means this was the key's first increment (previous value didn't exist), so set an expiry with `EXPIRE(key, windowSeconds)` — this needs its own follow-up call, and there's a subtle sub-race if the process crashes between INCR and EXPIRE, leaving a key with no expiry (mitigated by using `SET key value EX seconds NX` first to establish a TTL, or a periodic reaper); 3. then check whether `result > limit`, rejecting if so. Because the increment and the return value come back from a single atomic command, there is no window where two clients can both read a stale pre-increment value. ## When one command falls short For anything beyond a plain counter — token bucket with time-based refill math, sliding window counters that need to read and weight two window buckets, or any limiter requiring multiple reads and a conditional write in one logical unit — the standard solution is a **Lua script** executed via Redis's EVAL/EVALSHA. Redis guarantees that a Lua script runs to completion as a single atomic unit with respect to all other commands: no other client's command can be interleaved partway through the script's execution, even though the script itself may contain several GETs, arithmetic, and a conditional SET. This lets you implement, say, token bucket's "compute elapsed time since last refill, add tokens up to capacity, check if >= 1 token available, decrement if so, write back new state" entirely correctly in one round trip, with the atomicity guarantee coming from Redis's single-threaded command execution model rather than application-level locking. ## What scripts cost The trade-off of Lua scripts is **operational complexity**: - scripts must be deployed and versioned (often loaded once via `SCRIPT LOAD` and invoked by SHA to avoid re-sending source every call), - debugging is harder than plain commands, - a slow or buggy script blocks the single Redis event loop for its duration, so scripts must stay fast and simple. INCR-based approaches are simpler to reason about but only cover the fixed-window case directly. ## The second-order failure mode A second-order failure mode specific to distributed rate limiting is that Redis itself becomes a shared dependency and potential bottleneck or single point of failure: if every request round-trips to a central Redis instance to check its limit, Redis latency is added to every request's critical path, and a Redis outage either: - **fails open** (no limiting — dangerous), or - **fails closed** (all requests rejected — an outage of the whole API), depending on how the failure is handled, which is a deliberate design decision that must be made explicitly, not left as an accident of exception-handling code.
- Why use EVALSHA instead of EVAL for the Lua script in production?EVAL sends the full script source on every call, adding bandwidth and parse overhead; EVALSHA sends only the script's SHA1 hash after it's been loaded once via SCRIPT LOAD, so Redis can execute the cached compiled script without re-transmitting or re-parsing source on every rate-limit check, which matters at high request volume.
- What's the risk of setting EXPIRE only after checking that INCR returned 1?If the process crashes or the connection drops between the INCR and the follow-up EXPIRE call, the key is left with no TTL and will never expire, silently turning that client's counter into a permanent one that blocks them forever. A safer pattern uses SET key 0 EX windowSeconds NX before incrementing, or wraps both calls in a single Lua script so they can't be split by a crash.
- Would wrapping the GET-then-SET in a Redis MULTI/EXEC transaction fix the race?Not by itself — a plain MULTI/EXEC still executes the GET and SET as separate commands, and if the GET happens outside the transaction or the increment logic depends on a value read before MULTI started, the same stale-read race can occur. You'd need WATCH on the key for optimistic locking with retry on conflict, or more simply just use INCR/Lua directly, which sidesteps the need for a transaction entirely.
It's like two people checking a shared bank balance shows $50 available at the same instant, and both then withdraw $40 based on that stale read — the balance goes negative because the check and the withdrawal weren't one atomic step.
saying these in an interview costs you the question
- Doesn't recognize GET-then-SET as a race condition
- Suggests fixing it with application-level locking across multiple stateless server instances instead of using Redis's own atomicity
- Doesn't know INCR is atomic
- Has no opinion on fail-open vs fail-closed when Redis is unreachable