How does RedisRateLimiter enforce limits atomically across gateway instances, and what response headers and denial behavior does it produce?
answer
- Lua script = atomic read-refill-decrement-write
- Redis single-threaded -> no double-spend across nodes
- X-RateLimit-Remaining / -Burst-Capacity / -Replenish-Rate
- 429 + short-circuit on deny; TTL on idle keys
- Redis down -> tends to fail open (availability vs security)
basics
~20 sIt runs a Redis Lua script that atomically refills the token bucket by elapsed time and decrements it, so concurrent gateway nodes can't race. On success it adds X-RateLimit-* headers; when tokens run out it sets HTTP 429 and stops the request.
solid answer
~40 sRedisRateLimiter keeps two values per key (token count and last-refill timestamp) and evaluates a **Lua script** in Redis. Lua scripts run atomically in Redis's single-threaded execution, so the read-refill-decrement-write cycle is a single indivisible operation — multiple gateway instances hitting the same key can't double-spend or lose updates. The script computes tokens regenerated since the last call from `replenishRate`, caps at `burstCapacity`, subtracts `requestedTokens`, and returns whether the request is allowed plus tokens remaining. On allow, the filter adds informational headers: `X-RateLimit-Remaining`, `X-RateLimit-Burst-Capacity`, `X-RateLimit-Replenish-Rate` (and requested-tokens). On deny it sets status 429 and short-circuits without calling downstream. Keys get a TTL so idle buckets self-expire. A design consideration: if Redis is unreachable, the default behavior tends to fail *open* (allow), so you must decide whether that's acceptable.
code
yaml · 11 linesfilters:
- name: RequestRateLimiter
args:
redis-rate-limiter.replenishRate: 10
redis-rate-limiter.burstCapacity: 20
redis-rate-limiter.requestedTokens: 1
# informational headers (default true)
redis-rate-limiter.include-headers: true
# optional: override the denial status (default 429)
# status-code: TOO_MANY_REQUESTS
key-resolver: "#{@ipKeyResolver}"go deeper
Know it uses Redis and returns 429 with X-RateLimit headers.
Explain the token/timestamp state and the informational headers.
Articulate the distributed race and why the Lua script's atomicity solves it; note TTL and headers.
Weigh fail-open vs fail-closed on Redis outage, Retry-After ergonomics, and multi-region Redis topology.
**The consistency problem.** A gateway typically runs as N replicas. If each replica independently read a counter, decremented locally, and wrote it back, two replicas could read the same value concurrently and both allow a request that should have exhausted the bucket — a classic **read-modify-write race** that lets clients exceed the limit by up to N×. **How RedisRateLimiter solves it.** It stores, per bucket key, two fields: the current **token count** and the **timestamp** of the last refill. All the logic lives in a **Lua script** that Redis executes. Redis runs commands (and Lua scripts) on a **single thread**, and a Lua script executes **atomically** — no other command interleaves while it runs. The script does, in one shot: 1. Read stored tokens + last-refill time (and current time). 2. Compute how many tokens have regenerated: `elapsed_seconds * replenishRate`, added to the stored count, capped at `burstCapacity`. 3. If available tokens >= `requestedTokens`, subtract them and mark the request **allowed**; else **denied** (tokens unchanged, or only the timestamp/refill applied). 4. Write the new token count + timestamp back, and set/refresh a **TTL** on the keys. 5. Return `[allowed, tokensLeft]`. Because steps 1-4 are one atomic unit shared through Redis, every gateway instance sees a single authoritative bucket — no double-spend, no lost decrements. **Continuous refill.** Note the refill is *time-based*, computed from elapsed seconds, not a scheduled job. An idle bucket 'fills up' lazily on the next request. This is what makes it a proper token bucket rather than a fixed window. **Response headers.** On an allowed request RedisRateLimiter (when `include-headers` is true, the default) sets: - `X-RateLimit-Remaining` — tokens left in the bucket. - `X-RateLimit-Replenish-Rate` — configured replenishRate. - `X-RateLimit-Burst-Capacity` — configured burstCapacity. - `X-RateLimit-Requested-Tokens` — cost of this request. Header names are configurable via properties (e.g. `redis-rate-limiter.remaining-header`). You can disable them by setting `include-headers: false`. **Denial behavior.** When the script reports not allowed, the filter sets the response status to the configured code — **HTTP 429 Too Many Requests** by default (`status-code` property lets you change it) — and short-circuits the filter chain so the **downstream service is never invoked**. The remaining header typically shows 0. **TTL / memory.** Buckets carry an expiry so keys for clients that stop sending traffic are reclaimed automatically; otherwise Redis would accumulate a key per distinct KeyResolver value forever. **Failure mode — Redis down.** If Redis is unavailable or the script errors, the reactive pipeline errors; the common/default outcome is that the limiter lets the request through (**fail-open**) rather than failing all traffic closed. This is a deliberate availability trade-off but a security consideration — for sensitive endpoints you may want to fail closed. Know your version's exact behavior and consider a fallback. **Gotchas:** - Clock/time is taken consistently (the script uses a time source) so per-node clock skew doesn't corrupt refill math. - Headers reveal your limits to clients — fine for API ergonomics, but some prefer to hide them. - 429 responses ideally include a `Retry-After`; Gateway doesn't add it automatically, so add it via a filter if clients need it. - The atomic guarantee is per key; unrelated keys scale out naturally across Redis. **When this matters in interviews:** demonstrating you understand *why* Lua/atomicity is needed (the distributed race), the observable contract (429 + headers), and the operational risk (Redis dependency, fail-open) separates a senior answer from a config-recital.
- Why must the refill-and-decrement run as a Lua script rather than separate GET/SET commands?Separate commands allow interleaving between gateway instances, causing a read-modify-write race where two nodes both allow a request. A Lua script executes atomically on Redis's single thread, so the whole check-and-update is indivisible.
- What happens to rate limiting if Redis becomes unreachable?The default tendency is to fail open — requests are allowed — trading enforcement for availability. For sensitive routes you should decide whether to fail closed and design a fallback accordingly.
- Which headers does a client see, and can you turn them off?X-RateLimit-Remaining, X-RateLimit-Burst-Capacity, X-RateLimit-Replenish-Rate (and requested-tokens). Set redis-rate-limiter.include-headers: false to suppress them, and header names are configurable.
saying these in an interview costs you the question
- Claiming atomicity comes from a distributed lock the gateway holds, not the Redis Lua script
- Saying each gateway node keeps its own counter and they periodically sync
- Assuming Redis failure blocks all traffic (fail-closed) by default
- Thinking a background scheduler refills buckets rather than lazy time-based refill