In RedisRateLimiter, what do replenishRate and burstCapacity mean, and how does the token-bucket algorithm use them?
answer
- replenish = tokens/sec = sustained rate
- burst = bucket max = spike size
- each request costs requestedTokens (default 1)
- burst==replenish -> no bursting; burst==0 -> block all
- continuous refill beats fixed-window boundary burst
basics
~20 sreplenishRate is how many tokens (requests) are added to the bucket per second — the steady allowed rate. burstCapacity is the maximum tokens the bucket can hold — the biggest short burst allowed. Each request spends one token; an empty bucket means 429.
solid answer
~40 sRedisRateLimiter implements a **token bucket**. Think of a bucket that refills at `replenishRate` tokens per second and can hold at most `burstCapacity` tokens. Every request consumes `requestedTokens` (default 1). If a token is available the request passes; if the bucket is empty the gateway returns 429. `replenishRate` sets the sustained throughput; `burstCapacity` sets how big a momentary spike is tolerated after an idle period. Setting `burstCapacity == replenishRate` disables bursting (strict steady rate). Setting `burstCapacity = 0` blocks all traffic. A common config is burst = 2-3x replenish to absorb spikes while keeping the average bounded. Refill is computed lazily from elapsed time using a Redis Lua script, so tokens accrue continuously rather than in fixed windows, avoiding the boundary-burst problem of naive fixed-window counters.
code
yaml · 7 linesfilters:
- name: RequestRateLimiter
args:
# 5 requests/sec sustained, up to 10 in a burst, each request costs 1 token
redis-rate-limiter.replenishRate: 5
redis-rate-limiter.burstCapacity: 10
redis-rate-limiter.requestedTokens: 1go deeper
Know replenishRate = allowed rate, burstCapacity = max spike, empty bucket = 429.
Explain token bucket, continuous refill, requestedTokens, and the special cases.
Contrast with fixed-window boundary burst and discuss sizing burst to backend capacity.
Reason about the Lua-script atomicity, TTL/memory, and weighting endpoints via requestedTokens.
**Token-bucket model.** RedisRateLimiter uses the classic **token-bucket** algorithm. Imagine a bucket: - It is refilled at a constant rate of **`replenishRate`** tokens per second. - It can never hold more than **`burstCapacity`** tokens (overflow is discarded). - Each incoming request tries to remove **`requestedTokens`** tokens (default `1`). If enough tokens are present, they are removed and the request is allowed. If not, the request is denied with **HTTP 429**. **The two knobs:** - **`replenishRate`** — the *sustained* average request rate you permit, in requests per second. Over the long run a client cannot exceed this because that is the speed at which tokens regenerate. - **`burstCapacity`** — the *maximum burst*. Because unused tokens accumulate (up to this cap) during quiet periods, a client that has been idle can fire off up to `burstCapacity` requests almost instantly, then is throttled back to `replenishRate`. **`requestedTokens`** — an optional third arg letting a single request cost more than one token, useful when different endpoints have different weights (an expensive query might cost 5 tokens). **Worked example:** `replenishRate=10`, `burstCapacity=20`. - Steady state: ~10 requests/second sustainable. - After sitting idle, the bucket fills to 20; a client can burst 20 requests immediately. - After that burst the bucket is empty and refills at 10/s, so further requests are limited to that rate until it recovers. **Special configurations:** - `burstCapacity == replenishRate` → no bursting; a strict, smooth rate. - `burstCapacity == 0` → **all** requests are denied (a hard block; sometimes used to disable a route). Note `burstCapacity` must be >= `replenishRate` for the limiter to make sense; a burst smaller than the replenish rate is a misconfiguration. - Larger `burstCapacity` → friendlier to legitimate spikes but allows bigger momentary load on the backend. **Why token bucket over fixed window?** A naive **fixed-window counter** (e.g. 'max 10 per second, reset each second') suffers the *boundary burst* problem: a client can send 10 at 0.999s and 10 more at 1.001s — 20 requests in ~2ms. Token bucket refills *continuously* based on elapsed time, so it smooths this out and enforces a genuine average rate while still permitting a bounded burst. **Implementation detail — atomicity.** RedisRateLimiter runs a **Lua script** in Redis that atomically: reads the stored token count and last-refill timestamp, computes how many tokens have regenerated since then, adds them (capped at burstCapacity), tries to subtract the requested tokens, and writes the new state back. Running as a single Lua script makes the check-and-decrement atomic across concurrent gateway instances, avoiding race conditions. The keys carry a TTL so idle buckets expire and don't leak memory. **Configuration:** ```yaml filters: - name: RequestRateLimiter args: redis-rate-limiter.replenishRate: 10 redis-rate-limiter.burstCapacity: 20 redis-rate-limiter.requestedTokens: 1 ``` **Gotchas:** - Forgetting that burst allows short-term overshoot of the backend's true capacity — size `burstCapacity` to what the backend can survive. - Setting `burstCapacity` below `replenishRate` produces confusing behavior; keep burst >= replenish. - The limit is per **KeyResolver key**, not per route globally.
- How do you configure a strict, non-bursty rate?Set burstCapacity equal to replenishRate. With no room to accumulate spare tokens, requests are allowed only at the steady replenish rate.
- What is requestedTokens used for?It lets a single request consume more than one token, so you can weight expensive endpoints more heavily than cheap ones against the same bucket.
- Why is token bucket preferred over a fixed-window counter here?Fixed windows allow a boundary burst (double the limit across a window edge). Token bucket refills continuously by elapsed time, enforcing a true average while still permitting a bounded burst.
saying these in an interview costs you the question
- Swapping the two: saying burstCapacity is the per-second rate
- Thinking tokens refill in discrete one-second windows rather than continuously
- Believing each request always costs exactly one token with no way to change it
- Claiming setting burstCapacity below replenishRate is normal/useful