As an architect, how would you design a multi-tier rate-limiting strategy at the gateway, and what are the failure and fairness trade-offs of the RedisRateLimiter approach?
answer
- tiers: pre-auth by IP, authed by principal/tenant
- size to backend capacity; burst = survivable spike
- Redis = shared dep: latency, SPOF, fail-open vs fail-closed
- fairness: NAT vs per-user; combine keys
- defense in depth: CDN/WAF + gateway + per-service + circuit breaker
basics
~20 sLayer limits: coarse per-IP limits on unauthenticated routes and finer per-user/per-tenant limits on authenticated APIs, sized to each backend's capacity. Weigh Redis as a shared dependency (latency, single point of failure, fail-open vs fail-closed) and per-key fairness against real client identity.
solid answer
~50 sI'd rate-limit in tiers by endpoint class and identity: pre-auth endpoints (login, signup, password reset) limited by IP to blunt brute force; authenticated APIs limited by principal or API key for fair per-account quotas; multi-tenant systems by tenant with per-plan replenishRate/burstCapacity. Sizing: replenishRate to the backend's sustainable throughput divided across expected keys, burstCapacity to what the backend survives momentarily. Key trade-offs of RedisRateLimiter: (1) Redis is a shared dependency — adds a network hop per request and is a potential SPOF; decide fail-open (availability) vs fail-closed (protection) and provide HA Redis. (2) Fairness — per-IP punishes shared NAT, per-principal needs auth first and misses pre-auth floods; often combine keys. (3) Weighted costs via requestedTokens for expensive endpoints. (4) Observability — emit metrics on 429 rates and near-limit clients. Consider WAF/CDN limits in front and per-service limits behind as defense in depth.
code
java · 10 lines// Compose identity + route so users get fair per-endpoint buckets,
// falling back to client IP for unauthenticated (pre-auth) traffic.
@Bean
KeyResolver tieredKeyResolver() {
return exchange -> exchange.getPrincipal()
.map(p -> "user:" + p.getName())
.switchIfEmpty(Mono.fromSupplier(() ->
"ip:" + clientIp(exchange))) // trusted X-Forwarded-For parse
.map(id -> id + ":" + exchange.getRequest().getPath().value());
}go deeper
Beyond scope; know that limits can differ per endpoint and identity.
Can size replenish/burst and pick IP vs principal keys per endpoint.
Reasons about fairness, combined keys, headers/Retry-After, and Redis HA.
Owns the end-to-end strategy: tiering, failure posture, topology, observability, and defense-in-depth placement.
**Framing.** Gateway rate limiting is one layer of a defense-in-depth throttling strategy, not the whole thing. A principal answer treats it as a system design problem: *what* to limit, *by whom*, *how much*, *what happens on failure*, and *how you observe it*. **1. Tiered strategy by endpoint class and identity.** - **Unauthenticated / pre-auth endpoints** (login, signup, forgot-password, token issuance): there is no principal yet, so limit **by IP** (via a trusted `X-Forwarded-For` resolver). This blunts brute-force and credential-stuffing. Keep these limits tight. - **Authenticated APIs**: limit **by principal or API key** so each account gets a fair, isolated quota — abuse by one user can't starve others. Sized generously relative to normal usage. - **Multi-tenant / plan-based**: key **by tenant** and vary `replenishRate`/`burstCapacity` per subscription tier (free vs enterprise). This usually means selecting rate parameters dynamically rather than static YAML — you can supply a custom `RateLimiter` or configuration source. - **Expensive operations**: use `requestedTokens` to weight heavy endpoints so a few costly calls consume the budget of many cheap ones. **2. Sizing replenishRate / burstCapacity.** Work backward from **backend capacity**. If a service sustains ~1000 req/s and you expect 200 concurrently-active keys, a naive per-key `replenishRate` around 5/s keeps aggregate near capacity — but real traffic is uneven, so leave headroom and rely on `burstCapacity` for spikes sized to what the backend tolerates without tipping over. Over-generous burst defeats the purpose; too tight harms legitimate bursty clients (page loads firing many calls). **3. Redis as a shared dependency — the core trade-off.** - **Latency:** every limited request incurs a Redis round trip. Co-locate Redis with the gateway; use pipelining/connection pooling; the Lua script keeps it to one call. - **Availability / SPOF:** Redis outage affects all traffic. Decide the **failure posture**: the default tends to **fail open** (allow), maximizing availability but removing protection exactly when you might be under attack; **fail closed** protects backends but converts a Redis blip into a full outage. Often the right answer is *fail open for user-facing read APIs, fail closed for sensitive mutating/auth endpoints*, plus **HA Redis** (replication/cluster) and possibly a degraded local limiter fallback. - **Consistency vs. partitioning:** Redis gives strong per-key atomicity (Lua). In multi-region deployments, a single global Redis adds cross-region latency; regional Redis clusters trade global exactness for locality (a client could get its limit per region). Choose based on whether the limit must be globally exact. **4. Fairness pitfalls.** - **Per-IP:** shared NAT / corporate proxies / CGNAT group many users into one bucket (collateral throttling); IPv6 and mobile complicate it; forged `X-Forwarded-For` can spoof or poison buckets if trusted from untrusted hops. - **Per-principal:** fair per account but useless before auth and blind to distributed attacks from many accounts. - **Combined keys:** e.g. `principal + ':' + route` or `ip + ':' + user` to get both fairness and abuse resistance. **5. Client contract & ergonomics.** Return **429** with meaningful headers (`X-RateLimit-*`) and ideally a **`Retry-After`** (Gateway won't add it automatically — inject via a filter). Document limits so clients implement backoff instead of hammering. **6. Observability & tuning.** Emit metrics: 429 rate per route/key, buckets frequently at zero, top talkers. Feed this back into sizing. Alert on sudden 429 spikes (attack or a misbehaving client) and on Redis errors (limiter degraded). **7. Defense in depth.** Gateway limiting complements — not replaces — **CDN/WAF** edge limits (absorb volumetric attacks before your infra), **per-service** limits (protect a service even from internal callers), and **circuit breakers/bulkheads** (Resilience4j) for downstream protection. Rate limiting caps *arrival rate*; circuit breakers handle *downstream failure*. **8. Alternatives / when not to use it.** For extremely high throughput where a Redis hop per request is too costly, consider local token buckets with approximate coordination, or sticky routing so a key stays on one node. For simple single-instance deployments, an in-memory limiter avoids the Redis dependency entirely. **Bottom line:** the RedisRateLimiter gives correct, shared, atomic token-bucket limiting; the architectural work is choosing keys per endpoint class, sizing to backend capacity, and consciously picking the Redis failure posture and topology.
- For a login endpoint versus a read-heavy authenticated API, would you fail open or fail closed on a Redis outage, and why?Login: lean toward fail closed (or a strict local fallback) — losing brute-force protection during an outage is dangerous. Read API: often fail open to preserve availability, since the cost of over-serving reads is lower than a full outage. The point is to choose consciously per endpoint class rather than accept one global default.
- How would you give different tenants different limits without hardcoding YAML per route?Resolve the tenant/plan in the KeyResolver or a custom RateLimiter that looks up replenishRate/burstCapacity from a config source (DB/cache) keyed by plan, so limits are data-driven and changeable at runtime.
- Where does gateway rate limiting sit relative to a WAF/CDN and circuit breakers?CDN/WAF absorbs volumetric/edge attacks first; the gateway enforces per-identity fair-use; per-service limits protect individual services; circuit breakers (Resilience4j) handle downstream failure. They're complementary layers, not substitutes.
saying these in an interview costs you the question
- Treating gateway rate limiting as sufficient alone, ignoring WAF/CDN and per-service layers
- Not acknowledging Redis as a latency cost and single point of failure
- Applying one uniform limit/key to all endpoints regardless of auth state
- Ignoring the fail-open default and its security implications
- Assuming per-IP limiting is always fair despite shared NAT/CGNAT