skip to content

What is the RequestRateLimiter filter in Spring Cloud Gateway, and what happens when a client exceeds the limit?

level: juniorimportance: must knowfreq 60%

answer

  1. GatewayFilter -> RateLimiter -> KeyResolver
  2. token bucket in Redis
  3. 429 Too Many Requests, short-circuit
  4. state shared across gateway nodes
  5. per-key bucket (user or IP)

basics

~20 s

RequestRateLimiter is a built-in gateway filter that caps how many requests a client can make in a time window. When the client sends too many, the gateway rejects the extra requests with HTTP 429 Too Many Requests instead of forwarding them.

solid answer

~40 s

RequestRateLimiter is a GatewayFilter you attach to a route in Spring Cloud Gateway. It delegates to a RateLimiter implementation — the default is RedisRateLimiter, which uses a token-bucket algorithm stored in Redis. A KeyResolver decides which 'bucket' a request counts against (e.g. per user or per IP). On each request the filter asks the limiter whether a token is available: if yes the request proceeds to the backend; if the bucket is empty the filter short-circuits and returns HTTP 429 Too Many Requests without calling the downstream service. This protects backends from overload and abuse. Because state lives in Redis, the limit is shared across all gateway instances rather than being per-node.

code

yaml · 14 lines
yaml
spring:
  cloud:
    gateway:
      routes:
        - id: user_api
          uri: http://user-service
          predicates:
            - Path=/users/**
          filters:
            - name: RequestRateLimiter
              args:
                redis-rate-limiter.replenishRate: 10
                redis-rate-limiter.burstCapacity: 20
                # key-resolver bean referenced by SpEL, e.g. "#{@userKeyResolver}"

go deeper

for a junior

Know it's a filter that returns 429 and stops excess requests reaching the backend.

for a middle

Explain the three pieces (filter, RedisRateLimiter, KeyResolver) and why Redis is used.

for a senior

Discuss per-key buckets, response headers, and the deny-empty-key behavior.

for a principal

Frame it within an edge-protection strategy and multi-instance consistency guarantees.

**Spring Cloud Gateway** is a reactive API gateway built on Spring WebFlux; it routes incoming requests to backend services and can transform them via **filters**. Filters come in two kinds: global filters (apply to all routes) and per-route **GatewayFilters**. **RequestRateLimiter** is one of the built-in per-route GatewayFilters. **What it does:** It limits the rate of requests allowed through a route. If a caller sends more than the configured allowance, the gateway rejects the excess with **HTTP 429 Too Many Requests** and does *not* forward them downstream. This shields backends from traffic spikes, abusive clients, and accidental floods. **Three collaborating pieces:** 1. **RequestRateLimiter (the filter)** — wired onto a route. On every matching request it consults a rate limiter and either lets the request continue or aborts with 429. 2. **RateLimiter (the algorithm)** — an interface. The default implementation is **RedisRateLimiter**, which keeps counters in **Redis** using a **token-bucket** approach (tokens refill at `replenishRate` per second up to `burstCapacity`). 3. **KeyResolver** — decides *whose* limit applies by returning a key (a `Mono<String>`). Requests with the same key share one bucket. Common keys: the authenticated principal (per-user limiting) or the client IP. **Configuration example (YAML):** ```yaml spring: cloud: gateway: routes: - id: api uri: http://backend predicates: - Path=/api/** filters: - name: RequestRateLimiter args: redis-rate-limiter.replenishRate: 10 redis-rate-limiter.burstCapacity: 20 ``` This allows a steady 10 req/s with short bursts up to 20. **Why Redis?** A gateway usually runs as multiple instances behind a load balancer. Storing the counters in Redis makes the limit *global* across instances — a client hitting node A and node B still shares one bucket. A purely in-memory limiter would give each node its own allowance, effectively multiplying the real limit by the instance count. **When a request is denied:** the filter sets the response status to 429 and short-circuits — the downstream service is never invoked. By default RedisRateLimiter also adds informational headers (e.g. `X-RateLimit-Remaining`, `X-RateLimit-Burst-Capacity`, `X-RateLimit-Replenish-Rate`) so clients can see their remaining budget. **Edge cases / gotchas:** - You must provide a `KeyResolver` bean, or requests may be denied (or, if the key is empty, allowed through, depending on `deny-empty-key`). - Requires a Redis dependency (`spring-boot-starter-data-redis-reactive`). - The limit is per-key, not global for the whole route, unless your KeyResolver returns a constant. **When to use:** protecting expensive or fragile backends, enforcing per-tenant fair use, and blunting brute-force or scraping attacks at the edge.

  • Why store the rate-limit state in Redis instead of in memory?
    Gateways run as multiple instances. Redis gives a single shared counter so the limit is enforced globally; in-memory counters would give each node its own allowance and effectively multiply the true limit by the number of instances.
  • What HTTP status code does the client receive when limited, and is the backend called?
    HTTP 429 Too Many Requests. The filter short-circuits, so the downstream backend is not invoked.

saying these in an interview costs you the question

  • Saying it returns 503 or 403 instead of 429
  • Claiming the request is still forwarded to the backend but with a warning
  • Thinking a single default limit applies globally rather than per-key

context