skip to content

Rate Limiting

The rate-limiting filter uses a Redis token bucket with a replenish rate and burst capacity, keyed by principal or IP, returning 429 when exhausted. Interviewers ask why the counter lives in Redis rather than in memory.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

questions

5

What is the RequestRateLimiter filter in Spring Cloud Gateway, and what happens when a client exceeds the limit?

level: juniorimportance: must knowfreq 60%

answer

  1. GatewayFilter -> RateLimiter -> KeyResolver
  2. token bucket in Redis
  3. 429 Too Many Requests, short-circuit
  4. state shared across gateway nodes
  5. per-key bucket (user or IP)

basics

~20 s

RequestRateLimiter is a built-in gateway filter that caps how many requests a client can make in a time window. When the client sends too many, the gateway rejects the extra requests with HTTP 429 Too Many Requests instead of forwarding them.

solid answer

~40 s

RequestRateLimiter is a GatewayFilter you attach to a route in Spring Cloud Gateway. It delegates to a RateLimiter implementation — the default is RedisRateLimiter, which uses a token-bucket algorithm stored in Redis. A KeyResolver decides which 'bucket' a request counts against (e.g. per user or per IP). On each request the filter asks the limiter whether a token is available: if yes the request proceeds to the backend; if the bucket is empty the filter short-circuits and returns HTTP 429 Too Many Requests without calling the downstream service. This protects backends from overload and abuse. Because state lives in Redis, the limit is shared across all gateway instances rather than being per-node.

code

yaml · 14 lines
yaml
spring:
  cloud:
    gateway:
      routes:
        - id: user_api
          uri: http://user-service
          predicates:
            - Path=/users/**
          filters:
            - name: RequestRateLimiter
              args:
                redis-rate-limiter.replenishRate: 10
                redis-rate-limiter.burstCapacity: 20
                # key-resolver bean referenced by SpEL, e.g. "#{@userKeyResolver}"

go deeper

for a junior

Know it's a filter that returns 429 and stops excess requests reaching the backend.

for a middle

Explain the three pieces (filter, RedisRateLimiter, KeyResolver) and why Redis is used.

for a senior

Discuss per-key buckets, response headers, and the deny-empty-key behavior.

for a principal

Frame it within an edge-protection strategy and multi-instance consistency guarantees.

**Spring Cloud Gateway** is a reactive API gateway built on Spring WebFlux; it routes incoming requests to backend services and can transform them via **filters**. Filters come in two kinds: global filters (apply to all routes) and per-route **GatewayFilters**. **RequestRateLimiter** is one of the built-in per-route GatewayFilters. **What it does:** It limits the rate of requests allowed through a route. If a caller sends more than the configured allowance, the gateway rejects the excess with **HTTP 429 Too Many Requests** and does *not* forward them downstream. This shields backends from traffic spikes, abusive clients, and accidental floods. **Three collaborating pieces:** 1. **RequestRateLimiter (the filter)** — wired onto a route. On every matching request it consults a rate limiter and either lets the request continue or aborts with 429. 2. **RateLimiter (the algorithm)** — an interface. The default implementation is **RedisRateLimiter**, which keeps counters in **Redis** using a **token-bucket** approach (tokens refill at `replenishRate` per second up to `burstCapacity`). 3. **KeyResolver** — decides *whose* limit applies by returning a key (a `Mono<String>`). Requests with the same key share one bucket. Common keys: the authenticated principal (per-user limiting) or the client IP. **Configuration example (YAML):** ```yaml spring: cloud: gateway: routes: - id: api uri: http://backend predicates: - Path=/api/** filters: - name: RequestRateLimiter args: redis-rate-limiter.replenishRate: 10 redis-rate-limiter.burstCapacity: 20 ``` This allows a steady 10 req/s with short bursts up to 20. **Why Redis?** A gateway usually runs as multiple instances behind a load balancer. Storing the counters in Redis makes the limit *global* across instances — a client hitting node A and node B still shares one bucket. A purely in-memory limiter would give each node its own allowance, effectively multiplying the real limit by the instance count. **When a request is denied:** the filter sets the response status to 429 and short-circuits — the downstream service is never invoked. By default RedisRateLimiter also adds informational headers (e.g. `X-RateLimit-Remaining`, `X-RateLimit-Burst-Capacity`, `X-RateLimit-Replenish-Rate`) so clients can see their remaining budget. **Edge cases / gotchas:** - You must provide a `KeyResolver` bean, or requests may be denied (or, if the key is empty, allowed through, depending on `deny-empty-key`). - Requires a Redis dependency (`spring-boot-starter-data-redis-reactive`). - The limit is per-key, not global for the whole route, unless your KeyResolver returns a constant. **When to use:** protecting expensive or fragile backends, enforcing per-tenant fair use, and blunting brute-force or scraping attacks at the edge.

  • Why store the rate-limit state in Redis instead of in memory?
    Gateways run as multiple instances. Redis gives a single shared counter so the limit is enforced globally; in-memory counters would give each node its own allowance and effectively multiply the true limit by the number of instances.
  • What HTTP status code does the client receive when limited, and is the backend called?
    HTTP 429 Too Many Requests. The filter short-circuits, so the downstream backend is not invoked.

saying these in an interview costs you the question

  • Saying it returns 503 or 403 instead of 429
  • Claiming the request is still forwarded to the backend but with a warning
  • Thinking a single default limit applies globally rather than per-key

context

open as a page

What is a KeyResolver in Spring Cloud Gateway rate limiting, and how would you implement per-user and per-IP resolvers?

level: middleimportance: must knowfreq 55%

basics

~20 s

A KeyResolver returns a key that identifies which rate-limit bucket a request belongs to. Requests with the same key share one bucket. For per-user you return the user id/principal; for per-IP you return the client's IP address.

open as a page

In RedisRateLimiter, what do replenishRate and burstCapacity mean, and how does the token-bucket algorithm use them?

level: middleimportance: must knowfreq 65%

basics

~20 s

replenishRate is how many tokens (requests) are added to the bucket per second — the steady allowed rate. burstCapacity is the maximum tokens the bucket can hold — the biggest short burst allowed. Each request spends one token; an empty bucket means 429.

open as a page

How does RedisRateLimiter enforce limits atomically across gateway instances, and what response headers and denial behavior does it produce?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It runs a Redis Lua script that atomically refills the token bucket by elapsed time and decrements it, so concurrent gateway nodes can't race. On success it adds X-RateLimit-* headers; when tokens run out it sets HTTP 429 and stops the request.

open as a page

As an architect, how would you design a multi-tier rate-limiting strategy at the gateway, and what are the failure and fairness trade-offs of the RedisRateLimiter approach?

level: principalimportance: should knowfreq 25%

basics

~20 s

Layer limits: coarse per-IP limits on unauthenticated routes and finer per-user/per-tenant limits on authenticated APIs, sized to each backend's capacity. Weigh Redis as a shared dependency (latency, single point of failure, fail-open vs fail-closed) and per-key fairness against real client identity.

open as a page