skip to content

Reactive Routing on WebFlux

The gateway runs on Netty and WebFlux so it can proxy thousands of concurrent connections with backpressure and no thread per request. Interviewers ask why an edge component in particular has to be non-blocking.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

questions

5

Why does Spring Cloud Gateway run on the reactive WebFlux/Netty stack instead of the Spring MVC servlet stack?

level: juniorimportance: must knowfreq 70%

answer

  1. gateway = mostly waiting = I/O-bound
  2. servlet: thread-per-request, blocks while waiting
  3. Netty event loop: few threads, never block
  4. concurrency scales with active work not in-flight count
  5. bonus: end-to-end streaming + backpressure

basics

~20 s

A gateway spends its time waiting on downstream services, not computing. WebFlux on Netty handles many waiting requests with a few non-blocking threads, so the gateway scales far better than a servlet with one blocking thread per request.

solid answer

~50 s

A gateway is almost pure I/O: it forwards a request to a downstream service and waits for the response. On the classic Spring MVC servlet stack, each in-flight request pins one thread that blocks while waiting, so thousands of concurrent proxied calls need thousands of threads — expensive in memory and context switching. Spring Cloud Gateway is built on Spring WebFlux and runs on Netty, whose event loop uses a small fixed pool of threads that never block: while one request waits for a downstream response, that thread serves others. This lets a handful of threads carry very high concurrency, which is exactly the workload a gateway faces. It also lets the gateway stream and apply backpressure end to end rather than buffering whole bodies. That efficiency under fan-out I/O is the reason the reactive stack was chosen over servlets.

go deeper

for a junior

Should grasp the core intuition: a gateway mostly waits, and non-blocking threads let a few threads serve many waiting requests.

for a middle

Should name the servlet thread-per-request model vs Netty event loop and explain the resource-scaling difference.

for a senior

Should add streaming/backpressure benefits and the event-loop-blocking hazard.

for a principal

Should weigh the tradeoff (harder programming model, blocking-library risk) and note the MVC gateway variant exists for teams that want blocking code.

## What Spring Cloud Gateway is Spring Cloud Gateway is an API gateway — a reverse proxy that sits in front of backend services. A client calls the gateway, the gateway matches the request to a **Route** (predicates decide *which* route, filters transform the request/response), and it proxies the call to a downstream URI, then relays the response back. ## The two Spring web stacks Spring has two web runtimes: - **Spring MVC (servlet stack):** built on the Servlet API, typically Tomcat. The classic model is **thread-per-request** — one worker thread is dedicated to a request from start to finish. If the handler waits on I/O (e.g. a slow downstream), that thread is **blocked** and can do nothing else. - **Spring WebFlux (reactive stack):** built on Project Reactor (`Mono`/`Flux`) and runs by default on **Netty**, a non-blocking network server. Netty uses an **event loop**: a small, fixed pool of threads (roughly one per CPU core). Work is expressed as callbacks; when an operation would block (waiting for bytes from the network), the thread is released to do other work and resumes when data arrives. ## Why this matters for a gateway A gateway's job is overwhelmingly **I/O-bound, not CPU-bound**: for each request it opens a connection to a downstream service and *waits*. The dominant cost is waiting, not computing. - On the **servlet** model, N concurrent proxied requests need ~N threads, each parked while waiting. Threads cost ~0.5–1 MB of stack each and add context-switching overhead. A fan-out gateway handling tens of thousands of concurrent calls would need a huge thread pool or would queue and stall. - On the **reactive** model, those same N waiting requests are multiplexed over a handful of event-loop threads. A thread is only busy for the microseconds it takes to hand off or process a chunk; the rest of the time it serves other requests. Memory and CPU scale with *active* work, not with the number of *in-flight* requests. ## Extra benefits beyond thread count - **End-to-end streaming and backpressure:** request and response bodies flow as a `Flux<DataBuffer>`. The gateway can stream bytes through without buffering the whole payload, and reactor-netty propagates backpressure so a slow client or slow downstream throttles the other side instead of exhausting memory. - **Composable async pipeline:** filters return `Mono<Void>`, so pre/post logic composes naturally around the reactive proxy call. ## Gotchas / when it bites - The reactive advantage assumes **nobody blocks the event loop**. If a filter does blocking work (a JDBC call, `Thread.sleep`, blocking HTTP client), it stalls an event-loop thread and can collapse throughput for *all* requests on that loop. - The programming model is harder — you must think in `Mono`/`Flux`, and blocking libraries are unsafe unless offloaded to a bounded elastic scheduler. - There is also a newer **Spring Cloud Gateway Server MVC** variant that runs on the servlet stack for teams that prefer blocking code; but the classic, most common gateway is the reactive one and this question is about why *that* one is reactive. ## When the reactive choice pays off High connection concurrency with lots of waiting (many downstreams, slow backends, streaming/SSE, long-lived connections). For low-concurrency internal proxying the difference is small — the reactive stack shines precisely at gateway-scale fan-out I/O.

  • If the reactive gateway uses so few threads, what happens if one filter makes a blocking JDBC call?
    It parks an event-loop thread. Since the pool is tiny (about one per core), a few concurrent blocking calls can starve the loop and tank throughput for every request on it. Blocking work must be offloaded (e.g. Reactor's boundedElastic scheduler) or avoided.
  • Does reactive make a single request faster than servlet?
    No — per-request latency is roughly the same. The win is throughput and resource efficiency under high concurrency, because far fewer threads are needed to keep many waiting requests in flight.

saying these in an interview costs you the question

  • Claiming reactive makes each individual request faster (it's about concurrency/resource use, not single-call latency)
  • Saying Spring Cloud Gateway runs on Tomcat servlets by default
  • Thinking more threads always means more throughput

context

open as a page

Walk through the reactive request pipeline in Spring Cloud Gateway: how does an incoming request flow through ServerWebExchange and the GatewayFilterChain to become a proxied call?

level: middleimportance: must knowfreq 60%

basics

~20 s

A handler mapping matches the request to a route, then a filter chain of global + route filters runs. Each filter can modify the request before calling the next, and modify the response after. A routing filter actually proxies the call; the whole flow returns a Mono.

open as a page

The reactive gateway runs on a small Netty event-loop thread pool. What are the consequences of blocking inside a filter, and how do you handle work that must block?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Netty has only a few event-loop threads (about one per core). If a filter blocks — a JDBC call, Thread.sleep, a blocking HTTP client — it parks one of those threads, so many requests stall at once. Offload blocking work to a bounded elastic scheduler, or avoid it.

open as a page

How does Spring Cloud Gateway proxy request and response bodies in a backpressure-aware, streaming way rather than buffering entire payloads?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Bodies flow as a stream of DataBuffer chunks (a Flux), not one big blob. The gateway passes those chunks through reactor-netty, which only pulls more from the source when the destination is ready, so a slow client or downstream naturally throttles the fast side.

open as a page

Spring Cloud Gateway ships both a reactive (WebFlux/Netty) server and a servlet (Spring MVC) server. As an architect, how do you decide which to run, and what does the reactive stack actually buy you?

level: principalimportance: should knowfreq 30%

basics

~20 s

The reactive gateway scales huge concurrency with few threads and streams bodies with backpressure — ideal for high fan-out, slow backends, or streaming. The MVC gateway uses familiar blocking code and libraries. Choose reactive for scale/streaming; MVC when the team relies on blocking stacks and concurrency is moderate.

open as a page