skip to content

Why does Spring Cloud Gateway run on the reactive WebFlux/Netty stack instead of the Spring MVC servlet stack?

level: juniorimportance: must knowfreq 70%

answer

  1. gateway = mostly waiting = I/O-bound
  2. servlet: thread-per-request, blocks while waiting
  3. Netty event loop: few threads, never block
  4. concurrency scales with active work not in-flight count
  5. bonus: end-to-end streaming + backpressure

basics

~20 s

A gateway spends its time waiting on downstream services, not computing. WebFlux on Netty handles many waiting requests with a few non-blocking threads, so the gateway scales far better than a servlet with one blocking thread per request.

solid answer

~50 s

A gateway is almost pure I/O: it forwards a request to a downstream service and waits for the response. On the classic Spring MVC servlet stack, each in-flight request pins one thread that blocks while waiting, so thousands of concurrent proxied calls need thousands of threads — expensive in memory and context switching. Spring Cloud Gateway is built on Spring WebFlux and runs on Netty, whose event loop uses a small fixed pool of threads that never block: while one request waits for a downstream response, that thread serves others. This lets a handful of threads carry very high concurrency, which is exactly the workload a gateway faces. It also lets the gateway stream and apply backpressure end to end rather than buffering whole bodies. That efficiency under fan-out I/O is the reason the reactive stack was chosen over servlets.

go deeper

for a junior

Should grasp the core intuition: a gateway mostly waits, and non-blocking threads let a few threads serve many waiting requests.

for a middle

Should name the servlet thread-per-request model vs Netty event loop and explain the resource-scaling difference.

for a senior

Should add streaming/backpressure benefits and the event-loop-blocking hazard.

for a principal

Should weigh the tradeoff (harder programming model, blocking-library risk) and note the MVC gateway variant exists for teams that want blocking code.

## What Spring Cloud Gateway is Spring Cloud Gateway is an API gateway — a reverse proxy that sits in front of backend services. A client calls the gateway, the gateway matches the request to a **Route** (predicates decide *which* route, filters transform the request/response), and it proxies the call to a downstream URI, then relays the response back. ## The two Spring web stacks Spring has two web runtimes: - **Spring MVC (servlet stack):** built on the Servlet API, typically Tomcat. The classic model is **thread-per-request** — one worker thread is dedicated to a request from start to finish. If the handler waits on I/O (e.g. a slow downstream), that thread is **blocked** and can do nothing else. - **Spring WebFlux (reactive stack):** built on Project Reactor (`Mono`/`Flux`) and runs by default on **Netty**, a non-blocking network server. Netty uses an **event loop**: a small, fixed pool of threads (roughly one per CPU core). Work is expressed as callbacks; when an operation would block (waiting for bytes from the network), the thread is released to do other work and resumes when data arrives. ## Why this matters for a gateway A gateway's job is overwhelmingly **I/O-bound, not CPU-bound**: for each request it opens a connection to a downstream service and *waits*. The dominant cost is waiting, not computing. - On the **servlet** model, N concurrent proxied requests need ~N threads, each parked while waiting. Threads cost ~0.5–1 MB of stack each and add context-switching overhead. A fan-out gateway handling tens of thousands of concurrent calls would need a huge thread pool or would queue and stall. - On the **reactive** model, those same N waiting requests are multiplexed over a handful of event-loop threads. A thread is only busy for the microseconds it takes to hand off or process a chunk; the rest of the time it serves other requests. Memory and CPU scale with *active* work, not with the number of *in-flight* requests. ## Extra benefits beyond thread count - **End-to-end streaming and backpressure:** request and response bodies flow as a `Flux<DataBuffer>`. The gateway can stream bytes through without buffering the whole payload, and reactor-netty propagates backpressure so a slow client or slow downstream throttles the other side instead of exhausting memory. - **Composable async pipeline:** filters return `Mono<Void>`, so pre/post logic composes naturally around the reactive proxy call. ## Gotchas / when it bites - The reactive advantage assumes **nobody blocks the event loop**. If a filter does blocking work (a JDBC call, `Thread.sleep`, blocking HTTP client), it stalls an event-loop thread and can collapse throughput for *all* requests on that loop. - The programming model is harder — you must think in `Mono`/`Flux`, and blocking libraries are unsafe unless offloaded to a bounded elastic scheduler. - There is also a newer **Spring Cloud Gateway Server MVC** variant that runs on the servlet stack for teams that prefer blocking code; but the classic, most common gateway is the reactive one and this question is about why *that* one is reactive. ## When the reactive choice pays off High connection concurrency with lots of waiting (many downstreams, slow backends, streaming/SSE, long-lived connections). For low-concurrency internal proxying the difference is small — the reactive stack shines precisely at gateway-scale fan-out I/O.

  • If the reactive gateway uses so few threads, what happens if one filter makes a blocking JDBC call?
    It parks an event-loop thread. Since the pool is tiny (about one per core), a few concurrent blocking calls can starve the loop and tank throughput for every request on it. Blocking work must be offloaded (e.g. Reactor's boundedElastic scheduler) or avoided.
  • Does reactive make a single request faster than servlet?
    No — per-request latency is roughly the same. The win is throughput and resource efficiency under high concurrency, because far fewer threads are needed to keep many waiting requests in flight.

saying these in an interview costs you the question

  • Claiming reactive makes each individual request faster (it's about concurrency/resource use, not single-call latency)
  • Saying Spring Cloud Gateway runs on Tomcat servlets by default
  • Thinking more threads always means more throughput

context