skip to content

How does Spring Cloud Gateway proxy request and response bodies in a backpressure-aware, streaming way rather than buffering entire payloads?

level: seniorimportance: should knowfreq 40%

answer

  1. body = Flux<DataBuffer>, chunked not blob
  2. reactor-netty maps demand -> TCP flow control
  3. slow side throttles fast side, both hops
  4. constant memory for plain proxy, even GBs/SSE
  5. ModifyRequestBody/CacheRequestBody buffer = break streaming

basics

~20 s

Bodies flow as a stream of DataBuffer chunks (a Flux), not one big blob. The gateway passes those chunks through reactor-netty, which only pulls more from the source when the destination is ready, so a slow client or downstream naturally throttles the fast side.

solid answer

~50 s

In WebFlux a body isn't a byte array — it's a `Flux<DataBuffer>`, a stream of network chunks. Spring Cloud Gateway proxies by wiring the inbound request body `Flux` into reactor-netty's `HttpClient` outbound, and the downstream response body `Flux` back to the client via `NettyWriteResponseFilter` — chunk by chunk, without materializing the whole payload. Backpressure is the Reactive Streams demand signal: the consumer requests N items and the producer emits at most N. reactor-netty ties this to TCP: if the downstream (or client) can't accept bytes, it stops requesting, the source stops reading from the socket, and TCP flow control throttles the sender. So a slow client streaming a huge upload can't overwhelm gateway memory. The exception is filters that must see the full body (e.g. `ModifyRequestBody`/`ModifyResponseBody`): those buffer, trading streaming for the ability to transform.

code

java · 26 lines
java
import org.springframework.cloud.gateway.route.RouteLocator;
import org.springframework.cloud.gateway.route.builder.RouteLocatorBuilder;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;

@Configuration
public class GatewayRoutes {

    @Bean
    RouteLocator routes(RouteLocatorBuilder builder) {
        return builder.routes()
            // Plain streaming proxy: request/response bodies flow as Flux<DataBuffer>,
            // constant memory even for large uploads/downloads.
            .route("stream-proxy", r -> r.path("/files/**")
                .uri("http://storage-service:8080"))

            // Contrast: modifyResponseBody must buffer the WHOLE body to transform it,
            // which defeats end-to-end streaming for this route.
            .route("transform", r -> r.path("/api/**")
                .filters(f -> f.modifyResponseBody(String.class, String.class,
                    (exchange, body) -> reactor.core.publisher.Mono.just(
                        body == null ? "" : body.toUpperCase())))
                .uri("http://api-service:8080"))
            .build();
    }
}

go deeper

for a junior

Not expected in depth; may just know bodies stream in chunks.

for a middle

Should know body is a Flux<DataBuffer> and plain proxy doesn't buffer everything.

for a senior

Should explain demand-driven streaming, reactor-netty bridging to TCP flow control, and which filters break streaming by buffering.

for a principal

Should reason about DataBuffer/native-memory leaks, single-subscription semantics, memory budgeting under large-body high concurrency, and SSE/streaming implications.

## Bodies as streams, not blobs On the servlet stack a body is often read as a full `byte[]`/`InputStream` into memory. On WebFlux a body is a **`Flux<DataBuffer>`** — an asynchronous sequence of `DataBuffer` chunks, where each `DataBuffer` wraps a slice of bytes (backed by Netty's pooled `ByteBuf`). Nothing forces the whole payload into memory at once. ## What backpressure means here **Backpressure** is a core Reactive Streams mechanism: a `Subscriber` signals *demand* via `request(n)`, and the `Publisher` must not emit more than requested. It lets a slow consumer throttle a fast producer instead of being flooded. (Reactor/WebFlux fundamentals are covered by the sibling WebFlux topic — here we care about how the *gateway* leverages them.) ## How the gateway streams a proxy call 1. The client's request body arrives as a `Flux<DataBuffer>` on the `ServerHttpRequest`. 2. The routing filter (`NettyRoutingFilter`) hands that `Flux` to reactor-netty's `HttpClient` as the outbound body. reactor-netty writes chunks to the downstream connection **as the downstream accepts them**. 3. The downstream response body comes back as another `Flux<DataBuffer>`. `NettyWriteResponseFilter` writes it to the client's `ServerHttpResponse` — again chunk by chunk. At no point (for a plain proxy) is the entire body buffered. Memory use is bounded by in-flight chunk/window size, not payload size. This is why a reactive gateway can relay multi-gigabyte uploads/downloads or long-lived Server-Sent Events streams with tiny, constant memory. ## Backpressure end to end (the important part) reactor-netty bridges Reactive Streams demand to **TCP flow control**: - If the **downstream** is slow to accept the request body, reactor-netty stops requesting chunks from the client's body `Flux`; that in turn stops reading from the client socket; TCP's receive window fills and the client's OS throttles its send. The slow side sets the pace. - Symmetrically, if the **client** is slow to consume the response, the gateway stops pulling response chunks from the downstream, throttling it. So a fast producer can never outrun a slow consumer through the gateway — pressure propagates backward across both hops. This bounds memory and protects the gateway from being a buffering bottleneck. ## When streaming is broken on purpose Some filters need the **whole** body to work: - `ModifyRequestBody` / `ModifyResponseBody` transform the payload, so they aggregate the `Flux<DataBuffer>` into a full body, transform it, then re-emit. This **buffers** and defeats streaming for that route. - `CacheRequestBody` deliberately reads and caches the body so multiple filters can consume it (a `Flux<DataBuffer>` is generally single-subscription — you can't naively read it twice). - `RequestSize` rejects oversized bodies. Use these knowingly: they trade constant memory for transformation ability, and a large body plus a modify filter can spike heap. ## Gotchas - **DataBuffer leaks:** `DataBuffer`s wrap pooled Netty buffers with reference counts. If a custom filter consumes a body but forgets to release buffers (`DataBufferUtils.release`) or fails to fully drain the `Flux`, you leak native memory. Framework filters handle this; custom ones must be careful. - **Single subscription:** the request body `Flux` typically can be subscribed once. Caching (`CacheRequestBody`) is required if several filters need it. - **Losing backpressure with unbounded buffering:** operators like `.collectList()` or a modify-body filter drop backpressure by materializing everything — fine for small bodies, dangerous for large/streaming ones. - **Head-of-line / streaming semantics:** for SSE or chunked responses, don't insert a buffering filter or you break the stream's real-time nature. ## When it matters Large uploads/downloads, streaming APIs (SSE/NDJSON), high concurrency with big bodies. For tiny JSON payloads the streaming vs buffering distinction is negligible; it becomes critical as body size and concurrency grow.

  • What does a ModifyResponseBody filter do to the streaming/backpressure story?
    It aggregates the response `Flux<DataBuffer>` into the full body to transform it, then re-emits. That buffers the entire payload, defeating end-to-end streaming and constant memory for that route — cheap for small JSON, risky for large or streaming responses.
  • How is Reactive Streams backpressure connected to the actual network in a gateway?
    reactor-netty ties Reactive Streams demand to TCP flow control. When the consumer stops requesting chunks, reactor-netty stops reading from the socket; the TCP receive window fills and the OS throttles the sender. So backpressure isn't just in-JVM — it propagates to the wire on both the client and downstream hops.

saying these in an interview costs you the question

  • Believing the gateway always buffers the full request/response into a byte[]
  • Thinking backpressure is only an in-memory Reactor concept with no link to TCP
  • Assuming a request body Flux can be freely subscribed multiple times
  • Not realizing ModifyRequestBody/ModifyResponseBody buffer the whole payload

context