How do you stream a large, incrementally-produced dataset as a WebClient request body without buffering it all in memory, and what media type makes it a true stream?
answer
- Flux + body(flux, Type)
- subscribe = streaming upload
- default JSON = single array
- NDJSON = per-record flush
- chunked, unknown Content-Length
basics
~20 sPass a Flux to body(flux, Type.class) so WebClient writes elements as they are emitted instead of collecting them. Use a streaming media type like application/x-ndjson (or text/event-stream) so each element is flushed separately rather than buffered into one JSON array.
solid answer
~40 sProvide the payload as a Flux<T> and attach it via body(flux, T.class) (or BodyInserters.fromPublisher). WebClient subscribes to the Flux and encodes each element through the codec pipeline, applying reactive backpressure so you never materialize the whole dataset. The media type decides framing: with the default application/json the Jackson encoder still produces a single JSON array (elements are written incrementally but it's one document). To get discrete, independently-flushed records use a streaming type — application/x-ndjson (MediaType.APPLICATION_NDJSON, newline-delimited JSON) or text/event-stream (SSE). Set it via .contentType(...). This keeps client memory flat and lets the server begin processing before the stream ends. Note the receiving server must also understand the streaming type/framing.
code
java · 12 linesFlux<Measurement> measurements = sensorRepo.streamAll(); // never collected
Mono<Void> upload = webClient.post()
.uri("/ingest")
.contentType(MediaType.APPLICATION_NDJSON) // per-record framing
.body(measurements, Measurement.class) // WebClient subscribes
.retrieve()
.bodyToMono(Void.class);
// Server side (Spring WebFlux) that streams it in:
// @PostMapping(value = "/ingest", consumes = MediaType.APPLICATION_NDJSON_VALUE)
// Mono<Void> ingest(@RequestBody Flux<Measurement> body) { ... }go deeper
Knows to pass a Flux to body(...) instead of a List.
Explains backpressure + that media type controls array-vs-NDJSON framing.
Adds chunked encoding, server must consume the same framing, mid-stream error semantics.
Reasons about idempotency/retry when a partially-sent stream fails, proxy/HTTP-version constraints, ordering guarantees.
**Goal.** Upload data that is large or produced over time (e.g. a `Flux` from an R2DBC query, a Kafka consumer, or a generator) without first collecting it into a `List` and holding it all in memory. **Mechanism.** 1. Represent the data as a **`Flux<T>`** (a reactive stream of 0..N elements with backpressure). 2. Attach it with **`body(Flux<T>, Class<T>)`** — equivalently `BodyInserters.fromPublisher(flux, Class)`. 3. When the request executes, WebClient **subscribes** to the `Flux`. Each emitted element is passed to an `Encoder` (e.g. `Jackson2JsonEncoder`) which turns it into `DataBuffer`s written to the outbound `ClientHttpRequest`. Reactor Netty applies **backpressure**: if the network/server is slow, WebClient requests fewer elements from the `Flux`, so a fast producer won't overwhelm memory. **Why media type matters — framing.** Streaming the *elements* is not the same as streaming the *wire format*: - **`application/json`** (default for a POJO/collection): `Jackson2JsonEncoder` writes a **single JSON array** `[{...},{...}]`. It writes incrementally as elements arrive, but semantically it's one document; the receiver typically must parse the whole array. Elements are not independently flushed as separate records. - **`application/x-ndjson`** (`MediaType.APPLICATION_NDJSON`, newline-delimited JSON): each element is encoded as its own JSON object followed by `\n` and **flushed independently**. This is the canonical 'stream of records' format. Use `Jackson2JsonEncoder` in streaming mode, triggered by this content type. - **`text/event-stream`** (`MediaType.TEXT_EVENT_STREAM`, Server-Sent Events): elements are wrapped as SSE `data:` frames. More common for responses than uploads. Set it with `.contentType(MediaType.APPLICATION_NDJSON)`. **Requesting side vs receiving side.** Streaming only works end to end if the **server** accepts the same framing. A Spring WebFlux `@PostMapping(consumes = MediaType.APPLICATION_NDJSON_VALUE)` controller taking a `@RequestBody Flux<T>` will consume it as a stream. If the server buffers, you still save client memory but not server memory. **Gotchas.** - Don't `.collectList().block()` the Flux first — that defeats the purpose and blocks a reactive thread. - `Content-Length` is usually unknown for a stream, so WebClient uses **chunked transfer encoding**. Some servers/proxies dislike chunked uploads. - Errors mid-stream: if the `Flux` errors after some elements were sent, the request body is already partially written — the HTTP request may be aborted/reset; design idempotency accordingly. - The element `Class`/`ParameterizedTypeReference` is required because generics are erased. - Ordering: `Flux` preserves emission order; concurrency operators (`flatMap`) can interleave, so use `concatMap` if order matters.
- With the default application/json content type, is a Flux streamed element-by-element on the wire?The encoder writes incrementally, but it produces one JSON array document. For truly independent, flushed records you need application/x-ndjson or text/event-stream framing.
- What transfer encoding does WebClient use when the body length is unknown?Chunked transfer encoding, since Content-Length can't be computed ahead of a stream. Be aware some proxies/servers restrict chunked request bodies.
saying these in an interview costs you the question
- collectList().block() before sending (defeats streaming, blocks reactive thread)
- Assuming default application/json flushes each element as a separate record
- Thinking a Content-Length is set for a streamed Flux body
- Using flatMap and expecting strict ordering