What production concerns arise when running long-lived SSE endpoints at scale in Spring WebFlux, and how do you address them?
answer
- heartbeat comments beat idle timeouts (ALB 60s)
- proxy buffering off: X-Accel-Buffering: no, HTTP/2
- doOnCancel/doFinally = release cursors/subscriptions
- never block the Netty event loop
- multi-instance fan-out via Redis/Kafka broker
basics
~20 sLong-lived connections tie up sockets and cross proxies/load balancers that may buffer or time out idle streams. Address with heartbeats, sensible timeouts, disconnect cleanup (doOnCancel), non-blocking pipelines, resumption via Last-Event-ID, and disabling proxy response buffering.
solid answer
~50 sSSE endpoints hold connections open indefinitely, which shifts the bottleneck from CPU to concurrent-connection count and infrastructure behavior. Key concerns: (1) **Idle timeouts** — load balancers/proxies and Netty close idle streams; emit periodic comment heartbeats and set generous read timeouts. (2) **Proxy buffering** — reverse proxies (e.g., Nginx) may buffer the response and defeat streaming; disable buffering (`X-Accel-Buffering: no`) and ensure HTTP/1.1 chunked or HTTP/2. (3) **Resource leaks** — every abandoned connection can leak a subscription or DB cursor; use `doOnCancel`/`doFinally` to release. (4) **Event-loop starvation** — never block on a Netty thread; offload with `boundedElastic`. (5) **Delivery guarantees** — SSE is at-most-once unless you implement `Last-Event-ID` replay. (6) **Backpressure/memory** — bound hot sources with `onBackpressureLatest`/`sample`. (7) **Scaling/fan-out** — with multiple instances behind a LB, a client only sees its instance's events; use a shared broker (Redis/Kafka) or sticky routing.
code
java · 17 lines@GetMapping(path = "/notifications", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
public Flux<ServerSentEvent<Notification>> notifications(Principal user) {
// Events fanned out from a shared broker so any instance can serve any client
Flux<ServerSentEvent<Notification>> events = broker.subscribe(user.getName())
.onBackpressureLatest()
.map(n -> ServerSentEvent.<Notification>builder()
.id(n.id()).event("notification").data(n).build());
// Heartbeat keeps LB/proxy idle timers from closing a quiet stream
Flux<ServerSentEvent<Notification>> heartbeat = Flux.interval(Duration.ofSeconds(20))
.map(t -> ServerSentEvent.<Notification>builder().comment("keep-alive").build());
return Flux.merge(events, heartbeat)
.doOnCancel(() -> broker.unsubscribe(user.getName())) // release on disconnect
.doFinally(sig -> metrics.decrementOpenStreams());
}
// Also set response header X-Accel-Buffering: no at the proxy or via a filter for Nginx.go deeper
Know connections stay open a long time and need keep-alives.
List heartbeats, timeouts, and disconnect cleanup as concerns.
Address proxy buffering, event-loop blocking, resumption, and backpressure bounding concretely.
Architect multi-instance fan-out via a broker, define delivery-guarantee trade-offs, capacity-plan concurrent connections, and decide SSE vs WebSocket/RSocket/broker per requirements.
Running SSE at scale is less about the framework and more about the realities of many long-lived HTTP connections. Concerns and mitigations: **1. Idle timeouts and dead connections.** Load balancers (AWS ALB default 60s idle), reverse proxies, and even Reactor Netty have idle timeouts that close a stream with no traffic. A live feed that goes quiet looks idle. - *Mitigation:* emit periodic **heartbeat** frames — SSE comment lines (`: ping`) merged into the Flux via `Flux.merge(data, Flux.interval(...).map(t -> comment))`. Tune the LB idle timeout above the heartbeat interval. Also detect half-open connections; TCP keep-alive helps but app-level heartbeats are more reliable. **2. Reverse-proxy / CDN buffering.** Many proxies buffer upstream responses before forwarding, which destroys incremental delivery — the client gets a burst at the end instead of a stream. - *Mitigation:* for Nginx set `X-Accel-Buffering: no` (or `proxy_buffering off`); ensure the transport is HTTP/1.1 with chunked transfer-encoding or HTTP/2. Avoid gzip response compression on SSE (it can buffer). Bypass CDN caching for these paths. **3. Connection/resource limits.** Each open SSE stream consumes a socket/file descriptor and a live Reactor subscription. Thousands of concurrent streams are fine for non-blocking WebFlux, but you must raise OS FD limits (`ulimit -n`) and size Netty accordingly, and cap connections per user to prevent abuse. **4. Disconnect cleanup.** When a client vanishes (tab closed, network drop), a cancel signal propagates up the chain. Failing to react leaks upstream subscriptions, DB cursors (R2DBC), broker consumers, etc. - *Mitigation:* `doOnCancel` (client-initiated cancel) and `doFinally` (any terminal signal) to unsubscribe/close/release. **5. Event-loop starvation.** WebFlux serves all connections on a small pool of Netty event-loop threads. One blocking call (JDBC, `Thread.sleep`, synchronous HTTP) on that thread stalls every connection it hosts. - *Mitigation:* keep the pipeline non-blocking end-to-end; wrap unavoidable blocking work in `Mono.fromCallable(...).subscribeOn(Schedulers.boundedElastic())`. **6. Delivery guarantees & resumption.** SSE is inherently at-most-once during a disconnect window. Browsers auto-reconnect and send `Last-Event-ID`; you must implement server-side replay-from-id to get at-least-once. If you need strict ordering/exactly-once, SSE is the wrong tool — use a broker with acknowledgements. **7. Backpressure & memory bounding.** As covered elsewhere, hot sources need `onBackpressureLatest`/`onBackpressureDrop`/`sample`/`limitRate` so a slow client can't grow memory. **8. Horizontal scaling / fan-out.** Behind a load balancer with N instances, an event produced on instance A won't reach a client connected to instance B unless events are shared. - *Mitigation:* publish domain events to a **shared bus** (Redis Pub/Sub, Kafka, RabbitMQ) that every instance subscribes to and re-emits to its local SSE subscribers; or use sticky sessions (weaker). This decouples event production from the connection-holding instance. **9. Security & auth.** Browser `EventSource` can't set headers, so bearer-token auth is awkward — rely on cookies (with CSRF care) or a token query param (careful: it lands in logs). Enforce per-connection authorization and rate limits. **10. Observability.** Track concurrent open streams, per-stream lifetime, cancel rate, and heartbeat gaps as metrics (Micrometer). A rising open-connection count with low throughput often signals leaked/abandoned streams. **When to pick something else.** For bidirectional, high-frequency, or guaranteed-delivery needs, prefer WebSocket, RSocket (native reactive backpressure over the wire), or a message broker with a thin push layer.
- You scale to three instances behind a load balancer and users report missing events. What's happening and how do you fix it?Each SSE connection is pinned to one instance; an event produced on another instance never reaches that client. Fix by publishing events to a shared bus (Redis Pub/Sub, Kafka) that every instance subscribes to and re-emits to its local SSE subscribers, decoupling event production from the connection-holding node. Sticky sessions are a weaker alternative.
- A reverse proxy is buffering your SSE responses so clients get bursts instead of a live stream. What do you change?Disable proxy response buffering — for Nginx set the X-Accel-Buffering: no response header (or proxy_buffering off), ensure HTTP/1.1 chunked or HTTP/2 transport, and turn off gzip compression on the SSE path so bytes are forwarded as they're flushed.
- Why is a blocking JDBC call inside an SSE stream especially dangerous in WebFlux?WebFlux serves many connections on a few Netty event-loop threads. A blocking call parks that thread, stalling every connection it hosts — not just the current one. Offload blocking work to Schedulers.boundedElastic via subscribeOn, or use reactive R2DBC instead.
saying these in an interview costs you the question
- Assuming SSE gives exactly-once/guaranteed delivery
- Ignoring proxy/LB idle timeouts (no heartbeat)
- No disconnect cleanup, leaking subscriptions/cursors
- Blocking the event loop with JDBC/sleep
- Expecting cross-instance fan-out without a shared broker
- Enabling gzip/proxy buffering on the stream and wondering why it's not live