skip to content

What production concerns arise when running long-lived SSE endpoints at scale in Spring WebFlux, and how do you address them?

level: principalimportance: should knowfreq 30%

answer

  1. heartbeat comments beat idle timeouts (ALB 60s)
  2. proxy buffering off: X-Accel-Buffering: no, HTTP/2
  3. doOnCancel/doFinally = release cursors/subscriptions
  4. never block the Netty event loop
  5. multi-instance fan-out via Redis/Kafka broker

basics

~20 s

Long-lived connections tie up sockets and cross proxies/load balancers that may buffer or time out idle streams. Address with heartbeats, sensible timeouts, disconnect cleanup (doOnCancel), non-blocking pipelines, resumption via Last-Event-ID, and disabling proxy response buffering.

solid answer

~50 s

SSE endpoints hold connections open indefinitely, which shifts the bottleneck from CPU to concurrent-connection count and infrastructure behavior. Key concerns: (1) **Idle timeouts** — load balancers/proxies and Netty close idle streams; emit periodic comment heartbeats and set generous read timeouts. (2) **Proxy buffering** — reverse proxies (e.g., Nginx) may buffer the response and defeat streaming; disable buffering (`X-Accel-Buffering: no`) and ensure HTTP/1.1 chunked or HTTP/2. (3) **Resource leaks** — every abandoned connection can leak a subscription or DB cursor; use `doOnCancel`/`doFinally` to release. (4) **Event-loop starvation** — never block on a Netty thread; offload with `boundedElastic`. (5) **Delivery guarantees** — SSE is at-most-once unless you implement `Last-Event-ID` replay. (6) **Backpressure/memory** — bound hot sources with `onBackpressureLatest`/`sample`. (7) **Scaling/fan-out** — with multiple instances behind a LB, a client only sees its instance's events; use a shared broker (Redis/Kafka) or sticky routing.

code

java · 17 lines
java
@GetMapping(path = "/notifications", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
public Flux<ServerSentEvent<Notification>> notifications(Principal user) {
    // Events fanned out from a shared broker so any instance can serve any client
    Flux<ServerSentEvent<Notification>> events = broker.subscribe(user.getName())
        .onBackpressureLatest()
        .map(n -> ServerSentEvent.<Notification>builder()
                    .id(n.id()).event("notification").data(n).build());

    // Heartbeat keeps LB/proxy idle timers from closing a quiet stream
    Flux<ServerSentEvent<Notification>> heartbeat = Flux.interval(Duration.ofSeconds(20))
        .map(t -> ServerSentEvent.<Notification>builder().comment("keep-alive").build());

    return Flux.merge(events, heartbeat)
        .doOnCancel(() -> broker.unsubscribe(user.getName())) // release on disconnect
        .doFinally(sig -> metrics.decrementOpenStreams());
}
// Also set response header X-Accel-Buffering: no at the proxy or via a filter for Nginx.

go deeper

for a junior

Know connections stay open a long time and need keep-alives.

for a middle

List heartbeats, timeouts, and disconnect cleanup as concerns.

for a senior

Address proxy buffering, event-loop blocking, resumption, and backpressure bounding concretely.

for a principal

Architect multi-instance fan-out via a broker, define delivery-guarantee trade-offs, capacity-plan concurrent connections, and decide SSE vs WebSocket/RSocket/broker per requirements.

Running SSE at scale is less about the framework and more about the realities of many long-lived HTTP connections. Concerns and mitigations: **1. Idle timeouts and dead connections.** Load balancers (AWS ALB default 60s idle), reverse proxies, and even Reactor Netty have idle timeouts that close a stream with no traffic. A live feed that goes quiet looks idle. - *Mitigation:* emit periodic **heartbeat** frames — SSE comment lines (`: ping`) merged into the Flux via `Flux.merge(data, Flux.interval(...).map(t -> comment))`. Tune the LB idle timeout above the heartbeat interval. Also detect half-open connections; TCP keep-alive helps but app-level heartbeats are more reliable. **2. Reverse-proxy / CDN buffering.** Many proxies buffer upstream responses before forwarding, which destroys incremental delivery — the client gets a burst at the end instead of a stream. - *Mitigation:* for Nginx set `X-Accel-Buffering: no` (or `proxy_buffering off`); ensure the transport is HTTP/1.1 with chunked transfer-encoding or HTTP/2. Avoid gzip response compression on SSE (it can buffer). Bypass CDN caching for these paths. **3. Connection/resource limits.** Each open SSE stream consumes a socket/file descriptor and a live Reactor subscription. Thousands of concurrent streams are fine for non-blocking WebFlux, but you must raise OS FD limits (`ulimit -n`) and size Netty accordingly, and cap connections per user to prevent abuse. **4. Disconnect cleanup.** When a client vanishes (tab closed, network drop), a cancel signal propagates up the chain. Failing to react leaks upstream subscriptions, DB cursors (R2DBC), broker consumers, etc. - *Mitigation:* `doOnCancel` (client-initiated cancel) and `doFinally` (any terminal signal) to unsubscribe/close/release. **5. Event-loop starvation.** WebFlux serves all connections on a small pool of Netty event-loop threads. One blocking call (JDBC, `Thread.sleep`, synchronous HTTP) on that thread stalls every connection it hosts. - *Mitigation:* keep the pipeline non-blocking end-to-end; wrap unavoidable blocking work in `Mono.fromCallable(...).subscribeOn(Schedulers.boundedElastic())`. **6. Delivery guarantees & resumption.** SSE is inherently at-most-once during a disconnect window. Browsers auto-reconnect and send `Last-Event-ID`; you must implement server-side replay-from-id to get at-least-once. If you need strict ordering/exactly-once, SSE is the wrong tool — use a broker with acknowledgements. **7. Backpressure & memory bounding.** As covered elsewhere, hot sources need `onBackpressureLatest`/`onBackpressureDrop`/`sample`/`limitRate` so a slow client can't grow memory. **8. Horizontal scaling / fan-out.** Behind a load balancer with N instances, an event produced on instance A won't reach a client connected to instance B unless events are shared. - *Mitigation:* publish domain events to a **shared bus** (Redis Pub/Sub, Kafka, RabbitMQ) that every instance subscribes to and re-emits to its local SSE subscribers; or use sticky sessions (weaker). This decouples event production from the connection-holding instance. **9. Security & auth.** Browser `EventSource` can't set headers, so bearer-token auth is awkward — rely on cookies (with CSRF care) or a token query param (careful: it lands in logs). Enforce per-connection authorization and rate limits. **10. Observability.** Track concurrent open streams, per-stream lifetime, cancel rate, and heartbeat gaps as metrics (Micrometer). A rising open-connection count with low throughput often signals leaked/abandoned streams. **When to pick something else.** For bidirectional, high-frequency, or guaranteed-delivery needs, prefer WebSocket, RSocket (native reactive backpressure over the wire), or a message broker with a thin push layer.

  • You scale to three instances behind a load balancer and users report missing events. What's happening and how do you fix it?
    Each SSE connection is pinned to one instance; an event produced on another instance never reaches that client. Fix by publishing events to a shared bus (Redis Pub/Sub, Kafka) that every instance subscribes to and re-emits to its local SSE subscribers, decoupling event production from the connection-holding node. Sticky sessions are a weaker alternative.
  • A reverse proxy is buffering your SSE responses so clients get bursts instead of a live stream. What do you change?
    Disable proxy response buffering — for Nginx set the X-Accel-Buffering: no response header (or proxy_buffering off), ensure HTTP/1.1 chunked or HTTP/2 transport, and turn off gzip compression on the SSE path so bytes are forwarded as they're flushed.
  • Why is a blocking JDBC call inside an SSE stream especially dangerous in WebFlux?
    WebFlux serves many connections on a few Netty event-loop threads. A blocking call parks that thread, stalling every connection it hosts — not just the current one. Offload blocking work to Schedulers.boundedElastic via subscribeOn, or use reactive R2DBC instead.

saying these in an interview costs you the question

  • Assuming SSE gives exactly-once/guaranteed delivery
  • Ignoring proxy/LB idle timeouts (no heartbeat)
  • No disconnect cleanup, leaking subscriptions/cursors
  • Blocking the event loop with JDBC/sleep
  • Expecting cross-instance fan-out without a shared broker
  • Enabling gzip/proxy buffering on the stream and wondering why it's not live

context