skip to content

How would you architect and secure a horizontally-scaled STOMP-over-WebSocket system, and what failure modes matter?

level: principalimportance: nice to knowfreq 30%

answer

  1. clustered external relay = shared subscriptions
  2. two-layer security: handshake auth + per-frame authz
  3. LB: WebSocket-aware, sticky for SockJS, wss + Origin
  4. limits: message size, send buffer, send time
  5. failure: relay drop, backpressure, reconnect storm

basics

~20 s

Use an external STOMP broker relay so all app instances share subscriptions, authenticate and authorize at both the handshake and per-message level, plan for sticky/routable connections at the load balancer, and design for broker and relay failures with reconnection and backpressure handling.

solid answer

~40 s

At scale you replace the in-memory broker with enableStompBrokerRelay to a clustered RabbitMQ/ActiveMQ so any instance can deliver to any subscriber. Load balancers must handle long-lived upgraded connections (WebSocket-aware, idle timeouts raised) and often need sticky routing when using SockJS fallback. Security is two-layered: the HTTP handshake goes through the normal Spring Security filter chain (authenticate, set the Principal on the STOMP session), and individual STOMP frames are authorized with a MessageSecurityMetadataSource / AuthorizationManager (message-level @PreAuthorize-style rules on destinations) plus CSRF/same-origin checks on the handshake. Failure modes to design for: relay disconnect halts broadcasts (need reconnection + client resubscribe), broker backpressure and slow consumers, per-client broker TCP connection exhaustion, thundering-herd resubscribes after a deploy, and message ordering guarantees. Also cap message sizes and subscription counts to resist abuse.

code

java · 19 lines
java
@Configuration
@EnableWebSocketMessageBroker
class WsConfig implements WebSocketMessageBrokerConfigurer {
    @Override public void configureMessageBroker(MessageBrokerRegistry r) {
        r.setApplicationDestinationPrefixes("/app");
        r.enableStompBrokerRelay("/topic", "/queue")
         .setRelayHost("rabbit-cluster").setRelayPort(61613)
         .setSystemLogin("sys").setSystemPasscode("***")
         .setClientLogin("app").setClientPasscode("***");
    }
    @Override public void configureWebSocketTransport(WebSocketTransportRegistration reg) {
        reg.setMessageSizeLimit(64 * 1024)   // cap inbound frame size
           .setSendBufferSizeLimit(512 * 1024)
           .setSendTimeLimit(20_000);        // drop slow consumers
    }
    @Override public void registerStompEndpoints(StompEndpointRegistry reg) {
        reg.addEndpoint("/ws").setAllowedOrigins("https://app.example.com").withSockJS();
    }
}

go deeper

for a junior

Recognize that scaling needs an external broker and that WebSocket connections are long-lived.

for a middle

Explain the relay-for-scale requirement and basic handshake authentication.

for a senior

Detail two-layer security, transport limits, and common failure modes (relay drop, backpressure).

for a principal

Own the full architecture: clustered broker sizing, edge/LB behavior, per-frame authorization, abuse limits, reconnect-storm and ordering semantics, and the SSE-vs-STOMP tradeoff.

### Topology for scale **Shared broker is mandatory.** With more than one app instance you must use `enableStompBrokerRelay` pointing at a **clustered** external broker (RabbitMQ with the STOMP plugin, or ActiveMQ Artemis). Each app instance becomes a stateless relay; the broker holds the authoritative subscription registry and does fan-out, so a client on instance A receives a message published from instance C. The in-memory simple broker is disqualified because its subscription state is per-JVM. **Connection accounting.** The relay opens one 'system' TCP connection per app instance plus (depending on setup) connections associated with client sessions; a real broker also holds each browser's logical subscription. At tens of thousands of concurrent clients this drives broker sizing, file-descriptor limits, and memory. ### Load balancer / edge concerns - WebSocket connections are long-lived HTTP upgrades; the LB/proxy must support the `Upgrade` header and have generous idle timeouts, or connections get killed mid-session. - With **SockJS** fallback transports (XHR-streaming/long-polling), a logical session spans multiple HTTP requests, so you typically need **sticky sessions** (session affinity) to keep them on one instance. Native WebSocket is a single connection and less sensitive, but affinity still simplifies things. - Terminate TLS (`wss://`) at the edge; enforce allowed `Origin` to prevent cross-site WebSocket hijacking. ### Security — two layers 1. **Handshake authentication.** The initial HTTP(S) request that upgrades to WebSocket passes through the normal Spring Security filter chain. Authenticate it (session cookie, JWT, etc.) and ensure a `Principal` is bound to the WebSocket/STOMP session — this is what `convertAndSendToUser` and `@SendToUser` rely on. Configure `setHandshakeHandler`/interceptors if you carry auth in the STOMP `CONNECT` frame instead. 2. **Per-message authorization.** STOMP frames after connect are NOT re-checked by the HTTP filter chain. Use Spring Security's messaging support (`AbstractSecurityWebSocketMessageBrokerConfigurer` / the newer `AuthorizationManager`-based message security) to authorize by frame type and destination — e.g. only admins may SUBSCRIBE to `/topic/admin/**`, deny SEND to broker destinations directly, require authentication for `/app/**`. 3. **CSRF / same-origin.** The handshake is CSRF-protected by default in Spring Security; validate `Origin`. For token auth, prefer sending the token in the STOMP `CONNECT` headers over cookies to sidestep CSRF concerns. 4. **Abuse limits.** Cap inbound message size (`setMessageSizeLimit`), send buffer (`setSendBufferSizeLimit`), send time limit (`setSendTimeLimit`) via `configureWebSocketTransport`, and bound subscriptions per session to resist resource-exhaustion attacks. ### Failure modes to design for - **Relay connection loss:** if the app-to-broker TCP link drops, broadcasts stop until it reconnects; clients may need to resubscribe. Monitor the relay connection and surface health. - **Broker outage / failover:** clustered broker with failover; clients should auto-reconnect (SockJS/StompJS reconnect) and re-establish subscriptions idempotently. - **Slow consumers / backpressure:** a client that can't keep up backs up the outbound buffer; enforce send buffer/time limits so one bad client can't OOM the instance. The broker's flow control governs cross-instance backpressure. - **Thundering herd on deploy:** rolling restarts drop thousands of connections that reconnect at once; jitter client reconnect and scale the broker for the spike. - **Ordering & delivery semantics:** STOMP over a single connection preserves per-destination order for one session, but across reconnects or the relay you should assume at-most/at-least-once depending on broker config, not exactly-once; design idempotent clients. - **User-destination resolution across instances:** `/user/**` messages must reach the instance holding that user's session; the relay handles this via the broker, but with the simple broker it silently fails cross-instance. ### When NOT to go full STOMP If you only need server->client push (no client publishing, no complex routing), Server-Sent Events over HTTP/2 is simpler to scale and secure. Choose STOMP+relay when you genuinely need bidirectional pub/sub with routing and broker integration.

  • The HTTP filter chain authenticates the handshake. Why isn't that enough to secure the system?
    Individual STOMP frames after CONNECT bypass the HTTP filter chain, so you need message-level authorization (per-destination/per-frame-type) to control who can SUBSCRIBE or SEND where.
  • Why might you need sticky sessions at the load balancer?
    SockJS fallback transports span multiple HTTP requests for one logical session, so affinity keeps them on the same instance; it also simplifies session/principal handling.
  • How do you keep one slow client from taking down an instance?
    Set send buffer and send-time limits (configureWebSocketTransport) so the outbound buffer is bounded and slow consumers get disconnected rather than exhausting memory.
  • How do /user/** messages reach the right instance in a scaled deployment?
    Through the external broker relay, which shares user-destination routing across instances; the in-memory broker can't do this and silently drops cross-instance user messages.

saying these in an interview costs you the question

  • Assuming handshake authentication also authorizes every STOMP frame
  • Running the in-memory broker in a multi-instance deployment
  • Ignoring load-balancer idle timeouts and Origin/CSRF on the handshake
  • No message-size or send-buffer limits, allowing one client to exhaust resources
  • Expecting exactly-once delivery across reconnects

context