convertAndSendToUser works on a single instance but silently fails in a multi-node cluster. Why, and how do you fix it?
answer
- SimpUserRegistry is in-memory, per-JVM
- local resolve fails → silent no-op on other nodes
- simple broker = single node; use enableStompBrokerRelay
- setUserRegistryBroadcast (knowledge) + setUserDestinationBroadcast (delivery)
- need both; sticky sessions still required; broker = critical infra
basics
~20 sconvertAndSendToUser only knows sessions on the local node's SimpUserRegistry. If the target user is connected to a different node, the message goes nowhere. Fix it by using an external STOMP broker relay and enabling cross-node user registry and user-destination broadcast.
solid answer
~40 sUser-destination routing resolves the target's sessions from the in-memory SimpUserRegistry, which only holds sessions connected to the current JVM. In a cluster, if you call convertAndSendToUser on node A but the user is on node B, no local session matches and the message is dropped. The fix requires an external message broker: replace enableSimpleBroker with a STOMP broker relay (enableStompBrokerRelay) pointing at RabbitMQ/ActiveMQ, then enable two broadcasts — setUserRegistryBroadcast("/topic/simp-user-registry") so nodes share user/session info, and setUserDestinationBroadcast("/queue/unresolved-user-destination") so a message for a user not found locally is broadcast to other nodes, one of which resolves it. The simple broker cannot do this; it's inherently single-node. This is the standard scale-out story for WebSocket user messaging.
code
java · 18 lines@Configuration
@EnableWebSocketMessageBroker
public class ClusterWsConfig implements WebSocketMessageBrokerConfigurer {
@Override
public void configureMessageBroker(MessageBrokerRegistry registry) {
registry.enableStompBrokerRelay("/topic", "/queue")
.setRelayHost("rabbitmq").setRelayPort(61613)
.setSystemLogin("relay").setSystemPasscode("secret")
.setClientLogin("app").setClientPasscode("secret")
// cluster-wide knowledge of who is connected where:
.setUserRegistryBroadcast("/topic/simp-user-registry")
// delivery of user messages unresolved on the local node:
.setUserDestinationBroadcast("/queue/unresolved-user-destination");
registry.setApplicationDestinationPrefixes("/app");
}
// convertAndSendToUser(...) now reaches the user on whichever node holds their session
}go deeper
Know user sends may not reach users on other servers in a cluster.
Explain that the user registry is per-JVM and the simple broker is single-node.
Configure a STOMP broker relay with user registry and user-destination broadcast to fix it.
Reason about broker as critical infra, sticky sessions, broadcast overhead at scale, idempotent delivery, and partitioning strategies.
**Root cause — the SimpUserRegistry is local:** When a user connects, Spring records their sessions in a `SimpUserRegistry`. The default implementation is **in-memory and per-JVM** — it only knows about WebSocket sessions terminated on *this* instance. `convertAndSendToUser(user, dest, payload)` asks the registry "which sessions does `user` have?", rewrites the destination per session, and sends. On the node where the user *isn't* connected, the registry returns nothing, so the send is a **silent no-op** — no error, no delivery. Fine on one node; broken the moment you scale horizontally behind a load balancer. **Why the simple broker can't help:** `registry.enableSimpleBroker("/topic","/queue")` is an in-memory broker living inside each JVM. Two app instances have two independent brokers that don't share subscriptions or messages. So even `/topic` broadcasts don't cross nodes with the simple broker, let alone user destinations. **Fix — external broker relay + two broadcasts:** ```java @Override public void configureMessageBroker(MessageBrokerRegistry registry) { registry.enableStompBrokerRelay("/topic", "/queue") .setRelayHost("rabbitmq-host").setRelayPort(61613) .setUserRegistryBroadcast("/topic/simp-user-registry") .setUserDestinationBroadcast("/queue/unresolved-user-destination"); registry.setApplicationDestinationPrefixes("/app"); } ``` 1. **`enableStompBrokerRelay(...)`** — instead of an in-JVM broker, each app node connects (via STOMP) to a shared external broker (RabbitMQ with the STOMP plugin, ActiveMQ, etc.). All nodes now share `/topic` and `/queue` state through that broker. This alone fixes normal broadcasts across the cluster. 2. **`setUserRegistryBroadcast(dest)`** — nodes periodically **broadcast their user registry** over the given broker topic, and each node merges the others' data into a `MultiServerUserRegistry`. Now node A *knows* user X is connected (somewhere), even if not locally. Used for things like presence/`SimpUserRegistry.getUser(...)` queries cluster-wide. 3. **`setUserDestinationBroadcast(dest)`** — when a node receives a message for a user destination it **can't resolve locally**, it forwards (broadcasts) that unresolved message over this broker queue. Every node listens; the node where the user *is* connected resolves and delivers it. This is what actually makes `convertAndSendToUser` reach a user on another node. You need **both** broadcasts: the registry broadcast for cluster-wide *knowledge* of users, and the destination broadcast for *delivery* of unresolved user messages. **Operational considerations (principal-level):** - **Broker becomes critical infra:** you now depend on RabbitMQ/ActiveMQ availability, credentials, TLS, and the broker's own clustering/HA. The relay maintains a system TCP connection (`setSystemLogin`/`setSystemPasscode`, heartbeats). - **Sticky sessions still apply** at the HTTP/SockJS layer: a single logical WebSocket/SockJS session's transport requests must hit the same node; the broadcasts solve *cross-user* routing, not moving a live session between nodes. - **Scaling the fan-out:** registry broadcast traffic grows with node count (each node advertises its users to all); for very large clusters this chatter and the unresolved-destination broadcast add overhead. Alternatives include partitioning users by node (consistent hashing at the LB) so cross-node sends are rare, or using the broker's own routing. - **Ordering/at-most-once:** unresolved-destination broadcast is best-effort; if a user reconnects to a different node mid-flight, a message may miss. Design idempotent/replayable delivery for anything critical (persist notifications, let the client fetch on connect). - **Security:** the broker relay must be authenticated and network-isolated; destinations used for broadcasts (`/topic/simp-user-registry`, `/queue/unresolved-user-destination`) are internal — ensure your message-security rules and broker ACLs don't expose them to clients. **Diagnosis tip:** the symptom is 'works locally, nothing in prod behind a LB, no exception.' Check whether you're on `enableSimpleBroker` (single-node) and whether the two broadcast destinations are configured on the relay. `SimpUserRegistry` being a plain in-memory instance rather than a `MultiServerUserRegistry` confirms the gap.
- Why do you need BOTH setUserRegistryBroadcast and setUserDestinationBroadcast?setUserRegistryBroadcast shares user/session knowledge across nodes (so a MultiServerUserRegistry knows a user exists cluster-wide — useful for presence/queries). setUserDestinationBroadcast forwards messages that a node can't resolve locally to other nodes for delivery. Knowledge without delivery, or delivery without knowledge, leaves gaps; you want both.
- Does enabling the broker relay remove the need for sticky sessions?No. The broadcasts fix cross-user routing between nodes, but a single WebSocket/SockJS session (especially SockJS HTTP transports) still binds to one node's in-memory session state, so the load balancer must keep that session's requests on the same node via session affinity.
- How would you reduce the cross-node broadcast overhead at very large scale?Partition users to nodes (e.g. consistent hashing / affinity at the load balancer) so a user's producers and their session tend to share a node, making cross-node sends rare; or rely on the external broker's own routing/queues per user. Also make delivery idempotent/replayable so occasional misses during reconnects are harmless.
saying these in an interview costs you the question
- Believing convertAndSendToUser magically finds users across nodes with the simple broker
- Enabling only the broker relay but forgetting the two broadcast destinations
- Thinking a shared broker removes the need for sticky sessions
- Assuming the silent no-op is a code bug rather than a topology/broker limitation