skip to content

In a chat service whose room members are connected to many different gateway nodes, how do you broadcast one room message to all of them?

level: middleimportance: should knowfreq 44%

answer

  1. publish once, fan out twice
  2. which nodes actually hold a member
  3. subscribe on first, unsubscribe on last
  4. one channel, one bus node
  5. a fire-and-forget bus forgets

basics

~20 s

Publish the message once to the room's channel on a publish/subscribe bus. Each gateway node subscribes to that channel only while it holds a member of the room, then writes the message to its local members' sockets.

solid answer

~50 s

The chat service stores the message and publishes it **once** to a bus channel named for the room, such as `room:842`. Broadcasting every room message to every gateway, with each node filtering for its local members, is simple, but every node pays for every room's traffic. **Interest-based routing** is the usual choice: a node subscribes to a room's channel when its first local member joins and unsubscribes, after a short linger, when its last one leaves, so a message reaches only the nodes holding members, and each of those fans it out to its own sockets. A very hot room can overload the one bus node carrying its channel, so you **shard that channel** across several bus nodes. And because a plain publish/subscribe bus is usually at-most-once, a node that blips or subscribes late misses messages; a per-room sequence number lets clients spot the gap and fetch it from storage.

code

pseudocode · 18 lines
pseudocode
on member_joined(room, conn):
    local[room].add(conn)
    cancel(pending_unsubscribe[room])        # a member came back within the linger
    if room not in subscribed:
        subscribed.add(room)
        bus.subscribe("room:" + room, deliver)

on member_left(room, conn):
    local[room].remove(conn)
    if local[room] is empty:
        pending_unsubscribe[room] = after(30 seconds):
            if local[room] is empty:
                subscribed.remove(room)
                bus.unsubscribe("room:" + room)

deliver(room, msg):                           # msg carries its per-room sequence number
    for conn in local[room]:
        conn.send(msg)

go deeper

for a junior

Recall that a room's members sit on different gateway nodes, so the message is published once to a shared bus and each node delivers it to its own connected members.

for a middle

Explain broadcast versus interest-based routing, when a node subscribes and unsubscribes, why the linger exists, and the per-message node count each approach costs.

for a senior

Show you can run it: the join-time subscribe race, losses when a node's bus connection blips, slow-consumer drops, and sharding a hot room's channel across bus nodes.

for a principal

Decide where the line sits between broadcast and interest routing for your room-size distribution, and why a lossy bus cannot be the source of truth.

## One message, members on many nodes A **chat room** has members connected through long-lived sockets to a fleet of **gateway nodes**. Each member's socket lives on exactly one node, and the members of one room are usually scattered: a room of 8 people on a 200-node fleet can easily touch 8 different nodes. When one member posts, the message has to reach every other member's socket, which means it has to reach **every node that holds at least one member** of that room. Looking up each member one by one in a user-to-node registry works for a direct message, but for a room it turns one post into one lookup and one send per member. The usual design instead puts a **publish/subscribe bus** between the backend and the gateways: a broker that takes one published message on a named **channel** and hands a copy to every current subscriber of that channel. The chat service stores the message, gives it a per-room sequence number, and publishes it **once**; the bus does the cross-node fan-out, and each gateway does the last step to its own sockets. ## Two ways to route a room message | Approach | Which nodes receive a room message | Cost per message | |---|---|---| | **Broadcast to all nodes** | every gateway, on one fleet-wide channel; each node checks its local membership table and drops what it cannot use | one copy per node in the fleet | | **Interest-based routing** | only nodes subscribed to that room's channel, which a node does while it holds a member of the room | one copy per node that holds a member | - **Broadcast** is simple: no subscription bookkeeping and no late-subscription race. But every node pays for every room's traffic, so the total work is the fleet-wide message rate times the node count, and it grows each time you add a node. - **Interest-based routing** makes the work proportional to where members actually are. For the 8-member room on 200 nodes, each message reaches at most 8 nodes instead of 200, at least 25 times less. The price is subscription churn and a race at join time, covered below. - The two meet at the top: a room whose members sit on every node gets the same 200 copies either way. Interest routing pays off for the many small rooms, not for the giant one. ## Interest-based routing in practice Each gateway keeps a local table of `room -> connections on this node` and drives its bus subscriptions from it: 1. **First local member joins.** The node adds the connection and, if it is not already subscribed, subscribes to the room's channel, for example `room:842`. 2. **A message is published.** The bus delivers one copy to each subscribed node. 3. **Local fan-out.** The node writes the message to every local connection in that room. 4. **Last local member leaves.** The node schedules an unsubscribe after a short **linger**, such as 30 seconds, and cancels it if a member comes back. Without the linger, a member who reconnects to the same node every few seconds makes the node subscribe and unsubscribe in a loop. One race needs naming. Subscribing takes time to become active on the bus, so a message published between "the member joined" and "the subscription is live" never reaches that node. The client has to be able to notice that, which is the last section's point. ## Hot rooms: when one channel is too much On many buses a channel lives on one bus node, and that node does the work of `messages per second x subscribing gateway nodes`. A live-event room with 50 messages per second and members on all 200 gateways asks one bus node for 10,000 deliveries per second, on top of every other channel it carries. Two ways to **shard a hot room's channel** spread that work: - **Split by message.** Publish each message to one of several sub-channels, such as `room:842/0` to `room:842/3`, placed on different bus nodes; every interested gateway subscribes to all four. Each bus node carries a quarter of the room's messages. Order across sub-channels is no longer guaranteed, so clients order by the per-room sequence number. - **Split by subscriber.** Publish each message to several replica channels and let each gateway subscribe to exactly one, chosen by hashing its node ID. Each bus node serves a quarter of the gateways, at the price of four publishes per message. Either way the gateways still receive every message of the room; sharding moves load **off one bus node**, it does not reduce the fan-out itself. ## What the bus can lose A plain publish/subscribe bus is usually **at-most-once**: it delivers to whoever is subscribed at that moment and keeps nothing for anyone who is not. - A gateway's bus connection blips for two seconds: every room message published in that window is gone for that node's members. - A node subscribes late, as in the join race above. - A gateway reads too slowly, and the bus may drop messages or disconnect it to protect itself. None of these is visible to the publisher. This is why each message carries its **per-room sequence number**: a client that sees 41 followed by 44 knows it missed two and fetches that range from the message store, so the bus can stay a fast, lossy pipe rather than the source of truth.

  • When is broadcasting every room message to every gateway node actually the better choice?
    When the fleet is small or most rooms are large. With a handful of nodes, the wasted copies cost less than subscription bookkeeping and the join-time race. If most traffic belongs to rooms whose members already sit on nearly every node, interest-based routing delivers to almost every node anyway, so it saves little. It pays off with many small rooms spread over a large fleet.
  • A user joins a room and their first message from it arrives with sequence 3, not 1. What went wrong, and is it a bug?
    Probably the join race: the gateway subscribed to the room's channel, but messages 1 and 2 were published before the subscription was live on the bus, so the node never received them. It is expected behaviour of a lossy bus, not a bug in itself. The client sees the gap in the per-room sequence and fetches the missing range from the message store.

saying these in an interview costs you the question

  • Once a node subscribes to a room's channel, it cannot miss that room's messages.
  • A gateway node should subscribe to every room's channel at startup.
  • A plain publish/subscribe bus keeps messages for a node until it resubscribes.
  • Sharding a hot room's channel reduces how many gateway nodes receive each message.
  • Interest-based routing saves just as much for a room whose members are on every node.