skip to content

In a messaging app's gateway fleet, how does a backend service deliver a message to the specific node holding the recipient's connection?

level: seniorimportance: must knowfreq 52%

answer

  1. stateless sender, stateful receiver
  2. a channel addressed to one node
  3. one entry per device
  4. guarding the delete on disconnect
  5. liveness per gateway, not per entry

basics

~10 s

Gateways record user-to-gateway entries in a shared low-latency registry when connections open; the delivery service looks up the recipient and sends to that gateway's own addressed channel, and the gateway writes to the socket.

solid answer

~40 s

On connect, the gateway writes an entry such as `user -> {gatewayId, connectionId}` into a shared in-memory key-value store, one entry per device, and removes it on disconnect with a **conditional delete** so a late close cannot erase a newer connection's entry. The delivery service reads the recipient's entries and sends the message to each gateway's **own addressed channel**, such as a per-node subject on a message broker or a direct RPC, and the gateway writes it to the socket. Crashed gateways leave stale entries, so tie entries to a **gateway-level lease** and epoch: when a node stops renewing, all its entries count as dead, and delivery falls back to the offline path. Broadcasting every message to every gateway avoids the registry but multiplies the work by the node count.

code

json · 7 lines
json
{
  "user:48213": {
    "conn-7f3a": { "gateway": "gw-12", "epoch": 31, "device": "phone" },
    "conn-91cc": { "gateway": "gw-04", "epoch": 18, "device": "laptop" }
  },
  "gateway:gw-12": { "epoch": 31, "leaseExpiresAt": "2026-09-17T10:15:30Z" }
}

go deeper

for a junior

Recall that a message must reach the one server holding the recipient's connection, and that a shared lookup table usually records where that is.

for a middle

Explain the connect, deliver and disconnect paths, why each device gets its own entry, and why the delete must check the connection ID.

for a senior

Demonstrate failure handling: stale entries after crashes, gateway leases and epochs, reconnect races, and falling back to the offline path.

for a principal

Weigh the registry against broadcast and deterministic placement, and treat the registry as critical soft state with its own availability budget.

## The addressing problem In a messaging app, a message for a user is produced by a backend service, for example the chat service right after it stores the message. That service is stateless and knows nothing about sockets. The recipient's connection, however, lives on exactly one **gateway node** in a fleet of dozens. Something must answer *"which node holds this user's connection right now?"*, and the answer changes every time the user reconnects. The usual answer is a **user-to-gateway registry**: a shared, low-latency store, typically an in-memory key-value cache replicated and sharded by user ID, that gateways write to and delivery services read from. ## The registry record A user can be connected from several devices at once, so the record is a small set of entries rather than one value. Each entry names the gateway and the **gateway epoch**, a number the gateway increments every time it starts, and each gateway keeps a lease record of its own: ```json { "user:48213": { "conn-7f3a": { "gateway": "gw-12", "epoch": 31, "device": "phone" }, "conn-91cc": { "gateway": "gw-04", "epoch": 18, "device": "laptop" } }, "gateway:gw-12": { "epoch": 31, "leaseExpiresAt": "2026-09-17T10:15:30Z" } } ``` ## Write, read and delete paths 1. **Connect.** After authenticating the connection, the gateway adds an entry under the user's key with its gateway ID, current epoch and a connection ID. 2. **Deliver.** The delivery service reads the user's entries and, for each one, sends the message to that gateway's **own addressed channel**: a per-node subject on a message broker, or a direct RPC to the node. The gateway finds the connection ID in its local table and writes the message to the socket. 3. **Disconnect.** The gateway removes its entry with a **conditional delete** that succeeds only if the stored connection ID still matches. Without the condition, a slow close on node A could erase the entry a fresh connection on node B has just written for the same device, leaving a connected user unreachable. 4. **Miss.** If the user has no live entries, or every send fails, the message goes to the offline path, such as a mobile push notification. ## Keeping the registry honest Gateways crash without cleaning up, so entries go stale. The naive fix is a TTL on every entry that its gateway keeps renewing, and the arithmetic is unkind: renewing 10 million entries once every 60 seconds is 10,000,000 / 60, about **167,000 writes per second** of pure bookkeeping. A cheaper design ties liveness to the gateway instead of the entry: - each gateway renews one **lease record** every few seconds, a handful of writes per second across the whole fleet; - a reader treats an entry as dead when its gateway's lease has expired or the gateway's current epoch differs from the entry's; - a background sweeper deletes dead entries lazily, so readers never wait for cleanup; - a restarted gateway starts with a higher epoch, so entries left by its previous run stop counting as soon as readers see the new epoch. Other failure modes to plan for: - **Reconnect races.** A device can briefly have entries on two nodes. Delivering through both is harmless when the client de-duplicates by message ID. - **A node that just died.** The send times out or the node's channel has no subscriber. The delivery service re-reads the registry once, since the user may already be on another node, then falls back to the offline path. - **Registry outage.** The registry is **soft state**: gateways can re-register every live connection after a failover, so availability matters more than durability. ## Alternatives compared | Approach | How a message finds the node | Main cost | |---|---|---| | Registry lookup | read the user's entries, send to each node's channel | an extra lookup and a critical shared dependency | | Broadcast to all gateways | publish on a fleet-wide channel; every node checks its local table | every message costs work on every node, so load grows with fleet size | | Deterministic placement | compute the node by hashing the user ID | routing must read the user ID, and changing the node set moves users | Broadcast is a reasonable start for a small fleet, but with 50 nodes every message is examined 50 times, and the waste grows with each node added. Deterministic placement removes the lookup but brings its own trade-offs around scaling events. ## Routing is not presence The registry answers a routing question for machines. Whether a user appears *online* to friends is a separate, user-visible feature with its own heartbeat and fan-out rules. Keeping the two apart lets the registry stay small, fast and internal, and stops presence noise such as rapid online and offline flapping from touching the delivery path.

  • What should the delivery service do if it sends to a gateway that crashed a second ago?
    The send times out or the node's channel has no subscriber, so treat the message as not delivered. Re-read the registry once, because the device may already have reconnected to another node, and deliver there if a live entry appears. Otherwise hand the message to the offline path. Never block on the dead node; its expired lease makes its entries count as dead.
  • How do you keep the registry available when every delivery depends on it?
    Replicate it and shard it by user ID so no single failure removes all routing. Treat it as soft state: after a failover, gateways re-register every live connection from their local tables, so durability matters less than availability. Delivery services can cache lookups for a few seconds, accepting an occasional misroute that triggers a fresh read.

saying these in an interview costs you the question

  • Any gateway node can write to any user's socket.
  • One registry entry per user, overwritten by whichever device connected last, is enough.
  • Renewing a TTL on every registry entry adds negligible write load.
  • Showing a user as online and routing to their gateway are the same job.
  • Broadcasting each message to every gateway is simpler and scales fine.