skip to content

Why does a WebSocket or Server-Sent Events gateway tier not scale horizontally as simply as a stateless HTTP API tier?

level: juniorimportance: must knowfreq 62%

answer

  1. what the server must remember
  2. placed once, at connect time
  3. capacity unit for idle sockets
  4. which node holds this user?
  5. a restart is a mass event

basics

~20 s

Each long-lived connection is pinned to one server for its whole life, so that server holds state: capacity is counted in open connections, pushes must be routed to the right node, and restarts disconnect everyone on it.

solid answer

~40 s

A stateless API server handles a request in milliseconds and forgets it, so the load balancer can send the next request anywhere and servers can be added or removed freely. A WebSocket or SSE connection lives for minutes to days on **one** gateway node, and that node holds the socket, its buffers and the session. That changes three things: capacity is measured in **concurrent connections** and memory rather than requests per second; a backend that wants to push to a user must find **which node** holds that user's socket; and restarting or losing a node drops every connection on it, and those clients all reconnect at once. Adding nodes also does not rebalance existing load, because a connection is placed only when it opens.

go deeper

for a junior

Recall the root cause: a long-lived connection stays on one server for its whole life, so that server holds state a stateless API server never keeps.

for a middle

Explain the consequences one by one: connection-based capacity, finding the node that holds a user, and why new nodes receive only newly opened connections.

for a senior

Show you have operated such a tier: keep gateways thin, plan headroom for failed nodes, and treat every gateway restart as a reconnect event to be managed.

for a principal

Frame the tier as the one stateful layer in an otherwise stateless system, and argue how much logic it may own given that every change to it disconnects users.

## The stateless baseline A **stateless** API server keeps nothing about a client between requests. A request arrives, the server reads what it needs from shared storage, answers within milliseconds and forgets the exchange. Because no server remembers anything, a load balancer can send each request to any healthy server, a crashed server costs only its in-flight requests (which clients retry), and a new server adds capacity as soon as it passes health checks. Capacity is planned in **requests per second**. ## What a long-lived connection changes A **long-lived connection** is a single TCP connection kept open so the server can push data whenever it has some. A **WebSocket** is a full-duplex connection upgraded from an HTTP request; **Server-Sent Events (SSE)** is a one-way event stream carried in an HTTP response that does not finish. A chat, notification or live-feed product holds one such connection per open app or tab, and each may live for minutes, hours or days. The servers that hold these connections form the **connection tier**, often called the **gateway fleet**. Each connection belongs to exactly one gateway node for its whole life, and that node holds: - the socket itself, which is an operating-system **file descriptor**; - kernel send and receive buffers and, when it terminates TLS, the encryption state; - an application session: which user and device this is, what it subscribed to, and a small outbound queue. None of that can be handed to another node while the connection stays open, so the gateway is **stateful** whether the design admits it or not. ## Where the two tiers differ | Concern | Stateless API tier | Long-lived connection tier | |---|---|---| | Capacity unit | requests per second | concurrent open connections, plus the memory and descriptors they hold | | When placement happens | on every request | once, when the connection opens | | Who starts traffic | the client | often the server, pushing to one specific user | | Addressing | any server can answer | a push must reach the one node holding that socket | | Scale-out | new servers take load at once | new servers receive only newly opened connections | | Deploy or crash | in-flight requests fail and are retried | every connection on the node drops and reconnects | Three rows carry most of the interview answer: 1. **Capacity is counted in connections.** Most connections are idle most of the time, so requests per second badly understates the load. A node is full when its memory or descriptor budget is full, even while its CPU is quiet. 2. **Push needs addressing.** When a backend service wants to deliver a message to a user, it has to find which gateway node holds that user's connection. That usually means a shared **user-to-gateway registry** that gateways update as connections open and close. 3. **Restarts become events.** Restarting a node disconnects everyone on it at once, and those clients try to reconnect together. A careless deploy or a crash becomes a **reconnect storm** that can overload the rest of the fleet. The scale-out row matters as well. A load balancer places a connection only when it opens, and it cannot move an open TCP connection to a newly added node, so new nodes fill only as old connections close and their clients come back. ## Design consequences These differences shape the tier in predictable ways: - **Keep gateways thin.** A gateway terminates the connection, authenticates it, registers it and relays messages. Business logic lives in stateless services behind it, so everyday deploys do not disconnect anyone. - **Size for peak concurrent connections**, with headroom so the users of a failed node fit on the others. - **Maintain a way to find a user's node**, usually a registry, sometimes deterministic placement by hashing the user ID. - **Drain deliberately** on deploy, and have clients reconnect with randomized backoff. - **Detect dead connections**, because a client that vanishes without closing still holds a descriptor and memory until the server notices. ## Why interviewers ask it The question checks whether a candidate notices that "just add servers" stops being enough once the server must remember the client. Weak answers treat the gateway like any web tier and assume the load balancer keeps spreading load continuously. Strong answers name the pinned connection as the root cause and derive the consequences from it: connection-based capacity, addressing, slow rebalancing and disruptive restarts. SSE deserves a mention too: it looks like ordinary HTTP, but an open event stream pins one server exactly as a WebSocket does, so the same tier problems apply.

  • Why should business logic stay out of the gateway nodes?
    Restarting a gateway disconnects every client on it and triggers reconnects, so gateway deploys are expensive. Keeping the gateway thin, limited to terminating connections, authenticating, registering and relaying, lets business logic ship in stateless services behind it without disconnecting anyone. Gateway code then changes rarely, and the painful drain procedure runs rarely too.
  • Does Server-Sent Events avoid these problems because it runs over plain HTTP?
    No. An SSE stream is an HTTP response that stays open, so it pins one server for its lifetime exactly like a WebSocket. It consumes a descriptor and memory while idle, needs addressing for server pushes, and drops on restart. SSE changes the direction and format of the traffic, not the fact that the tier is stateful.

A stateless API is a call centre where any free agent takes the next call; a connection tier is a switchboard where each caller stays on the line with one operator, so a message for that caller must reach that operator.

saying these in an interview costs you the question

  • A WebSocket tier scales like a REST tier: just add servers behind the balancer.
  • The load balancer spreads existing connections onto newly added nodes automatically.
  • Sizing the gateway fleet by requests per second is enough.
  • Any gateway node can push a message straight to any connected user.
  • Restarting a gateway node is harmless because clients reconnect anyway.