skip to content

How do you estimate how many WebSocket gateway nodes a messaging app needs to hold 10 million concurrent connections?

level: middleimportance: must knowfreq 55%

answer

  1. peak concurrent, not daily users
  2. bytes held by an idle socket
  3. every socket costs a descriptor
  4. the four-tuple and ephemeral ports
  5. headroom for a lost node

basics

~20 s

Divide the memory budget by per-connection memory to find a node's ceiling, check descriptor, port and CPU limits, then plan well below it: 10 million connections at 500,000 per node is 20 nodes plus spares.

solid answer

~40 s

Start from peak **concurrent** connections, not users. Measure memory per idle connection, covering socket buffers, TLS state and the session object; assuming about 16 KB, a node with 16 GB to spare tops out near 1 million connections. Then check the other ceilings: every socket is a **file descriptor**, so raise the per-process and system-wide limits; a proxy that opens its own upstream connection per client to one gateway address is capped at roughly 64,000 ephemeral ports per source and destination pair; and CPU is consumed by TLS handshakes and fanout bursts rather than idle sockets. Plan at about half the ceiling, say 500,000 per node, so a failed or draining node's users fit elsewhere: 10,000,000 / 500,000 = 20 nodes, plus a few spares for deploys and zone loss.

code

pseudocode · 7 lines
pseudocode
peak_connections = 10,000,000
bytes_per_conn   = 16 KB                           # measured on idle sockets
memory_budget    = 16 GB                           # per node, after OS and headroom
ceiling          = memory_budget / bytes_per_conn  # about 1,048,576
target           = 500,000                         # about half the ceiling
nodes            = ceil(peak_connections / target) # 20
fleet            = nodes + 4                       # deploy batch and zone loss

go deeper

for a junior

Remember that the number to plan around is how many connections are open at the same time, and that each open connection uses memory even when idle.

for a middle

Walk through the arithmetic out loud: memory per connection, the node ceiling, the planning target and the node count, then name descriptor and port limits.

for a senior

Show operational judgment: measure under realistic traffic, alert at the planning target, reclaim half-open connections, and size CPU for reconnect waves.

for a principal

Treat headroom as a policy decision that trades hardware cost against how many node or zone failures, and how fast a deploy, the fleet must survive.

## Start from the right demand number The input to connection-tier sizing is **peak concurrent connections**, not registered or daily active users. A messaging app with 50 million daily users might see 10 million connections open in its busiest minute. Count connections rather than people: a user with a phone, a laptop and two browser tabs holds four. This worked example assumes **10 million peak concurrent connections**. ## Memory per connection Every open connection costs memory even when it is silent. The main pieces are: - **kernel socket buffers** for sending and receiving, which the operating system may enlarge under traffic; - **TLS state**, when the gateway terminates encryption; - the **application session**: user and device identity, subscriptions and a small outbound message queue. An illustrative budget, not a measured limit of any particular stack: | Component | Assumed size | |---|---| | Socket buffers (idle, tuned small) | ~8 KB | | TLS session state | ~4 KB | | Session object and outbound queue | ~4 KB | | **Total per connection** | **~16 KB** | Real figures vary widely with buffer tuning and runtime, so measure them: open a large number of idle test connections against one node and divide the memory growth by the count, then repeat at a realistic message rate, because buffers and queues grow once traffic flows. With 16 GB of a node's memory set aside for connections, the memory ceiling is 16 GB / 16 KB = 2^34 / 2^14 = 2^20, about **1,048,576 connections**. ## The other ceilings Memory is rarely the only limit. Check each of these: | Limit | What it caps | Typical fix | |---|---|---| | Per-process file descriptors | one descriptor per socket; defaults can be as low as about a thousand | raise the process limit well above the target | | System-wide file table | all descriptors on the host | raise the host limit together with the process limit | | Ephemeral ports | connections from one source IP to one destination IP and port: roughly 64,000 at most, often fewer as configured | add source IPs or destination ports, or forward at L4 | | CPU | TLS handshakes and fanout bursts, not idle sockets | size for the reconnect rate, not the idle count | | Bandwidth | bursts of messages to many sockets at once | cap the per-node fanout rate | Two points trip candidates up. First, **a server's listening port does not cap its client connections**. A TCP connection is identified by the four-tuple of source IP, source port, destination IP and destination port, so one listening port can accept connections from millions of distinct client address and port pairs. The ephemeral-port limit bites only where one machine opens many connections to the same destination, such as an intermediate proxy that opens its own upstream connection per client to a single gateway address. Second, idle connections barely use CPU, but **TLS handshakes are expensive**, so the CPU ceiling shows up during reconnect waves rather than in steady state. ## From ceiling to fleet size 1. Take the lowest ceiling: here memory, at about 1 million connections per node. 2. Choose a **planning target** well below it, about half, rounded to **500,000 per node**. The headroom absorbs a failed node's users, nodes being drained during deploys, and buffer growth under bursts. 3. Divide: 10,000,000 / 500,000 = **20 nodes**. 4. Add spares for a deploy batch and the loss of a failure zone, for example 4 more, giving **24 nodes**. The same numbers become a monitor: alert when a node passes its planning target, long before it reaches its ceiling. ```pseudocode peak_connections = 10,000,000 bytes_per_conn = 16 KB # measured on idle sockets memory_budget = 16 GB # per node, after OS and headroom ceiling = memory_budget / bytes_per_conn # about 1,048,576 target = 500,000 # about half the ceiling nodes = ceil(peak_connections / target) # 20 fleet = nodes + 4 # deploy batch and zone loss ``` ## Reclaiming dead connections A client that disappears without closing, such as a phone losing signal or a laptop going to sleep, leaves a **half-open** connection that still holds a descriptor and memory. Over hours these can inflate the connection count well past the number of real users. The tier should: - record last activity per connection and close any that miss the heartbeat interval or pass an idle timeout; - track expirations with a cheap structure such as a timing wheel instead of one timer per socket; - count only live connections on capacity dashboards. ## What interviewers look for A good answer states its assumptions, does the division, and explains why the plan sits well below the ceiling. Stronger candidates bring up descriptor limits, correct the myth of a 65,535-connection server limit, and note that CPU becomes the constraint when many clients reconnect at once.

  • How does the connection tier find and reclaim dead connections?
    Clients often vanish without closing, leaving half-open sockets that still hold a descriptor and memory. The gateway records last activity per connection and closes any that miss the heartbeat interval or exceed an idle timeout. Tracking expirations in a timing wheel rather than one timer per socket keeps the sweep cheap at a million connections, and dashboards should count only live connections.
  • Why plan at about half the memory ceiling instead of close to it?
    When a node fails, its users reconnect to the rest, so each survivor must have room for a share of them. Deploys drain nodes and push their users onto others, and buffers and outbound queues grow under message bursts beyond what idle measurements show. Running near the ceiling turns any of these into a cascading overload.
  • Why does CPU, rather than memory, often become the limit during an incident?
    Holding an idle encrypted connection costs almost no CPU, but establishing one requires a TLS handshake with expensive key-exchange work. When many clients reconnect together, handshake cost dominates and a node can saturate its CPU while memory is still available. Size CPU for the peak reconnect rate you intend to admit.

saying these in an interview costs you the question

  • Size the gateway fleet directly from daily active users.
  • Idle connections cost nothing, so a node can hold unlimited sockets.
  • A server's listening port limits it to about 65,000 client connections.
  • Run each node near its measured ceiling to save money.
  • Memory measured on idle connections stays the same under a message burst.