skip to content

When a server hits the operator-set cap on concurrent WebSocket connections it will hold, what does the next client actually observe?

level: seniorimportance: nice to knowfreq 33%

answer

  1. refused at the door
  2. no socket, so no close
  3. HTTP status, not a frame
  4. established connections keep running
  5. handshake path or accept path

basics

~20 s

A failed handshake, not a closed socket. The refusal lands while the exchange is still HTTP — an error status instead of 101 Switching Protocols — or, enforced lower down, a connection dropped with no response at all.

solid answer

~40 s

A concurrency cap is refused at the door, so the client never reaches the state where WebSocket semantics exist. If the cap is enforced in the handshake path, the server answers the upgrade request with an HTTP error status such as `503 Service Unavailable` and never sends `101 Switching Protocols`. If it is enforced beneath that — at the accept queue — the connection is simply not accepted, or is accepted and dropped, and the client gets a transport failure with no HTTP response at all. In neither case is there a Close control frame, because no socket ever existed. Existing connections are untouched: the cap is about how many are open at once, not how long any one of them lives.

code

http · 10 lines
http
GET /occupancy HTTP/1.1
Host: board.example
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13

HTTP/1.1 503 Service Unavailable
Retry-After: 5
Content-Length: 0

go deeper

for a junior

Know that a connection ceiling stops new WebSocket handshakes from completing and leaves connections that are already open running normally.

for a middle

Explain that the refusal happens while the exchange is still HTTP, so the client sees an error status instead of 101 Switching Protocols and never a close status code.

for a senior

Compare the enforcement points: the handshake path yields a loggable, alertable HTTP refusal, while refusing at the accept path gives the client an unexplained failure and you almost nothing to diagnose with.

for a principal

Own what the refusal means to the fleet: a countable refusal is a capacity signal and a client backoff contract, so decide the status, the retry guidance and the alert threshold once rather than per service.

## A cap is a concurrency setting, not a lifetime An operator running a socket server sets a ceiling on how many connections one node will hold at once. It exists because every open socket costs resources for as long as it is open, whether or not it is carrying traffic — which is the thing that makes a socket server different from a request server, where the cost is bounded by how long a request takes. The important behavioural consequence is the one people get backwards: reaching the cap does **not** evict anybody. Connections already established keep running. What changes is that new arrivals are refused, and the interesting question is what "refused" looks like from the outside. ## Where the refusal lands There are two places the ceiling can be enforced, and they look completely different to a client: 1. **In the handshake path.** The upgrade request arrives, the server counts its open sockets, decides it is full, and answers with an HTTP error status — `503 Service Unavailable` is the natural one — with no `101 Switching Protocols`. The client sees a well-formed HTTP response that is simply not an upgrade. 2. **Beneath HTTP.** The ceiling is enforced on accepting connections at all. Nothing reads the request; the connection is never accepted, or is accepted and closed at once. The client gets a transport-level failure and no response whatsoever. | enforcement point | what the client gets | what the server can log | |---|---|---| | handshake path | an HTTP error status, no `101` | the request, the origin, the reason | | accept path | a connection failure, no response | a counter at best | | not enforced at all | a `101` and a socket the node cannot afford | nothing until the node degrades | The third row is the one to avoid. A node with no ceiling does not refuse anything; it accepts until something else gives way, and then every connection on it suffers together rather than one new arrival being turned away cleanly. ## What the client can and cannot tell From the client's side, a refusal at the cap is indistinguishable from several other failures unless the server made it distinguishable. A rejected upgrade and a node that is not running both produce "no socket". This is why the choice of enforcement point is an operational decision and not an implementation detail: - refusing in the handshake path gives you an HTTP status you can count, alert on and hand to a client team; - refusing beneath HTTP gives you one number on one node and a client-side mystery. Note the thing that is **not** happening: there is no WebSocket Close control frame and no close status code, because a close belongs to a socket and no socket was created. An answer that reaches for a close code here has skipped a step. ## Retrying without making it worse A refused client will try again, and the shape of that retry is what decides whether the cap protects the node or amplifies the problem. A fleet of clients that all retry immediately converts one full node into sustained load against every node. A refusal that carries an HTTP status at least lets the client distinguish "full, come back shortly" from "this endpoint is wrong" and back off differently for each. ## On the occupancy board The board's node is sized for the city's operator consoles plus a public display feed. The operator sets the ceiling deliberately — how that number is chosen is a capacity question with its own arithmetic and is not what this question is about — and then makes two choices that are: 1. enforce it in the handshake path, so a refusal is an HTTP response with a reason; 2. count refusals as a first-class signal, because a rising count is the earliest honest sign that the node is at its ceiling — earlier than latency, and much earlier than a user reporting a board that stopped updating. The summary an interviewer wants: at the cap the server stops completing handshakes, existing sockets are unaffected, the refusal is an HTTP-level event rather than a WebSocket-level one, and where you enforce it decides whether you can see it happening.

  • Why prefer refusing in the handshake path over refusing at the accept path?
    Because the refusal becomes an HTTP response you can count, alert on, and explain to a client team, and the client can tell "full, try later" from "wrong endpoint". Refusing beneath HTTP leaves the client with an unexplained connection failure and you with a single counter.
  • A client reports the socket 'closed' when the node was at its cap. Why is that wording wrong?
    Nothing closed — nothing ever opened. A close is a WebSocket event on an established connection; a capped node refuses before the `101`, so the correct description is a failed handshake. The distinction matters because the two have completely different causes and fixes.
  • Does hitting the cap affect connections that are already open?
    No. A concurrency ceiling governs admission, not lifetime, so established sockets keep sending and receiving normally. If existing connections are also dying, something else is wrong — an idle allowance on a hop, a restart, or resource exhaustion the cap was meant to prevent.

saying these in an interview costs you the question

  • Says the server sends a close status code when it is full
  • Thinks reaching the cap evicts existing connections
  • Cannot distinguish a refused handshake from a closed socket
  • Assumes the client can always tell a full node from a dead one
  • Treats a concurrency cap as a limit on handshakes per second