A GraphQL WebSocket endpoint is closing sockets with codes 4401 and 4408 — what does each mean?
answer
- Numbers the application chose, not the transport
- One code is about timing, one about order
- What happened between open and the first message
- Reconnect paths produce a distinctive burst
- An intermediary only ever reports closed
basics
~20 s4408 means the client never sent connection_init inside the server's initialisation window; 4401 means it sent an operation before the connection was acknowledged, or its init payload was rejected. Both are the graphql-ws subprotocol's own application-range codes, not transport ones.
solid answer
~40 sWebSocket reserves the 4000–4999 range for applications, and the `graphql-ws` subprotocol defines its codes there. **4408** is a timing statement: the socket opened and the client's `connection_init` did not arrive before the server's initialisation window expired — usually because the client does work such as fetching a token *after* opening the socket, and that work got slower. **4401** is a sequencing statement: a `subscribe` arrived while the connection was still unacknowledged, classically a reconnect path that replays stored subscriptions on the socket's open event instead of on `connection_ack`. Servers also use 4401 to reject an init payload. Nearby codes worth knowing are 4429, a second `connection_init`, and 4409, a reused live id. Because the codes are application-defined, an intermediary reports only 'closed' — you must record the code and reason yourself.
code
pseudocode · 19 lines# 4408-prone: the socket races an async credential fetch
socket = open(endpoint)
token = await fetchToken() # slow under load -> 4408
socket.send({ type: "connection_init", payload: { authorization: token } })
# safe: nothing awaited between open and the first message
token = await fetchToken()
socket = open(endpoint)
socket.onOpen():
send({ type: "connection_init", payload: { authorization: token } })
# 4401-prone: replaying streams on open instead of on the acknowledgement
socket.onOpen():
for s in storedSubscriptions: send(s) # not acknowledged yet -> 4401
# safe:
onMessage(msg):
if msg.type == "connection_ack":
for s in storedSubscriptions: send(s)go deeper
Recall that these are the subprotocol's own codes in WebSocket's application range, not HTTP statuses: 4408 is a missing or late connection_init, 4401 an operation sent before the acknowledgement.
Explain the client behaviour behind each — work done between opening the socket and the first message causes 4408, replaying subscriptions on the open event instead of on connection_ack causes 4401 — and name 4409 and 4429 as the neighbouring state bugs.
Show the diagnosis: cluster shape, open-to-close latency, whether the connection was ever acknowledged, and how many client messages preceded the close. Be ready to argue why reordering the client beats lengthening the server's window.
Own the visibility problem. Application-range close codes are invisible to every generic layer in front of the endpoint, so the organisation only ever learns about this class of failure if someone decided in advance to record the code, the reason and the socket's lifetime.
## Where these numbers come from WebSocket reserves the 4000–4999 close-code range for applications to define, and the `graphql-ws` subprotocol uses it. So `4401` and `4408` are not transport-level codes with a general meaning you could look up in a networking reference — they are this subprotocol's own vocabulary, and a server that speaks a different GraphQL WebSocket protocol will use different numbers or none at all. That has a practical consequence before you read either code: a generic proxy, load balancer or metrics layer sitting in front of the endpoint will report only "socket closed". If you are not recording the close code and the reason string yourself, on both ends, this whole class of failure is invisible. ## What each code means **4408 — connection initialisation timeout.** The socket opened and then the client said nothing, or said something else, until the server's initialisation window expired. The window is a server setting, typically a few seconds; the protocol fixes no value. A 4408 is always a statement about *timing*: the client's first message was late. **4401 — unauthorized.** The client sent an operation before the connection was acknowledged. `subscribe` arriving while the connection is still in the not-acknowledged state is closed with 4401 rather than answered. Servers also use 4401 when they reject the `connection_init` payload outright. Two neighbours are worth knowing because they show up in the same logs: **4429**, a second `connection_init` on the same socket, and **4409**, a `subscribe` reusing an id that is already live. Both usually indicate a client that has lost track of its own state, typically across a reconnect. ## Reading a 4408 spike 4408 is the interesting one, because a client that has been correct for months can start producing it under load without changing a line. The registry's public results dashboard opens a socket per browser tab. During a results-release window the endpoint saw a **1,200-request-per-minute** peak of new connections, and roughly one in nine of them closed with 4408 within seconds. The client code was, in outline: open the socket, then `await` a fresh access token, then send `connection_init` with it. Under normal load the token call returned in about 40 ms and the sequence was invisible. Under the same peak the token service — sharing capacity with everything else — pushed past the server's initialisation window, so a slice of connections were closed before they had said their first word. The shape of that diagnosis generalises. A 4408 cluster means one of: * the client does work *between* opening the socket and sending `connection_init` — a token fetch, a config load, a lazy module import — and that work got slower; * the client's own event loop is contended, so the message is composed late even though nothing is awaited; * something in the path is buffering, so the first message leaves on time but arrives late; * the server's window was shortened, or the server itself is slow enough that it is not reading the socket promptly. The fix in this case was ordering, not tuning: acquire the credential **first**, open the socket only once it is in hand, and send `connection_init` immediately on open. Raising the server's window would have hidden the problem at the cost of leaving half-open sockets around longer at exactly the moment there were most of them. ## Reading a 4401 spike 4401 is almost always a client-sequencing bug, and it is usually a **reconnect** bug specifically. A client that keeps a table of active subscriptions and replays it when the socket reopens will produce a burst of 4401s if it replays on the socket's open event rather than on `connection_ack`. The signature is distinctive: 4401s that appear only after a network disruption, in bursts proportional to how many subscriptions each client had open, from clients that connected cleanly the first time. The other 4401 source is a rejected `connection_init` payload — an expired or malformed credential. You can usually tell these apart from the timing: a rejection arrives after the client has sent `connection_init`, so the socket carries one client message before closing, while a sequencing bug closes on a `subscribe`. ## Instrumenting so you can answer this at all The senior half of this question is not the code table, it is that you had the data. Record, per closed socket: the close code, the reason string, the elapsed time from open to close, how many messages the client sent, and whether the connection was ever acknowledged. Those five fields separate every case above without a debugger. Add a distribution of open-to-`connection_init` latency and a 4408 spike explains itself the moment it appears. And treat clusters, not individuals. One 4408 is a client on a bad network. Two hundred in a minute, all under three seconds from open, is a deploy or a dependency.
- A 4408 spike appears only at peak traffic, from a client that has been unchanged for months. What is the likely cause?Something the client does between opening the socket and sending `connection_init` got slower — most often an awaited credential fetch, sometimes a lazy module load or a contended event loop. Under normal load it completes in tens of milliseconds and is invisible; at peak the same dependency is slow enough to push the first message past the server's initialisation window. The fix is ordering rather than tuning: acquire the credential first, open the socket second, and send `connection_init` immediately on open.
- How would you tell a rejected init payload from a premature subscribe when both close with 4401?Count the client messages the socket carried before it closed. A rejection arrives after `connection_init`, so the socket saw exactly one client message and the close follows it. A sequencing bug closes on a `subscribe`, so either no `connection_init` was sent at all or one was sent and the client did not wait for the acknowledgement. Recording whether the connection was ever acknowledged separates the two cleanly without a debugger.
- Why not simply raise the server's initialisation window until the 4408s stop?Because it hides the cause and buys a worse failure. The window exists to stop half-open sockets accumulating; lengthening it means more of them survive longer, at exactly the peak where you have the most connections. It also does nothing for the underlying dependency that got slow, which is likely hurting more than the handshake. Raise it deliberately if measurement says the window was simply too tight, not as a response to a spike.
- What do close codes 4409 and 4429 indicate?4409 is a `subscribe` reusing an id that is already live on that socket; 4429 is a second `connection_init` on a socket that already handshook. Both usually mean the client has lost track of its own state, and both typically appear after a reconnect where a stale handler and a fresh one are running against the same connection. Neither is a load problem, so they should be triaged as client bugs rather than capacity.
saying these in an interview costs you the question
- Reads 4401 and 4408 as HTTP 401 and 408
- Assumes an intermediary will surface the close code
- Raises the initialisation window instead of finding the cause
- Treats every 4401 as an expired credential
- Thinks the initialisation window is fixed by the protocol
- Diagnoses single closes rather than clusters