A GraphQL subscription socket opens but no events ever arrive — how do you diagnose it?
answer
- Everything looks healthy; that is the clue
- Read the negotiated identifier on both ends
- How fast it closed tells you where it broke
- Connected is not the same as subscribed
- Check what terminates the socket in between
basics
~10 sCheck the negotiated WebSocket subprotocol on the open connection first. An empty or mismatched identifier explains an open-but-silent socket more often than resolvers or authorization do, and it costs one field to rule out.
solid answer
~50 sStart at the connection, not at the resolver. Read the negotiated subprotocol identifier on the open socket from both ends: empty means nothing was agreed and neither side will ever understand the other, and a value the client did not expect means it is speaking the wrong vocabulary. That single check separates a subprotocol mismatch — which looks exactly like a healthy connection to every network probe — from the real alternatives: the subscription never registered, the event source produced nothing matching the filter, or authorization silently dropped the subscriber. Watch the shape of the failure too. A mismatch that dies at connect closes almost immediately, while a mismatch that survives the opening exchange stays open and simply ignores the operation. Then check whether an intermediary in the path is stripping the identifier before it reaches your server.
go deeper
Know the first check: look at the negotiated subprotocol on the open connection before assuming the server or the resolver is broken. An open socket does not mean the two ends understand each other.
Explain the two failure shapes — an immediate close versus an open, silent connection — and why the second happens: the vocabularies coincide at the opening exchange and diverge on the first real operation.
Drive the diagnosis in order: negotiated identifier, then connected-versus-subscribed, then authorization and filters, then the event source, with a synthetic subscriber per subprotocol as the instrument. Suspect intermediaries when the failure is uniform.
Own the observability gap this exposes: a transport whose failures produce no status code and no error payload needs deliberate instrumentation — results per connection, negotiated identifier as a dimension, and synthetic probes — funded as platform work, not per team.
## Why this failure is worth a senior question A GraphQL subscription that is silently dead is one of the least observable failures in an API estate, because almost every signal you own says the system is fine. The socket is open. Connection count is normal. There is no HTTP status, because there is no HTTP response after the connection is established. There is no `errors` array, because no result was ever produced. On a parcel-locker dashboard the symptom is a screen that arrives half-empty: the initial query paints the locker list and every compartment's last known state, and then the live tiles never move — no error, no spinner timeout, just yesterday's data presented as though it were current. Operators do not report it for hours because nothing looks broken. ## Check the cheapest hypothesis first The negotiated subprotocol identifier is visible on the open connection at both ends: the client can read which identifier came back, and the server records which one it attached its handler to. Three readings, three conclusions: - **Empty on the client.** No vocabulary was agreed. The socket is open and both sides are waiting for the other to say something they can parse. This is the classic silent case, and it is common because an open socket with no agreed subprotocol is not, on its own, a network error. - **A value the client did not offer or does not implement.** Rare, and normally rejected outright, but if you see it, the client has been handed a vocabulary it cannot speak. - **`graphql-ws` where the client's code expects the current protocol, or the reverse.** The identifier and the implementation disagree — usually because someone read a project name as a wire string. Remember the crossing: `graphql-transport-ws` is the current protocol, and `graphql-ws` selects the legacy subscriptions-transport-ws vocabulary. One field, three answers, and you have either found the bug or eliminated an entire class of them. ## Read the timing of the failure How long the socket survives tells you where the mismatch bit. A connection that is **rejected or closed within milliseconds** of opening means the mismatch was caught during establishment: the server was offered only an identifier it does not serve, or your own server closed a socket that had no agreed vocabulary. This is the good outcome, because it is loud — the client sees a close and can report it. A connection that **stays open indefinitely and delivers nothing** means the mismatch survived the opening exchange. That happens because the two vocabularies coincide at the very first message and its acknowledgement, so a mismatched pair can connect, appear to complete a handshake, and diverge only when the first real operation is sent under a message name the other side does not recognise. The give-away is a connection-lifetime histogram with a population of long-lived connections that have never produced a single result — count results per connection, not just connections. ## Ruling out the alternatives Once the identifier is confirmed correct on both ends, the remaining causes are ordinary and are distinguished by where the trail stops: 1. **The operation never registered.** The server has a connection but no active subscription for it. A per-connection count of active subscriptions separates "connected" from "subscribed" — they are not the same state, and conflating them is the most common reason this diagnosis stalls. 2. **Authorization dropped it.** The subscriber connected but was not permitted to receive, and the refusal went to a log rather than to the client. 3. **Nothing matched.** The event source is producing events, but none pass the subscription's filter — wrong locker identifier, wrong tenant, a stale value in the variables the client stored. 4. **Nothing was produced.** The upstream event source is genuinely idle. Prove it with a synthetic event rather than by waiting. A synthetic subscriber is the fastest instrument for all four: open a connection with a known identifier, subscribe to a known-active stream, and assert that a result arrives within a bounded window. Run it per subprotocol. On a platform whose subscription service is one subgraph among 62, that probe is what turns "someone says live updates are broken" into "the legacy vocabulary path stopped delivering at 09:14". ## The intermediary you forgot If the client offers an identifier and the server insists it received none, look at what sits between them. A reverse proxy, ingress or load balancer that terminates and re-opens the socket must carry the subprotocol selection through in both directions. One that forwards the upgrade but drops the selection produces exactly the empty-identifier symptom, on every connection, with correct code on both ends. The tell is that it fails uniformly rather than for one client build, and that connecting directly to the service bypasses it. ## What to change afterwards Make the failure loud next time. Close sockets that negotiated nothing instead of holding them open; alarm on connections that live longer than a threshold without producing a result; and put the negotiated identifier on connection metrics and log lines so the next occurrence is a dashboard filter rather than an investigation.
- Why does a subprotocol mismatch sometimes close instantly and sometimes hang forever?An instant close means the mismatch was caught while the connection was being established — the server was offered only identifiers it does not serve, or refused a socket with nothing agreed. A hang means the connection was established anyway and the divergence appeared later: the two vocabularies coincide on the opening message and its acknowledgement, so a mismatched pair can complete that exchange and only fail when the first real operation is sent.
- What metric would have caught this before a user reported it?Results delivered per connection, not connection count. A connection that has been open for hours and produced zero results is the signature of both a mismatch and a dead event source, and neither shows up in connection totals, error rates or HTTP status distributions. Pair it with the negotiated identifier as a dimension, and a synthetic subscriber per subprotocol that asserts a result arrives inside a bounded window.
- The client offers an identifier but the server reports none was agreed — what next?Suspect whatever terminates the socket between them. A reverse proxy, ingress or load balancer that forwards the upgrade but drops the subprotocol selection produces exactly this, uniformly, with correct code at both ends. Confirm it by connecting straight to the service: if the identifier arrives intact there and not through the front door, the fix is in the intermediary's configuration, not in the client or the server.
saying these in an interview costs you the question
- Starts by debugging resolvers, not the connection
- Treats an open socket as proof the protocol matched
- Assumes a mismatch always fails at connect time
- Counts connections but never results per connection
- Never suspects an intermediary dropping the identifier
- Confuses a connected client with a subscribed one