How can one GraphQL subscription endpoint serve both graphql-transport-ws and graphql-ws clients?
answer
- Both vocabularies, one execution engine
- The choice is known before any operation
- Branch at open time, keep the handler
- Never guess the protocol from a message
- Count connections by negotiated identifier
basics
~20 sRegister both identifiers on the endpoint and dispatch per connection on whichever one was negotiated. Each identifier gets its own message-vocabulary adapter over one shared subscription execution, so the schema and resolvers are written once.
solid answer
~50 sAccept both identifiers at the endpoint and branch on the negotiated value, not on message content. Because the subprotocol is agreed once and fixed for the connection's lifetime, the server knows from the first moment which vocabulary a socket speaks and can attach the matching adapter to it. Underneath, both adapters drive the same subscription execution and the same resolvers — only the wrapper differs — so nothing about the schema is duplicated. Two rules matter. Never sniff the shape of an incoming message to guess the protocol: the two vocabularies overlap at the opening message, so sniffing appears to work and then fails on the first operation. And record the negotiated identifier as a dimension on connection metrics; that count is the only honest evidence of how many clients still need the legacy vocabulary, and therefore of when it can be retired.
code
pseudocode · 11 linesonConnectionOpen(connection):
switch connection.negotiatedSubprotocol:
case "graphql-transport-ws":
connection.handler = CurrentProtocolAdapter(executor)
case "graphql-ws":
connection.handler = LegacyProtocolAdapter(executor)
default:
connection.close(reason = "no supported subprotocol")
metrics.increment("subscription.connections",
tags = { subprotocol: connection.negotiatedSubprotocol })go deeper
Know that a single endpoint can accept both identifiers and that the server picks one per connection. You are not expected to design the adapter layer, only to recognise that both can be served at once.
Explain the dispatch: the negotiated identifier is known at open time and fixed, so the server attaches a vocabulary adapter per connection over one shared execution. Say why message sniffing fails late rather than early.
Demonstrate the operational half — tag connection metrics by negotiated identifier, keep an end-to-end test per adapter so the legacy path does not rot, and make retirement a decision driven by which clients remain.
Own the cost of carrying two vocabularies indefinitely against the cost of stranding devices that cannot be upgraded, and set the platform default so every new service registers both consistently rather than each team choosing.
## Why dual support is the normal answer A subscription client is not a browser tab you can force-refresh. It is often an installed application, a kiosk, a device or a third-party integration on its own release cadence. If a parcel-locker platform has 3,140 shipped locker kiosks holding open subscriptions, a flag day where the server stops answering the legacy identifier at 02:00 means every kiosk that has not been updated goes dark and stays dark. So the practical shape is: serve both identifiers for as long as the slower half of the fleet needs, and make the retirement decision on evidence. ## The mechanic that makes it cheap One fact does all the work: **the subprotocol is agreed once, when the connection is established, and is fixed for that connection's lifetime.** The server therefore knows, before a single operation arrives, which of the two vocabularies this particular socket speaks. That turns dual support from a protocol-parsing problem into a dispatch problem — pick a handler at open time and keep it. ```pseudocode onConnectionOpen(connection): switch connection.negotiatedSubprotocol: case "graphql-transport-ws": connection.handler = CurrentProtocolAdapter(executor) case "graphql-ws": connection.handler = LegacyProtocolAdapter(executor) default: // No agreed vocabulary: nothing can be understood on this socket. connection.close(reason = "no supported subprotocol") metrics.increment("subscription.connections", tags = { subprotocol: connection.negotiatedSubprotocol }) ``` Everything below `executor` — schema, resolvers, subscription execution, authorization, the event source — is written once. The two adapters are thin: they translate an incoming operation into a subscription execution, and translate each produced result, each field error and each stream ending into whichever message names their vocabulary uses. ## The rule you must not break: dispatch on the negotiated value, never on message shape The tempting shortcut is to accept any connection and infer the protocol from the first message that arrives. It is a trap, and it is a trap in a specific, delayed way: the two vocabularies agree on the opening initialization message and its acknowledgement. Sniffing therefore succeeds in every smoke test — the connection opens, the handshake message is recognised, the acknowledgement goes back — and then diverges when a real operation arrives under a message name the guessed adapter does not know. You have converted a connect-time failure into a runtime one that only shows up on the second message, in production, on the clients you have least visibility into. The same rule holds for the reverse case: if no identifier was agreed at all, do not carry on optimistically. A socket with no agreed vocabulary cannot be understood; close it with a clear reason rather than leaving it open and silent. ## Preference, and who actually decides A client may offer several identifiers, conventionally in preference order — current first, legacy second. The selection, however, is the server's: it answers with exactly one. A server that supports both should prefer the current identifier whenever a client offers it, so a client that has been upgraded stops using the legacy vocabulary the moment it reconnects, with no coordination. That single policy is what makes a fleet drain itself as clients update. ## Making the retirement decision on evidence The metrics dimension in the snippet above is not decoration. Without it, "can we drop the legacy identifier?" is answered by opinion. With it, you can watch a real curve: on a parcel-locker platform whose subscription service is one subgraph of a 62-subgraph estate, the legacy share fell from 100% to 412 of 3,140 kiosks over 47 days of staged firmware rollout, and the remaining 412 turned out to be two hardware revisions that could not take the update at all. That is a completely different conversation from a percentage — it names the clients, so someone can decide whether to buy the replacement, keep the legacy adapter indefinitely, or route those kiosks somewhere else. Two refinements are worth having. Record the identifier on connection *close* as well as open, with the close reason, so a spike of short-lived legacy connections is visible as churn rather than healthy usage. And keep an integration test per adapter that drives a real subscription end to end; the legacy path is the one nobody exercises by hand, so it is the one that silently rots while everyone develops against the current vocabulary. ## Where dual support does not reach Dual support ends at your own endpoint. Anything in the path that terminates or forwards the socket — a reverse proxy, an ingress, a load balancer — must also pass the negotiated identifier through faithfully. If it does not, the endpoint sees no agreed vocabulary no matter how many adapters you registered, and every connection lands in the `default` branch above.
- Why is inferring the protocol from the first message received a bad idea?Because the two vocabularies agree on the opening initialization message and its acknowledgement, so the inference looks correct through connect and handshake and only fails when the first real operation arrives under an unfamiliar message name. That converts a clean connect-time rejection into a delayed runtime failure on exactly the clients you can least easily debug. The negotiated identifier is already known at open time and is unambiguous — use it.
- If a client offers both identifiers, which should the server pick?The current one, whenever it is offered. The server makes the selection, so preferring graphql-transport-ws means an upgraded client moves off the legacy vocabulary automatically on its next reconnect, with no coordinated cutover. It also makes the legacy connection count a true measure of clients that genuinely cannot speak the current protocol, rather than a mix of those and ones that simply asked in an unlucky order.
- How do you know when the legacy adapter can finally be deleted?Tag connection open and close metrics with the negotiated identifier and watch the legacy count, not a percentage. Percentages hide a small, permanent tail. The useful answer names the remaining clients — a firmware revision, an integration, an internal tool — so the decision becomes whether those specific clients can be upgraded, replaced, or supported indefinitely. Delete the adapter when that count is zero and has stayed zero across a full client release cycle.
saying these in an interview costs you the question
- Sniffs the first message to guess the protocol
- Thinks supporting both means two schemas
- Leaves a socket open when nothing was negotiated
- Assumes clients can be force-upgraded on a flag day
- Prefers the legacy identifier when both are offered
- Never measures which identifier clients actually use