skip to content

Why does the graphql-ws subprotocol define its own ping and pong JSON messages?

level: middleimportance: nice to knowfreq 22%

answer

  1. Two heartbeats on one connection
  2. What the browser is not allowed to reach
  3. Application data counts as traffic everywhere
  4. Symmetric: either side may ask
  5. No id field on either message

basics

~20 s

Because a browser's WebSocket API gives page JavaScript no way to send or observe the transport's own liveness probe. The subprotocol therefore defines ping and pong as ordinary JSON messages that either side may send, carrying no operation id.

solid answer

~40 s

WebSocket already has a liveness mechanism at the transport level, but a browser does not expose it: page JavaScript can open a socket, send and receive application data and observe open, message, error and close, and nothing more. A client that wants to know whether its connection is still alive therefore has to ask at the application layer. The `graphql-ws` subprotocol supplies that: either side may send `ping` once the connection is acknowledged, the receiver **must** answer with `pong` promptly, both may carry an optional free-form payload, and an unsolicited `pong` with no preceding `ping` is legal as a one-way heartbeat. Neither carries an `id` — they concern the connection, not any stream. As ordinary application messages they also count as traffic everywhere, which keeps an idle connection from being reaped.

code

json · 10 lines
json
[
  {
    "type": "ping",
    "payload": { "sentAt": 1775894124317 }
  },
  {
    "type": "pong",
    "payload": { "sentAt": 1775894124317 }
  }
]

go deeper

for a junior

Recall that ping and pong are ordinary JSON messages in this subprotocol, that either side may send a ping, and that the receiver is expected to answer with a pong. Know they carry no subscription id.

for a middle

Explain why a second heartbeat exists at all: the browser API does not expose the transport's own, so a page has no other way to test liveness. Be able to state that an unsolicited pong is legal and that both may carry a free-form payload.

for a senior

Show the operational uses — detecting a severed connection quickly, keeping an idle socket from being reaped, and deriving round-trip latency from an echoed payload — and be able to reason about choosing an interval against the shortest idle timeout on the path.

for a principal

Own the distinction between connection liveness and stream liveness. A fleet that treats a healthy heartbeat as proof that subscriptions are delivering will miss silent stream death entirely; decide what independent signal proves data is still flowing.

## Two heartbeats, one connection WebSocket has a liveness mechanism built into the transport itself. The `graphql-ws` subprotocol nevertheless defines `ping` and `pong` as ordinary JSON **messages**, sent through the same channel as `subscribe` and `next`. Having two heartbeats on one connection looks redundant until you ask who is able to use each one, and that is exactly the question an interviewer is opening. ## The reason: the browser cannot reach the transport one A page's JavaScript gets a very narrow WebSocket API. It can open a connection, send and receive application data, and observe open, message, error and close. It cannot send the transport's own liveness probe, and it cannot observe one arriving. That machinery lives inside the browser and is not exposed. The consequence is concrete. A browser client that wants to answer "is this connection still alive, or has something in the middle quietly gone away?" has no transport-level tool to do it with. Its only option is to send something at the application layer and see whether a reply comes back. So the subprotocol gives it one. A second, more mundane reason: application messages are unambiguously traffic. Every layer that counts bytes, resets an idle timer or logs a message sees them. That is not always true of a transport-level probe, which some intermediaries handle without ever surfacing it to the application above. ## The rules, and they are short * Either side may send `ping` at any time, once the connection is acknowledged. * The receiver **must** reply with `pong` as soon as it reasonably can. * Both messages may carry an optional `payload` object, free-form. If a `ping` carried one, a courteous `pong` echoes it back — useful for correlating a reply with the probe that caused it. * A `pong` may be sent **unsolicited**, with no `ping` preceding it. That makes a one-way heartbeat legal: a client can simply emit a `pong` every N seconds purely as traffic. * Neither message carries an `id`. They are properties of the connection, not of any stream, and they are entirely independent of how many subscriptions are open. That last point is the one candidates get wrong most often. A `ping` says nothing about whether a particular subscription is still producing events; it says the socket and the two protocol implementations on either end of it are still talking. ## What it buys you in practice Three things, and it is worth naming them separately. **Detecting a dead peer.** A connection that has been severed somewhere in the middle — a NAT table entry expired, a machine went away without a close — looks exactly like an idle healthy one from either end. A client that pings on an interval and gives up if no `pong` returns within a deadline can reconnect in seconds instead of hanging until the operating system eventually gives up. **Keeping an idle connection from being reaped.** A subscription can legitimately produce nothing for a long time. In a clinical-trial registry graph, a feed of adverse events for one small study might emit twice a week; every layer in between sees a socket that has said nothing for hours and is entitled to conclude it is abandoned. Periodic `ping`/`pong` traffic prevents that judgement, and because it is application data rather than a transport probe, every layer counts it. **Measuring round-trip latency.** Send a `ping` with a timestamp in the payload, get it echoed in the `pong`, and you have a per-connection latency number that costs nothing and needs no instrumentation on the server. ## Choosing an interval The interval is a client and server setting, not a protocol constant, and the sensible way to pick one is to make it comfortably shorter than the shortest idle timeout anywhere on the path. The failure mode is asymmetric: too frequent costs a few bytes per connection, while too infrequent costs a reconnect storm when a proxy starts reaping sockets. At a peak of 1,200 new connection attempts per minute against the registry's public endpoint, a heartbeat that is a few seconds too slow is not a nuisance — it is the difference between a stable pool of long-lived sockets and a churn of clients reconnecting and re-subscribing. Do not confuse the heartbeat interval with the initialisation window, which is a different timer entirely: that one governs how long the server waits for the first message on a brand-new socket. ## The historical footnote The older `subscriptions-transport-ws` protocol approached this differently, with a server-emitted keep-alive message and no client-initiated probe. The move to a symmetric `ping`/`pong` pair, where either side can ask and either side must answer, is one of the meaningful improvements in the newer subprotocol — it lets the *client* be the one that decides the connection is dead, which is the side that can actually do something about it. None of this appears in the GraphQL specification, which defines no transport and therefore no heartbeat.

  • Is a pong ever legal without a preceding ping?
    Yes. The subprotocol explicitly permits an unsolicited `pong`, which makes a one-way heartbeat legal: a client can emit one every few seconds purely as traffic, with no expectation of a reply and no state to track. That is the cheapest way to keep an idle connection from being judged abandoned, though it tells you nothing about the peer, since you never learn whether anything came back.
  • Does a successful ping tell you a particular subscription is still producing events?
    No, and conflating the two is the usual mistake. `ping` and `pong` carry no `id`; they are properties of the connection. A healthy `pong` says the socket and the two protocol implementations at either end are still talking. A stream that has gone quiet because its source stopped, or because a server-side error killed it silently, looks exactly the same from the heartbeat's point of view.
  • How would you choose the ping interval?
    Make it comfortably shorter than the shortest idle timeout on the path, then leave headroom for a missed beat. The failure modes are asymmetric: too frequent costs a few bytes per connection, while too infrequent costs a reconnect storm when something starts reaping idle sockets — at peak, precisely when you have the most of them. It is a client and server setting, not a protocol constant, and it is a different timer from the server's initialisation window.

The transport's own heartbeat is a signal the building's wiring can send; the JSON ping is you knocking on the wall. Only one of them is available from inside the room the browser puts your code in.

saying these in an interview costs you the question

  • Says ping and pong carry a subscription id
  • Claims a pong proves a given stream is alive
  • Thinks only the server may initiate a ping
  • Believes browser JavaScript can send the transport's own probe
  • Confuses the ping interval with the initialisation window

context