A WebSocket to a turbine maintenance console has carried no traffic for an hour and still looks connected; how do you establish whether it is alive?
answer
- silence is not evidence
- the write fails, not the read
- probe faster than the shortest idle timeout
- count consecutive misses, not one
- only a Pong resets the counter
basics
~20 sSend Ping frames on a timer and require a Pong within a deadline. A silent TCP connection cannot distinguish a healthy idle link from one whose state something along the path discarded, so only a probe — judged over several consecutive misses — settles it.
solid answer
~50 sSilence proves nothing. A connection with no traffic on it looks identical whether the peer is idle and healthy or gone entirely, and your side will keep reporting the connection as established until a write finally fails after the transport gives up retransmitting. That can be an hour after the fact. The fix is to ask: send a Ping control frame on a timer and require the matching Pong within a deadline. Two numbers carry the design. The **interval** must sit comfortably below the shortest idle timeout anywhere on the path, so your own probe traffic keeps the connection from being reaped for silence. The **deadline** should span several consecutive unanswered probes, because one lost frame under load is not a dead peer. When the count is exceeded, attempt a Close frame — best effort, since the path is probably gone — and tear the connection down locally.
code
pseudocode · 13 linesPROBE_INTERVAL = shortest_idle_timeout_on_path / 3
MISSES_TOLERATED = 2
every PROBE_INTERVAL:
if outstanding_probes > MISSES_TOLERATED:
send control frame CLOSE with status 1001 // best effort
release connection locally
return
send control frame PING with payload = next_sequence_number()
outstanding_probes = outstanding_probes + 1
on receive control frame PONG:
outstanding_probes = 0 // only a PONG resets the countergo deeper
Know that a connection carrying no traffic tells you nothing, and that a Ping frame answered by a Pong frame is how a WebSocket endpoint checks its peer is still there.
Explain the mechanics of half-open: reads block silently because nothing is arriving, and the error only appears when a write exhausts the transport's retransmissions.
Show the operating judgement — an interval set against the shortest idle timeout on the path, a tolerance of several consecutive misses, and a best-effort Close frame before releasing the connection.
Weigh detection latency against probe cost across a whole fleet of idle connections, and decide where protocol-level liveness is enough and where an application acknowledgement that proves work is progressing must replace it.
## Why silence is not evidence A WebSocket sits on one TCP connection for its entire life, and TCP is a stateful abstraction held in memory at both ends and, invisibly, in the middleboxes between them. When no bytes flow, none of those parties learns anything. The connection is not "idle" in any observable sense — it is simply not being exercised. So a connection that has carried nothing for an hour can be in any of these states, and they look the same from your side: - the peer is connected, healthy and has had nothing to say; - the peer's process died without writing a Close frame; - an intermediary on the path discarded the flow's state and will drop or reset anything that arrives afterwards; - a network path changed and the peer can no longer reach you at all. The local view — an established connection, a read that is quietly blocked — is identical in all four. ## What half-open costs A **half-open** connection is one where one side believes the connection exists and the other does not. Two consequences follow, and both are what makes this an operational question rather than a theoretical one: 1. **The failure surfaces on a write, not a read.** A read on a half-open connection does not fail; it simply never returns anything, because nothing is coming. The endpoint discovers the truth only when it next writes and the transport exhausts its retransmissions, or a reset comes back. On a feed that is silent for long stretches, that can be the *first message you actually needed to deliver*. 2. **Everything above it is misinformed.** A dashboard drawn from connection counts shows healthy sessions. Routing that assumes a live socket keeps handing work to it. The resources behind each connection stay held. For a maintenance console the failure mode is precise and ugly: the operations room sees a connected console and no new readings, and reads the absence of alarms as good news. ## The probe loop 1. On a timer, send a Ping control frame with a payload you can correlate, and increment a count of outstanding probes. 2. On receiving a Pong, reset that count to zero. Only a Pong resets it — resetting on any inbound frame would mean a chatty peer masks a broken return path, and would make "consecutive unanswered probes" untrue. 3. If the count exceeds the tolerance before a Pong arrives, declare the peer dead. 4. Attempt a Close frame on the way out. It is best effort — if the path really is gone it will never arrive — but it costs nothing and, when the peer is merely slow rather than absent, it ends things cleanly. 5. Release the connection locally regardless, and let whatever owns reconnection take over. ## Choosing the two numbers | Setting | What it trades | A reasonable starting point | |---|---|---| | Probe interval | Detection latency against per-connection cost across the whole fleet | Comfortably below the shortest idle timeout on the path — a fraction of it, not just under it | | Missed-probe tolerance | False positives against how long a dead peer is believed alive | More than one, few enough that detection stays within your latency budget | | Deadline per probe | Sensitivity against tolerance of a loaded peer | Generous: the specification only recommends that a Pong be sent as soon as is practical | The interval is the number people get wrong. An intermediary that reaps connections carrying no bytes will cut yours at its own threshold, so a probe interval **longer** than that threshold does not merely detect the death late — it *causes* the death, by leaving the connection silent long enough to be reaped. The probe traffic is doing two jobs at once: proving liveness and keeping the path warm. ## What each signal actually proves | Signal | Proves | |---|---| | The connection is in an established state locally | Only that your own side has not torn it down | | Transport-level keepalive | The peer's networking stack answers; its defaults are typically measured in hours | | A Pong answering your Ping | The peer's process read a frame and wrote one back, so both directions work | | An application acknowledgement of a real message | The above, plus that the peer's own work path is progressing | The last row is the honest upper bound: a protocol-level probe proves the socket, not the application behind it. Where the cost of a silently stalled consumer is high, the acknowledgement is the signal worth building — and the Ping is what catches the case where nothing at all is coming back.
- Why does a half-open WebSocket surface as a failed write rather than a failed read?A read has nothing to fail on — no bytes are arriving, which is indistinguishable from an idle peer, so it simply blocks. A write forces the issue: the transport retransmits, exhausts its attempts or receives a reset, and only then reports the error. On a connection that writes rarely, that report can be an hour late.
- Why not just rely on transport-level keepalive?Two reasons. Its default intervals are typically measured in hours, which is far too coarse for a live feed, and tuning them is an operating-system concern rather than the application's. More importantly it probes the peer's networking stack, not its process: a stack can answer while the application above it has stopped reading entirely.
- What does declaring the peer dead after several missed probes cost if you are wrong?A connection torn down under a peer that was merely slow, and whatever reconnection costs on your system. That is the trade-off the tolerance sets: one missed probe is a normal loss event under load, several consecutive misses over a window is evidence. Tighten the tolerance only where a stale connection costs more than a spurious teardown.
saying these in an interview costs you the question
- Treats an established transport connection as proof the peer is alive.
- Relies on transport keepalive defaults measured in hours.
- Declares the peer dead after a single unanswered probe.
- Sets a probe interval longer than the path's shortest idle timeout.
- Expects a read to fail first when a connection goes half-open.
- Assumes the Close frame will still reach a peer already gone.