skip to content

The browser WebSocket API performs no automatic reconnection and gives scripts no way to send a protocol-level ping. What must an application-level reconnect-and-heartbeat layer get right for a browser WebSocket client?

level: seniorimportance: must knowfreq 58%

answer

  1. the socket object is single-use
  2. distinguish your close from their close
  3. random delays, not lockstep ones
  4. expect traffic, notice its absence
  5. a new socket knows nothing

basics

~20 s

Reconnect with capped exponential backoff plus jitter, skip reconnecting after a deliberate client close or a permanent application close code, and resubscribe and resync state on every new socket. Heartbeats must be ordinary application messages on a timer, since scripts cannot send WebSocket ping frames.

solid answer

~60 s

The platform gives you one signal — the `close` event — and a dead object, since a `WebSocket` is single-use and reconnecting means constructing a new one. A correct layer distinguishes intentional closes (set a flag before calling `close(1000)`) from drops, and treats codes your server reserves in the 4000–4999 range as permanent so a rejected client stops hammering. Delays need exponential growth, a ceiling, and jitter, or every client returns in lockstep after a server restart. Reset the delay only after a socket has stayed open for a while, otherwise a server that accepts then immediately closes produces a hot loop. Listen for the window's `online` event to retry immediately instead of waiting out a backoff. For heartbeats, scripts cannot send ping frames, so send your own small message on an interval and close the socket yourself if nothing arrives in time — that is the only way to notice a half-open connection where `readyState` still says OPEN. Finally, clear every timer and drop every listener on the old socket, and resubscribe plus request a delta after each reconnect.

code

javascript · 34 lines
javascript
const PERMANENT = new Set([4001, 4003]); // codes this protocol defines as fatal
let attempt = 0;
let intentional = false;
let socket = null;
let pending = null;
let lastSeq = 0;

function connect() {
  pending = null;
  socket = new WebSocket('wss://example.com/feed');
  const controller = new AbortController();
  const opts = { signal: controller.signal };
  let stable = null;

  socket.addEventListener('open', () => {
    stable = setTimeout(() => { attempt = 0; }, 10000); // reset only once stable
    socket.send(JSON.stringify({ type: 'resume', since: lastSeq }));
  }, opts);

  socket.addEventListener('close', (event) => {
    clearTimeout(stable);
    controller.abort(); // drop every listener on this dead socket at once
    if (intentional || PERMANENT.has(event.code)) return;
    if (pending) return; // never stack reconnect attempts
    const base = Math.min(30000, 250 * 2 ** attempt++);
    pending = setTimeout(connect, base * (0.5 + Math.random() / 2));
  }, opts);
}

window.addEventListener('online', () => {
  if (pending) { clearTimeout(pending); connect(); } // do not wait out the backoff
});

connect();

go deeper

for a junior

Know that the browser never reconnects a WebSocket for you, that a closed socket cannot be reused, and that reconnecting means constructing a new WebSocket and reattaching its handlers.

for a middle

Explain exponential backoff with a ceiling and why jitter matters, how the close code decides whether to retry at all, and why a heartbeat has to be an application message rather than a protocol ping.

for a senior

Show the failure modes you have actually hit: hot loops from resetting backoff too early, stacked sockets from missing teardown, false positives from throttled timers in hidden tabs, and half-open connections that readyState never reveals.

for a principal

Own the client-server contract around reconnection: which close codes mean stop, how clients resume from a sequence number, and what a fleet of thousands of jittered clients does to your servers during a rolling deploy.

## What the platform actually gives you Very little, and that is the point of the question. A `WebSocket` object is single-use: the ready state only moves forward and `CLOSED` is terminal, so "reconnect" always means constructing a new object and rewiring its handlers. The browser will not do this for you, will not tell you *why* a connection failed beyond a close code, and will not queue anything you tried to send while disconnected. Everything else described here is code you write. ## Deciding whether to reconnect at all The `close` event is the whole input. Three cases need separating. **You closed it.** Set a flag immediately before calling `close(1000, …)` — during a logout, a route change, or component teardown — and check it in the close handler. Without that flag the layer cheerfully reconnects a socket the application deliberately shut down, which is the most common bug in hand-rolled reconnect code. **The server closed it permanently.** Refusing a connection is useless for signalling *why*, because a rejected connection surfaces as an opaque abnormal close. So the server accepts, then closes cleanly with a code your protocol reserves — say 4001 for a bad token, 4003 for a suspended account. Those must map to "stop and surface an error to the user", not to backoff. **Everything else.** Retry. ```js socket.addEventListener('close', (event) => { if (intentional) return; if (PERMANENT_CODES.has(event.code)) { surfaceFatal(event); return; } scheduleReconnect(); }); ``` ## Backoff, jitter, and the reset trap Delays must grow — 250 ms, 500 ms, 1 s, up to a ceiling of perhaps 30 seconds — and every delay must carry randomness. Jitter is not a nicety: after a server restart, thousands of clients all disconnect within the same second, and a deterministic backoff brings them all back in the same instant, re-creating the overload that dropped them. Randomising each delay across a range spreads the return. The subtle part is *when to reset the counter*. Resetting on the `open` event looks right and is wrong: a server that accepts connections and immediately closes them — because it is overloaded, or the token is stale — turns the backoff into a hot reconnect loop at full speed. Reset only after the socket has stayed open for some minimum period, or after the first successful application message. ## Reacting to the environment instead of only to the clock Waiting out a thirty-second backoff when connectivity returned five seconds ago is a poor experience. The window fires `online` and `offline` events, and `navigator.onLine` reports the browser's current belief. Treat both as hints — `online` can be true behind a captive portal — but use `online` as a trigger to attempt a reconnect immediately, collapsing the remaining delay. Conversely, when the browser reports offline there is little value in burning attempts. Background tabs matter too. Browsers heavily throttle timers in hidden tabs, so a `setInterval` heartbeat scheduled every fifteen seconds may fire roughly once a minute once the tab is backgrounded. A dead-connection detector tuned for fifteen-second intervals will then produce false positives, or notice failures far too late. Take the tab's visibility into account: relax the detector while hidden, and revalidate the connection when the page becomes visible again. ## Heartbeats, and why they are yours to build The WebSocket API exposes no method for sending a control ping and no event for a pong; the browser handles peer pings itself, invisibly to script. So a heartbeat is an ordinary application message — `{"type":"ping"}` or a single byte — sent on an interval, with the server replying in kind. The purpose is to detect a **half-open connection**: an intermediary such as a NAT device, proxy or load balancer silently discards a connection it considers idle, and neither endpoint receives any notification. The browser still reports `readyState === WebSocket.OPEN` and fires no `close` event, possibly for a very long time. The only way to notice is to expect traffic and act on its absence: arm a timer when you send the heartbeat, cancel it whenever *any* message arrives, and if it fires, call `close()` yourself so the normal reconnect path runs. Choose an interval comfortably below the shortest idle timeout in your path — many proxies default to around a minute — and remember that the heartbeat also keeps that path warm. ## Teardown is where reconnect layers leak Each reconnect creates a new socket, and every listener and timer attached to the old one must go. The classic failure is stacking: a drop schedules a reconnect, a second event schedules another, and each generation doubles the number of live sockets until the page is opening dozens per second — a self-inflicted connection storm that looks, from the server, exactly like an attack. Guard with a single pending-reconnect handle, refuse to connect while one attempt is in flight, and clear heartbeat and detector timers in the close handler. Wiring listeners with an `AbortController` and passing `{ signal }` to `addEventListener` lets you remove all of them in one call. ## Resuming state, not just the connection A new socket is a blank slate. Subscriptions do not survive it, and messages the server produced while you were away were not held anywhere. So the layer must, on every `open`: re-send its subscriptions, and ask for what it missed — typically by sending the last sequence number or timestamp it processed and letting the server reply with a delta or a full snapshot. Consumers must also tolerate duplicates, since a message may have been delivered just before the drop and replayed after it. Designing that resume path is the difference between a reconnect layer that restores the connection and one that restores the application.

  • Why reset the backoff delay after a period of stability rather than on the open event?
    Because a server that accepts a connection and immediately closes it — overloaded, or rejecting a stale token — would reset the counter on every attempt, turning the backoff into a full-speed loop. Requiring the socket to stay open for some minimum time, or to complete one successful exchange, means a flapping server still gets exponentially spaced attempts.
  • Why does readyState stay OPEN on a connection that is actually dead?
    Because a middlebox that drops an idle connection sends nothing to either endpoint. With no close frame and no transport-level error surfaced, the browser has no reason to change state, so the socket reports OPEN indefinitely. Only an application heartbeat with a response deadline turns that silence into an observable failure you can act on.
  • What breaks if your reconnect layer forgets to clear the old socket's listeners and timers?
    Generations stack. The dead socket's handlers still fire, each scheduling another reconnect, so attempts multiply and the page opens an accelerating number of sockets — indistinguishable from an attack at the server. A single in-flight guard plus removing listeners, ideally via one AbortController per socket, prevents it.
  • How should the client's message subscriptions survive a reconnect?
    They cannot survive it — a new socket carries no server-side state. The layer must keep its own record of active subscriptions and replay them in the open handler, then send the last sequence number or timestamp it processed so the server can return a delta. Consumers must be idempotent, because messages near the drop can be delivered twice.

saying these in an interview costs you the question

  • Reconnects after the application deliberately closed the socket
  • Retries on a fixed interval with no backoff or jitter
  • Resets the backoff on open, causing a hot loop
  • Believes JavaScript can send a WebSocket ping frame
  • Assumes readyState OPEN proves the connection is alive
  • Forgets that subscriptions do not survive a new socket

context