skip to content

How do you decide how much replay a GraphQL subscription API should promise its clients?

level: principalimportance: should knowfreq 26%

answer

  1. an argument is a promise you cannot withdraw
  2. climb only until the data is satisfied
  3. bounded means a number you can state
  4. per-subscriber buffers are the expensive shape
  5. an in-process buffer dies with the deploy

basics

~20 s

Decide per root field, and stop at the cheapest rung that fits the data: no replay with refetch on reconnect, self-healing payloads, or an explicit cursor argument with a stated retention window you can actually operate and test.

solid answer

~50 s

The moment a subscription root field takes a position argument, replay is a published contract you must operate, alert on and test — so the question is which data needs it, not whether it would be nice. Work up a ladder. Rung 0: no replay; clients resubscribe and refetch, which costs one query per reconnect and no server state. Rung 1: self-healing current-state payloads, pure schema design, shrinking exposure to one inter-event interval. Rung 2: a cursor argument replaying from a **bounded** retention window — and none of it comes from GraphQL or either subprotocol, so you must define the window, the out-of-range behaviour (an explicit resync signal, never a silent jump to the present), whether buffers are per-topic or per-subscriber, and what survives a redeploy. Split the schema instead of picking one guarantee: state-shaped fields at rung 1, transition-shaped data as a paginated query over an append-only log.

code

graphql · 8 lines
graphql
type Subscription {
  """Current state. No replay: resubscribe and refetch after a gap."""
  sensorReading(paddockId: ID!): SensorReading!

  """Replays from `since` within a 15-minute window, then goes live.
  A `since` older than the window fails with CURSOR_EXPIRED; resync in full."""
  valveCommands(paddockId: ID!, since: Cursor): ValveCommand!
}

go deeper

for a junior

Take away the default: most subscriptions promise no replay at all, and clients are expected to refetch current state after a reconnect rather than ask for missed events.

for a middle

Be able to explain why a replay window must be a stated bound — a time or a count — and what the server has to keep in order to honour a position a client hands back to it.

for a senior

Argue the operational side: buffer shape and memory cost, behaviour when a cursor falls outside retention, and what a redeploy does to anything held in process memory.

for a principal

Own the contract decision. Choose per root field rather than per API, keep the guarantee at the cheapest rung the data tolerates, and refuse to publish one nobody has tested by deliberately breaking it.

## Start by noticing that "replay" is a public API promise The moment a subscription root field takes a `since:` argument, replay is part of your published schema and part of your support contract. Arguments are hard to withdraw, clients build on them immediately, and the guarantee you advertise becomes a thing you must operate, alert on and test. So the decision is not "would replay be nice" — it is "which data genuinely needs it, and what am I signing up to run for the next several years". Work up the ladder and stop at the first rung that satisfies the data. ## The ladder, cheapest first **Rung 0 — no replay; clients resubscribe and refetch.** The default, and correct for most data. It costs one query per reconnect and no server state at all. Adequate whenever the client needs to be right *now* rather than to know everything that happened. **Rung 1 — self-healing payloads.** Each result carries the identified object's current state, so the next event repairs the gap. Pure schema design, no infrastructure, and it shrinks the exposure at rung 0 from "until reload" to "one inter-event interval". **Rung 2 — an application-level cursor argument.** The subscription root field accepts an opaque position and the server replays from it out of a bounded retention window before joining the live stream. This is the first rung that costs real infrastructure and real semantics — and note that nothing in GraphQL or in either WebSocket subprotocol gives you any of it. Everything here is yours to define. **Rung 3 — durable per-subscriber positions.** The server remembers where each subscriber got to, which requires identity, storage, expiry and a garbage-collection story for subscribers that never return. Very few product surfaces earn this; the ones that do are usually better served by a different API shape entirely. ## What rung 2 forces you to define, explicitly - **A retention bound**, in time or in count, per topic. "We keep 15 minutes" is a promise you can size and test; "we keep a buffer" is not. - **The out-of-range behaviour.** When a cursor is older than retention, the server must say so distinctly, so the client performs a full resync. Silently starting from the present is the worst possible answer: it converts an honest error into invisible data loss. - **The buffer's shape.** Per-topic buffers are cheap and shared; per-subscriber buffers are not. With 1,900 concurrent dashboards, retaining 3,148 events each at roughly 512 bytes is about 3 GB of live memory, while the same retention held once per topic across 19 paddocks is a rounding error. Per-topic is almost always the right shape, and it constrains the cursor to be a position in a topic rather than in a subscriber's private view. - **Survival across restarts.** An in-process buffer dies with the process. If you deploy several times a day, a 15-minute retention promise is fiction for every subscriber reconnecting during a rollout — retention either lives outside the process or the guarantee shrinks to your uptime. - **Ordering.** Replay implies a total order per topic that you now must actually provide and preserve through whatever event backbone sits behind the resolver. Do not promise a position in a stream that has no stable order. ## Decide per root field, not per API The useful move is to split the schema rather than pick one guarantee for everything. Current-state fields (`sensorReading`) sit at rung 1 and document plainly that reconnecting subscribers should refetch. Event-log fields (`valveCommandLog`) either get a cursor argument with a stated window, or — often better — stop being subscriptions and become a paginated query over an append-only log, with a subscription that merely signals "there is more". That split gets you an auditable stream with real durability, backed by storage you already know how to operate, and it keeps the socket doing the one thing sockets are good at. ## How you would know it works A replay promise you have never broken on purpose is untested. The verification is deliberately hostile: kill connections during load, restart the process mid-stream, present a cursor deliberately older than retention, and assert that clients converge on the same state as a full refetch. If nobody owns that test, the honest engineering decision is to stay at rung 1 and say so in the schema documentation, because an unverified guarantee is worse than a stated absence — clients build on it exactly as hard.

  • What should the server do when a client presents a cursor older than the retention window?
    Fail the subscription with a distinct, documented signal that tells the client to perform a full resync. The dangerous alternative is starting from the present as if nothing were wrong, which converts an honest, observable error into invisible data loss — and clients will happily build on it, because from their side the stream looks perfectly healthy.
  • How does a deploy cadence interact with a retention promise?
    An in-process buffer dies with the process, so a fifteen-minute window is fiction for anyone reconnecting during a rollout. If you deploy several times a day, either retention lives outside the process — in storage you already operate — or the honest guarantee shrinks to your uptime and should be documented as such rather than advertised.
  • When is the right answer to stop using a subscription for that data entirely?
    When every transition matters and must be auditable. An append-only log exposed as a paginated query with a cursor gives real durability on storage you already run, and a subscription that merely signals *there is more* keeps the socket doing the one thing it is good at. That is usually cheaper than growing replay guarantees inside the subscription itself.
  • How would you validate a replay guarantee before publishing it?
    By breaking it deliberately: kill connections under load, restart processes mid-stream, present cursors deliberately older than retention, and assert that clients converge on the same state as a full refetch. If nobody owns that test, the defensible decision is to stay at self-healing payloads and document the absence, because an unverified guarantee is worse than a stated one.

saying these in an interview costs you the question

  • Adds a replay argument before any data needs it
  • Promises replay with an unbounded in-memory buffer
  • Silently resumes from now when a cursor expires
  • Buffers per subscriber rather than per topic
  • Ignores that a redeploy empties the retention window
  • Applies one delivery guarantee to the whole schema

context