How does a GraphQL client resync its data after a subscription reconnects?
answer
- the socket is not the authority
- one order manufactures a second gap
- duplicates are recoverable, absences are not
- compare a marker before overwriting
- the resync document is not the page document
basics
~10 sResubscribe first, then run a query for current state and apply the buffered events on top of it. The subscription is a change notification; a query is the source of truth on every reconnect.
solid answer
~50 sBecause no protocol-level replay exists, the client rebuilds state from a query whenever continuity is in doubt. The order is **subscribe, then refetch**: reopen the connection and start the subscription, buffering anything that arrives; run a query loading the current state of everything that subscription feeds; write the query result, then apply the buffered events over it. Querying first and subscribing after reopens the hole — anything published between the snapshot and the subscribe lands in neither. Subscribing first can only duplicate information, which is recoverable. To make the overlap safe, both the query and the subscription must select the same identifying field (`id` with `__typename`) and a monotonic marker such as `lastReadingAt` or `version`, so a buffered older event never overwrites fresher query data. Scope the resync query deliberately: it runs on every reconnect, so make it a small, purpose-built document.
code
graphql · 20 linesfragment SensorState on Sensor {
id
__typename
soilMoisture
batteryPercent
lastReadingAt
}
query ResyncPaddock($paddockId: ID!) {
paddock(id: $paddockId) {
id
sensors { ...SensorState }
}
}
subscription WatchPaddock($paddockId: ID!) {
sensorReading(paddockId: $paddockId) {
sensor { ...SensorState }
}
}go deeper
Know the shape of the fix: after a reconnect, ask the server for the current state with a query instead of trusting whatever the socket has delivered so far.
Be ready to walk the sequence and defend its order, and to name the two schema requirements that make the overlap safe — a stable identifier and a monotonic marker on both payloads.
Show you have costed it: a resync runs on every reconnect and flapping connections multiply it, so talk about a purpose-built resync document scoped to visible data rather than a full page refetch.
Own the boundary of the pattern. Decide which parts of the graph are state-shaped and safely resynced, and which are transition-shaped and must not be modelled as a subscription at all.
## The pattern: events notify, queries state the truth The durable answer to a gap you cannot see is not to close the gap on the wire but to stop depending on it. Treat the subscription as a **change notification channel** and a query as the **source of truth**. Client state is then rebuilt from a query whenever the channel's continuity is in doubt — on first load, on every reconnect, and optionally on a slow timer — while events keep it fresh in between. ## The reconnect sequence, and why its order matters The sequence is: **resubscribe first, then refetch.** 1. Re-open the connection and start the subscription. Buffer, in the client, any result that arrives from now on. 2. Run the query that loads the current state of everything the subscription feeds. 3. When the query result lands, write it, then apply the buffered events on top of it. Doing it the other way — query first, subscribe after — reopens exactly the hole you are trying to close. Anything published between the query's snapshot and the subscribe lands in neither, and you have manufactured a second, smaller gap on every reconnect. Subscribing first can only produce *duplicate* information, which is recoverable; querying first produces *missing* information, which is not. ## Making the overlap harmless Because step 1 and step 2 overlap, the client will apply some facts twice and may briefly hold an event older than the query result. Two things make that safe, and both are schema decisions: - **Every payload identifies its object.** The subscription result and the query result must select the same identifying field (typically `id`, with `__typename`) so a client can tell it is looking at the same `Sensor` twice, rather than at two unrelated blobs. - **Every payload carries a monotonic marker** — `lastReadingAt`, or an integer `version` — that the client compares before overwriting. Apply an event only if its marker is newer than what is already held. Without one, a buffered event from 14:02:11 can silently overwrite the fresher 14:02:24 value the query just returned. A convenient consequence of GraphQL's document model: the query and the subscription can select fields through the **same named fragment**, so both write exactly the same field set for the same type, and the merge is a field-for-field replace rather than a partial patch. ## Scoping the refetch The refetch is a full query execution on every reconnect, and a flapping connection turns it into a repeated one. Two disciplines keep that affordable. First, scope the refetch to what the subscription actually feeds — the 19 paddocks currently on screen, not every sensor in the account. Second, make the refetch cheap by design: a list of identified objects with a small field set is a far better resync document than the page's original deep query, and it can be a different, purpose-built operation. ## Where the pattern is not enough Resubscribe-and-refetch restores *current state*. It cannot recover *transitions*. If the domain needs every event — an irrigation valve's open/close log, a metered total, an audit trail — then a snapshot of the present does not tell you what happened while you were away, and no amount of refetching will. That data needs either an explicit application-level replay contract or, more often, to be modelled as a paginated query over an append-only log the client reads with a cursor rather than as a subscription at all. ## What a client library does and does not do for you Reconnection logic is the part that is usually already written for you: a socket closes, something reopens it with a delay, and the stored operations are started again. Treat that as the easy half. It restores the *channel*; it says nothing about the *data*, because nothing on the server kept the results published while the channel was down. The resync is application code either way — the query to run, the scope it covers, the buffer for in-flight events and the marker comparison are all decisions the library cannot make, and a client that stops at "the library reconnects for us" is a client that has automated the reopening of a hole. ## What an interviewer is listening for Three things. That you name a query, not the socket, as the authority. That you get the ordering right and can say *why* subscribe-then-query is the safe order. And that you can state the cost — one extra query per reconnect, sized deliberately — instead of pretending the pattern is free. A candidate who answers "the library handles reconnection for us" has answered a different question: reconnecting is the easy half, and the client library reopening a socket does nothing about the events that were published while it was closed.
- Why subscribe before refetching rather than after?Because the two failure modes are not symmetrical. Subscribing first means some facts arrive twice, and a marker comparison makes that harmless. Refetching first means anything published between the snapshot and the subscribe reaches neither path, so every reconnect quietly manufactures a second, smaller gap — exactly the loss the resync exists to repair.
- The resync query is expensive across 19 paddocks and thousands of sensors. What do you do?Scope it rather than skip it. Refetch only what the subscription actually feeds and what the user can currently see, and build a dedicated resync document — a flat list of identified objects with a handful of fields — instead of replaying the page's original deep query. Off-screen data can wait until it is displayed, at which point it is fetched anyway.
- What does this pattern still fail to recover?Transitions. A snapshot tells you the present, never what happened while you were away, so a valve's open/close history or a metered total cannot be reconstructed by refetching. That data needs an explicit replay contract or, more usually, to be modelled as a paginated query over an append-only log rather than as a subscription.
saying these in an interview costs you the question
- Refetches first and subscribes afterwards
- Says the client library's reconnect handles data loss
- Applies buffered events without comparing a marker
- Replays the whole page query on every reconnect
- Treats the socket rather than a query as authoritative
- Assumes duplicates are as damaging as missed events