After a push channel drops and reconnects, why are the cache entries it fed wrong, and what restores them?
answer
- the gap's events are already gone
- wrong, and nothing marks it wrong
- resubscribe, then read
- scope the read to keys with readers
- connection state is ordinary state
basics
~20 sBecause the changes made while disconnected were never delivered, so entries sit at pre-gap values with nothing marking them stale. Reconnect must resubscribe and trigger a catch-up read of the keys that still have readers.
solid answer
~50 sA reconnected channel resumes the *future*; it does not restore the past. Whatever changed during the gap was emitted while nobody was listening, so the affected entries hold pre-gap values and look exactly as fresh as correct ones — the failure is silent, which is what makes it dangerous. The reconnect path therefore has two jobs: resubscribe the keys that still have readers, and **invalidate those keys so a catch-up read refills them** from the server. Some transports can resume from a marker and replay part of the gap, but a client must not depend on it; the read is what guarantees convergence. The third piece is visibility: hold the connection state in the store like any other value, so the view can mark the screen as reconnecting or possibly stale instead of confidently rendering old data.
go deeper
Remember that a reconnected stream starts again from now. Anything that changed while you were disconnected was never sent, so the app has to ask the server again.
Explain that the entries are wrong and unmarked, and that reconnect means resubscribe plus invalidate-and-read for the keys that still have readers, while continuing to serve the current value.
Show the production detail: scope and coalesce the catch-up read so a mass reconnect is not a herd, expect the collision with resumed events, and render connection state so a degraded channel is visible.
Set the policy: which screens may keep rendering possibly-stale data, which must degrade or disable actions, and what the client is allowed to assume about replay — none of which the transport decides for you.
## A gap corrupts the cache silently When a live channel drops, nothing in the cache changes — and that is the problem. The entries the channel was feeding keep their last pushed values, keep whatever freshness marker they had, and keep rendering. Meanwhile the server goes on changing the underlying records and emitting events into a connection nobody is receiving. When the transport comes back, it resumes the flow of *new* events. The events from the gap were emitted once, to an audience of nobody. So after a reconnect the cache contains entries that are: - **wrong**, by an unknown amount; - **indistinguishable** from correct entries, because nothing marked them; - **self-healing only by accident**, if the record happens to change again soon. A screen can sit for an hour showing a value that stopped being true during a twenty-second dropout. ## The catch-up read The cure is a read, not a cleverer channel. On reconnect the bridge should walk the registry of keys that still have readers and, for each one: 1. **resubscribe** it on the new connection, so future events land again; 2. **invalidate** it so the next read refetches, or refetch it immediately if the screen is visible; 3. **keep serving the current value** while that read runs, so the user sees stale-then-correct rather than an empty screen; 4. **apply the collision rule** as the read lands, because the catch-up read and the resumed events are racing by construction. Step four is the part teams forget. Reconnect is the moment when every subscribed key has both a request and a push in flight, so whatever ordering rule the cache has — a revision comparison, or a refetch-after-collision fallback — is exercised here first and hardest. ## Resume markers do not remove the read Some transports can tell the server where the client stopped and receive part of the gap back. When it works it is cheaper than refetching. A client should still not build correctness on it: | Reason | Consequence for the cache | |---|---| | Retained history is finite | a long gap exceeds it and the replay is incomplete without saying so | | Resumption is negotiated, not guaranteed | a failed resume looks like a normal reconnect | | Replayed events face the same merge limits | a notification-only event still forces a read | Treat replay as an optimisation that may reduce the catch-up read's scope, never as a substitute for having one. ## Make the channel visible A live screen is built on a connection the tree does not control, and the tree should say so. Hold the connection state — connected, reconnecting, offline — in the store as ordinary state, so it is read like any other value and rendered like any other state: - a persistent but quiet indicator while reconnecting, not a modal; - a clear "data may be out of date" treatment once the gap has outlasted a threshold you choose; - a manual refresh affordance, which is both an escape hatch and an honest admission; - for anything a user acts on financially or destructively, a stronger degrade: disable the action rather than let them act on a value you cannot vouch for. The anti-pattern is the confident screen: a live dashboard with no channel state, rendering numbers it has no reason to believe. Users discover the dropout by comparing with someone else's screen, which is the worst way to learn it. ## Getting the read's scope right Refetching every key the client ever subscribed to turns every dropout into a thundering herd, and if many clients dropped together — the usual cause — they all do it at the same instant. Bound it: - refetch only keys that **still have readers**, and prioritise **visible** ones; - **coalesce** duplicate invalidations into one request per key; - let entries nobody is reading be refetched lazily on their next read; - spread the herd, since a reconnect storm is a self-inflicted load spike on a server that has just had a bad minute. ## What a strong answer contains Name the silence: the gap's events are gone, and nothing marks the entries. Name the cure: resubscribe plus a catch-up read scoped to keys with readers. Name the race: the catch-up read and resumed events collide by construction. Name the visibility: connection state is ordinary state, so a degraded channel is on screen rather than hidden. That is the whole mechanism, and it is the part that distinguishes someone who has run a live screen in production from someone who has only opened a connection.
- Why is a reconnect the most likely moment to hit the push-versus-response race?Because the catch-up read and the resumed event flow start together for every subscribed key. Each entry has a request and a push in flight at once, so an ordering rule that is never exercised in normal use runs on every key simultaneously — and the responses involved were assembled just as the events they conflict with were emitted.
- How do you keep a catch-up read from becoming a thundering herd?Scope and spread it. Refetch only keys that still have readers, prioritise the visible ones, coalesce duplicate invalidations into one request per key, and leave unread entries to refetch lazily. Clients usually drop together, so a spread across clients keeps a recovering server from being hit by all of them at the same instant.
- What should the UI show while the channel is down?A quiet, persistent indication that updates are paused, escalating to an explicit staleness treatment once the gap outlasts a threshold, plus a manual refresh. For destructive or financial actions, disable rather than warn: acting on a value you cannot vouch for is worse than making the user wait.
A live broadcast you had the radio off for. Nothing was held for you, and the show simply carries on when you switch back — so you ask for a summary of where things stand rather than assume you missed nothing.
saying these in an interview costs you the question
- Assuming a reconnected channel replays what was missed
- Trusting entries after a gap because nothing marked them stale
- Refetching every key the client ever subscribed to on reconnect
- Clearing the cache on disconnect so the screen empties out
- Keeping connection state outside the store where no view can read it
- Letting a live screen render confidently with no channel indicator