How do you decide which pushed signals belong in the shared data cache and which stay ephemeral?
answer
- ask who reads it back later
- records persist, signals expire
- churn wakes every subscriber of a key
- keyed by session, not by resource
- signals drop first under pressure
basics
~20 sAsk whether anything reads the value back later. Server-owned records a later-mounting component must see belong in the cache. Presence, typing flags and live cursors are true only now, churn constantly, and belong in a short-lived ephemeral store.
solid answer
~50 sThe test is retention, not origin: **will a component that mounts a minute from now need to read this value?** If yes, it is server-owned data and belongs in the cache with a key, a merge rule and an invalidation story. If it is only meaningful right now — who is present, who is typing, where a cursor is — it belongs in a small ephemeral store beside the cache, not persisted, not invalidated, and throttled at the bridge because it arrives far faster than anything worth re-rendering for. Mixing them is expensive: high-frequency signals churn shared entries and wake every subscriber of that key, they inflate a store meant for retention, and they blur what a stale entry means. The grey cases — "last seen", "currently editing" surfaced in an audit trail — are the ones to decide explicitly, usually by asking who else queries them.
go deeper
Not everything the server pushes is data to keep. Things like who is online or who is typing only matter right now and are not stored with the records.
Explain the retention test and the cost of mixing: a high-frequency signal in a cache entry notifies every subscriber of that key, and entries nothing collects grow for the whole session.
Show the operational rules for ephemeral state: throttle at the bridge, key by session, expire on a timer because the stop event is the one a dropped connection loses, and keep it out of write paths.
Own the classification rule, the per-client budget, the order things degrade in when the channel is struggling, and the retention call on signals that describe people rather than resources.
## Two kinds of thing arrive on one channel A live connection typically carries both **state changes** — a record was created, updated, removed — and **ephemeral signals** — presence, typing indicators, cursors, transient progress. They arrive the same way, which is why they end up in the same place by default. They should not be. ## The retention test One question separates them: **will something read this value back later?** - A record another component will mount and read is **server-owned data**. It needs a key, a merge rule, an invalidation story and a catch-up read after a gap. - A signal that is only true at this instant is **ephemeral**. Reading it back later is meaningless; a five-second-old "typing" flag is not stale data, it is noise. | Property | Cached data | Ephemeral signal | |---|---|---| | Read by a component mounting later | yes, that is the point | no, and it would be wrong | | Survives navigation | expected | irrelevant, usually harmful | | Refilled after a disconnect | by a catch-up read | by the next signal, or not at all | | Event rate | change rate of the domain | often continuous while a user acts | | Wrong value costs | user acts on false data | a stale avatar or indicator disappears | | Retention question | invalidation policy | dropped when the subscription ends | ## What mixing them costs 1. **Churn.** A cursor position at input rate writes the same entry continuously; every subscriber of that key is notified each time, so a component rendering the *record* re-renders because someone moved a pointer. In a runtime with fine-grained tracking the blast radius is smaller but non-zero; in one that re-runs the component function it can dominate the frame. 2. **Unbounded growth.** Presence for everyone who has ever appeared, kept under keys nothing collects, is a store that only grows across a session. 3. **A blurred contract.** If "stale" can mean either "the record may have changed" or "this person stopped typing a while ago", nobody can reason about freshness, and the invalidation rules stop making sense. 4. **Retention you did not intend.** Presence and cursors describe *people*, not resources. Persisting them, restoring them from storage or shipping them into diagnostic tooling turns an interaction hint into a record of who was where and when — a decision that deserves to be made on purpose. ## Where ephemeral signals go instead A small store beside the cache, with its own rules: - **keyed by session or participant**, not by resource, and cleared when the subscription ends; - **throttled or sampled at the bridge**, because the useful update rate for a human-visible indicator is far below the event rate; - **not persisted** and not restored on reload; - **expiring on their own**, since the "stopped" event is exactly the one a dropped connection loses — a presence entry with no timeout leaves ghosts on screen forever; - **never a dependency of business logic**; an indicator may read it, a write path may not. ## The grey cases Some signals are genuinely both. "Last seen at" is ephemeral as a live dot and durable as a profile field the server stores. "Currently editing" is an indicator, but if it also appears in an audit view it has become data. The resolution is to ask **who else can query it**: if the server persists it and another surface can read it back, it is data with a key, and the live signal is a faster path to the same value. If only the connected clients know it, it is ephemeral. Deciding by asking which store is convenient is how a presence hint quietly becomes a record nobody meant to keep. ## What a lead actually owns here - **The classification rule**, written down next to the event contract, so each new event type is placed by a principle rather than by whoever wired it. - **The budget**: how many live keys and how many signals per second a client may hold, and what happens when a busy document exceeds it. - **The degradation order**: when the channel is struggling, ephemeral signals are the first thing to drop and state changes the last, because losing a cursor is invisible and losing a record change is a correctness bug. - **The retention decision** for anything that describes people rather than resources — because the cheap default is to keep it, and the cheap default is the wrong one.
- Why must ephemeral presence entries expire on their own?Because the event that says someone left is exactly the one a dropped connection never sends. If a presence entry only disappears on an explicit stop signal, every unclean disconnect leaves a ghost on screen forever. A short expiry refreshed by ongoing signals makes absence the default and presence the thing that must keep proving itself.
- A live cursor is stored in the same entry as the document record. What does that cost?Every subscriber of that entry is notified at pointer rate, so components rendering only the document re-render continuously. A fine-grained runtime narrows the blast radius to the parts that read the cursor field, but the write path and notification still run; a runtime that re-runs the whole component pays the full cost.
- How would you decide a genuinely ambiguous signal such as 'currently editing'?Ask whether the server persists it and whether any other surface can query it back. If an audit or history view can, it is data: give it a key and treat the live signal as a faster path to the same value. If only connected clients know it, keep it ephemeral and let it expire.
saying these in an interview costs you the question
- Putting every pushed signal into the cache because it arrived the same way
- Keying presence by resource so entries accumulate all session
- Writing cursor positions into the record entry components read
- Relying on an explicit stop event to remove a presence entry
- Letting business logic branch on a transient indicator
- Persisting presence or cursor history without deciding to