A console reconnects to your Server-Sent Events feed with a `Last-Event-ID` older than anything the server still holds — what should it do?
answer
- the protocol stops at the header
- a gap the client cannot detect on its own
- bounded work beats proportional work
- say it on the wire with a named event
- the window is retention, not a setting
basics
~20 sTreat it as a signalled gap rather than a fresh subscriber: emit a distinctly named event marking the break plus a current snapshot, then continue live. Silently streaming from now on hides missing readings behind a healthy-looking stream.
solid answer
~50 sThe protocol stops at the header, so this is entirely your decision — and the wrong one is the tempting one. Streaming live events from now on is what happens if you write no branch at all: the reconnect succeeds, the warning console fills in again, and the gauge trace has a hole in it that nobody can see. The honest handling is to say so on the wire: a distinctly named `event:` marking that the resume failed, then a snapshot of current state for every gauge, then live events. That bounds the work — a snapshot costs the same whether the client was away for a minute or a week — and it gives the receiving application something to act on. The window itself is a promise: it is only as long as you keep the events your ids address, and it costs retention plus the lookup path on every reconnect.
code
pseudocode · 16 lineson open stream request:
resume_from = request header Last-Event-ID
oldest = oldest id still held in the replay store
if resume_from is absent:
send event named snapshot with current state of every gauge
else if resume_from is older than oldest:
send event named gap with field oldest_available = oldest
send event named snapshot with current state of every gauge
count one window miss
else:
for each stored event after resume_from:
send event with its own id
flush
then stream live events as they occurgo deeper
The takeaway is that a server is not obliged to replay from the id it is given, and when it cannot, the client has no way to find that out by itself.
Describe the branch concretely: missing header, usable id, unresumable id, and what each writes before live events begin.
Argue for bounded work — a named gap event plus a snapshot — and show that you instrument window misses so the retention decision is driven by data rather than by guesswork.
Own the promise: how long the feed claims to be resumable is a contract on the store behind it, and every consumer's catch-up logic is written against the number you publish.
## The specification hands you a value and stops A client that re-opens a dropped stream sends `Last-Event-ID` and nothing more. There is no field for "how far back can you go", no error the client understands as "too old", and no way for it to ask again with a different value. Whatever your handler writes next **is** the resumption, correct or not. So the first thing to be clear about is that this is an emitter design decision with no protocol default. ## Four things a handler can do, and what each costs | response to an unresumable id | what the client ends up with | cost to you | |---|---|---| | stream live events only | a silent gap it cannot detect | none, and that is why it happens | | replay the oldest events still held | wrong events presented as the missing ones | a lookup plus misleading data | | named gap event, then a snapshot, then live | correct current state and knowledge that it missed some | one snapshot per affected reconnect | | end the response and refuse | a client that re-opens on its delay with the same stale id and loops | a reconnect storm of your own making | Only the third is defensible for a warning surface. The first is the default behaviour of a handler that forgot to branch; the second is the failure mode of a replay store that clamps an out-of-range cursor to its oldest entry instead of reporting that it is out of range; the fourth trades a data problem for an availability problem, because a client that cannot progress will keep coming back with the same value. ## Why "gap event plus snapshot" is the shape to reach for A snapshot is **bounded**. Replay is proportional to how long the client was away; a snapshot of current level and current threshold state for every gauge in the catchment is the same size whether the outage lasted ninety seconds or a fortnight. That matters precisely when the window has been exceeded, which is to say when replay would have been largest. The named event matters just as much: 1. It lets the receiving application distinguish "you are caught up" from "you are current but missed some", which are different facts on an operational console. 2. It gives you somewhere to put the boundary — the oldest id you can still serve — so an operator debugging a hole in a trace can see where it starts. 3. It is observable at the emitter: counting how often you emit it tells you how often your window is actually too short, which is the only honest input to lengthening it. ## What the window really is, and what it costs The replay window is not a setting; it is a consequence. It is exactly as long as you keep the events that your ids address, in a form you can enumerate forward from a given id. Three costs follow: - **Retention.** Every extra hour of promised resumability is an hour of events kept somewhere addressable. - **Lookup.** Every reconnect that carries an id does a seek, and reconnects arrive in bursts — a link that drops for a whole district drops for every console in it at once. - **Write-path coupling.** Ids must be assigned where the events are durable, or the window you advertise is longer than the data you can actually produce. Publish the number, and publish the smaller of "what we retain" and "what we are willing to serve". A window you advertise but cannot honour under load is worse than a short one, because consumers will have written their catch-up logic against it. ## A first request carries no header at all Worth folding into the same branch: a client whose remembered id is empty sends no `Last-Event-ID` field. That is the same code path as an out-of-window resume — send current state, then live events — and treating both identically keeps one tested path instead of two. The difference is only in what you record: a missing header is a new subscriber, an unresumable id is a window miss, and conflating them in your metrics hides the signal that the window is too short.
- How can the receiving application tell that a gap happened at all?Only because you told it. The protocol has no gap signal: a resumed stream and a restarted one are the same bytes apart from what you choose to send. A distinctly named event, or a field inside the data, is the only way the difference travels — which is why the branch must write something rather than just streaming live.
- What actually bounds how far back you can promise to resume?The retention of the store your ids address, and your willingness to serve a seek into its far end during a reconnect burst. Advertise the smaller of the two. An in-memory ring that outlives nothing bounds it to one process; a durable log bounds it to that log's retention.
saying these in an interview costs you the question
- Streams live events and assumes the client will notice the gap.
- Clamps an out-of-range id to the oldest entry and replays from there.
- Thinks the protocol defines an error for an expired resume point.
- Treats the replay window as a tunable independent of what is retained.
- Refuses the request, causing the client to loop on the same stale id.
- Assumes replay cost is bounded regardless of how long the client was away.