What makes the `id:` you attach to each Server-Sent Events event actually usable for resuming a stream?
answer
- an address, not a name
- what came after this value?
- must survive a restart and another instance
- one value cannot describe twelve positions
- a cursor into the store you replay from
basics
~20 sAn id is resumable only if it addresses the store the server replays from: assigned by the server, ordered so that everything-after-it is computable, stable across instances and restarts, and short enough to ride back in a request header on every reconnect.
solid answer
~40 sThe grammar will accept any opaque string, so the constraint is not syntactic — it is that the value must answer one question later: *what came after this?* That requires four properties. It is **server-assigned**, because only the emitter knows the order. It is **ordered against the store you replay from**, so a cursor, an offset or a committed sequence rather than a bare random identifier. It is **stable beyond one process**, or a reconnect that lands elsewhere resumes into a different position while looking successful. And it is **small and printable**, because it travels in a request header field on every reconnect from every client. A wall-clock timestamp is the usual near-miss: ties and clock adjustment make it ambiguous exactly when two readings arrive together.
code
http · 7 linesid: 20481
event: level
data: {"gauge":"upper-weir","seq":7,"metres":3.41}
id: 20482
event: level
data: {"gauge":"mill-bridge","seq":12,"metres":1.08}go deeper
The id is chosen by the server, not the client, and it needs to mean something the server can look up later — a position, not just a unique-looking string.
Explain why order matters: resumption is the set of events after a value, so a scheme with no order cannot answer the question unless the store indexes it.
Bring up the failures that only appear in production: a counter that restarts, an instance that resumes into a different position, a merged feed whose id describes one source.
Frame it as a contract on storage: the id scheme fixes what must be retained and enumerable, so it constrains the replay window and the write path for the life of the feed.
## An id is an address, not a name It is easy to read `id:` as a label for an event, in the way a message might have a message id. It is not doing that job. The only use the value ever has is that a client hands it back later and the server must turn it into the set of events that followed it. So the test for a candidate scheme is a single question: **given only this value, can the emitter enumerate what came next, on any instance, for as long as the window promises?** ## Four properties, and how schemes fail them | scheme | ordered | resolvable after a restart | resolvable on another instance | verdict | |---|---|---|---|---| | store cursor or committed offset | yes | yes | yes | what to use | | per-process counter | yes | no | no | resumes into the wrong position | | random identifier per event | no | only as a lookup key | only if the store indexes it | needs a store that orders it | | wall-clock timestamp | approximately | yes | yes | ambiguous on ties and clock changes | | a hash of the payload | no | no | no | not an address at all | The per-process counter deserves the most suspicion, because it works perfectly in every test that runs one instance. A catchment feed served by several processes gives out `41`, `42`, `43` from each of them; a client that resumes from `42` against a different process is handed events that came after a completely different `42`, with no error and no visible symptom. Random identifiers are not disqualified, but they are only half a scheme: they carry no order themselves, so the store must be able to locate the entry and walk forward from it. If your replay path is "find this id, then read the next N", a random id is fine and has the pleasant property of leaking nothing about volume. If your replay path is "read everything greater than this value", it is unusable. ## The merged-feed trap The other trap is specific to feeds that multiplex. A district console reads one response carrying twelve gauges. It is natural for each gauge to have its own sequence number — and equally natural to write that number into `id:`. Then the ids on the wire read `7`, `12`, `7`, `8`, `13`, and a reconnect returns exactly one of them. The client sends back **one** value. It cannot describe twelve positions. So the id must address the merged order that this response is emitting, not any one source's position within it. The choices are: 1. **Assign a merged sequence** at the point where the streams are combined, and keep it durably so a restart continues it. 2. **Encode all the positions** into one opaque value, which the client will faithfully return in a header field on every reconnect — workable, but the value grows with the number of sources. 3. **Split the feed**, one response per source, so each has a one-dimensional id of its own. This trades header size for connection count. Choosing (1) has a consequence worth stating: the merged sequence is the thing you must retain, because it is what your window is measured in. ## What the header field allows Two practical limits come from the value's return trip. It is carried as a request header field encoded as UTF-8, and it cannot contain a line break — a line ending is what ended the field on the way out, so that shape never existed. Beyond that it is opaque, and the constraint that bites is size: it is sent on every reconnect by every client, and reconnects arrive in bursts when a shared link drops. Two more habits are worth adopting: - **Do not put anything secret or personal in it.** It travels in request headers and will appear in access logs along the path. - **Do not let it be guessable as an authorisation.** Possession of an id is not permission to read from that position; the request must still be authorised on its own terms. ## The one-line check Before shipping a scheme, hand a colleague one id value and a cold instance of the service, and ask them to produce the events that followed it. If they need to know which process emitted it, when it was emitted, or which of twelve gauges it belonged to, the id is not an address yet.
- Is a random identifier per event acceptable as an id?Only if your replay path can locate that entry and walk forward from it. A random value carries no order of its own, so a store that answers queries as everything-greater-than cannot use it. Where the store indexes it, a random id works and has the side benefit of revealing nothing about event volume.
- Why is a wall-clock timestamp a poor choice even though it is ordered and portable?Because ties are common and ordering under them is undefined: two gauge readings committed in the same millisecond cannot be separated, so a resume either repeats one or skips one. Clock adjustment makes it worse by allowing the sequence to move backwards, which no cursor should ever do.
- Can an id double as proof that the client is allowed to read from that position?No. It is an address, not a capability, and it travels through request logs along the path. Authorise the reconnect request on its own terms and treat the id purely as a position, or a client that guesses a neighbouring value reads a stream it was never granted.
saying these in an interview costs you the question
- Uses a per-process counter and only ever tests one instance.
- Puts a single source's sequence number in a merged feed's id.
- Assumes any unique value is enough because the grammar accepts it.
- Treats a wall-clock timestamp as a total order across events.
- Lets the client choose or derive the id value itself.
- Treats possession of an id as authorisation to replay from it.