Your text/event-stream feed emits a line with no colon and a field name the grammar never defines - how does a conforming client treat each?
answer
- no fatal parse error in the body
- no colon means empty value
- unknown field ignored, not passed through
- exactly one leading byte order mark stripped
- tolerance protects uptime, not correctness
basics
~20 sBoth are tolerated, not rejected. A non-empty line with no colon is processed as a field name with an empty value, and a field name the grammar does not define is ignored outright. Neither ends the stream.
solid answer
~50 sA Server-Sent Events client is deliberately forgiving about input it cannot make sense of. A non-empty line containing no `U+003A COLON (:)` is processed as though the whole line were a field name with the empty string as its value - so a bare `data` line contributes an empty value rather than being thrown away. A field name that is not one the grammar defines is ignored completely, along with its value. And the `UTF-8` decode that feeds the parser strips exactly one leading `U+FEFF BYTE ORDER MARK`; a second mark later in the stream is not stripped and simply ends up inside a field name, which then goes unrecognised and is ignored like any other. The important part is what does not happen: none of these cases raises an error, drops an event, or ends the stream. A conforming client has no notion of a fatal parse error in the body at all.
code
http · 7 linesHTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
platform
severity: minor
data: {"train":"IC 4712","status":"delayed 6 min"}go deeper
Remember the headline: a client tolerates lines the grammar does not define. It does not fail, it does not stop reading, and it does not tell anyone.
Be able to state each rule exactly: whole line as field name with an empty value, unknown field ignored, one leading byte order mark stripped by the decode.
Show that you know tolerance hides emitter defects, and that the only place to catch them is a byte-level assertion on the emitting side, because the receiver never raises one.
Weigh the extensibility the ignore rule buys against the silent-corruption surface it opens, and decide where in your systems a feed gets validated rather than passed through.
Streams in production are rarely clean. A feed may be produced by a much older system, re-emitted by a component nobody owns, or extended by a team that added a line of its own. The specification anticipates this: the parsing rules are written so that a client can meet input the grammar does not define and keep going. Knowing exactly which input is survivable is what separates someone who has read the rules from someone who has only used a client library. ## Tolerance is a rule, not a courtesy There is no fatal parse error in the body of an event stream. Whatever arrives, the client keeps reading and keeps dispatching whatever it could build. The three cases worth knowing: - **A non-empty line with no `U+003A COLON (:)`** - the whole line is processed as a field name, with the empty string as the field value. A line reading `data` behaves as `data:` with nothing after it. - **A field name the grammar does not define** - the field is ignored. `severity: minor` contributes nothing at all; it is not passed through, not exposed to the receiving application, and not an error. - **A leading `U+FEFF BYTE ORDER MARK`** - the `UTF-8` decode removes exactly one, and only at the very start of the stream. A mark appearing later is ordinary text: it becomes the first character of whatever field name follows it, which is then unrecognised and ignored. | input the client meets | what it does | what it never does | |---|---|---| | non-empty line, no colon | whole line is the field name, value is empty | discard the line, or end the stream | | unrecognised field name | ignores field and value | pass it to the receiving application | | one leading byte order mark | strips it during decode | strip a second one later | | a later byte order mark | leaves it in the text | report an encoding error | ## Why the specification chose this Two reasons, and both matter operationally. First, **forward compatibility**: because unknown fields are ignored rather than rejected, a server may emit a field that older clients have never heard of without breaking any of them. The extension point is free, and it is the ignore rule that pays for it. Second, **an unattended client**: the receiving side reconnects by itself and often runs for days with nobody watching it. A client that treated a stray line as fatal would take an entire fleet down over a formatting slip in one event. ## The cost of that generosity Silence cuts both ways. Because nothing complains, a mistake on the emitting side can run for months: - a misspelled field name - `date:` where `data:` was meant - is dropped without a sound, and the event arrives missing whatever that line carried; - a line intended as a comment but written without its leading colon becomes a field name with an empty value instead of being skipped; - a system that concatenates two streams may put a byte order mark in the middle of the combined output, where it silently corrupts the field name that follows it. None of these produces an error anywhere. The client is behaving exactly as specified, the receiving application sees an event that is merely incomplete, and the only place the defect is visible is in the bytes on the wire. ## What to do about it The check has to live on the side that emits, because the side that receives will never raise one. 1. Build the stream with one function that writes fields, rather than by string-concatenating lines at call sites. Field names then cannot be misspelled per call site. 2. Assert on the wire in tests: capture the raw response body of a streaming endpoint and compare bytes, including the blank line that ends each event. A test that asserts on parsed objects cannot see a field that was ignored. 3. When you bridge a feed from an upstream system whose output you do not control, validate it at the bridge and record what you dropped. Anything you pass through unexamined, the client will silently ignore for you. 4. Treat a byte order mark as a hazard whenever streams are produced by concatenation or by a tool that writes one by default. Exactly one, at the very start, is invisible; anywhere else it is corruption you will not be told about. The rule to carry out of this: the tolerance rules protect availability, not correctness. They guarantee a malformed line will not take the stream down. They guarantee nothing at all about the event arriving intact.
- A team adds a custom field to the feed and asks how clients will expose it. What do you tell them?Clients will not expose it at all - an unrecognised field name is ignored, along with its value, and never reaches the receiving application. Anything the application must read has to travel inside the `data` payload. The ignore rule makes adding the field harmless, which is exactly why it is also useless as a transport.
- Does any malformed input make a conforming client stop reading the stream?Not in the body. Reading stops when the response body ends or the connection breaks, and the client then reestablishes the connection; or when the response itself was unacceptable at open, which fails the connection outright. Bad lines inside an otherwise good stream are absorbed by the parsing rules and never terminate anything.
- Why would a byte order mark ever appear in the middle of a stream?Usually because the stream was assembled from more than one source - a bridge concatenating an upstream feed's output, or a tool that writes a mark at the start of everything it emits. Only one mark, at the very start, is removed by the decode; any later one stays in the text and corrupts the field name that follows it.
saying these in an interview costs you the question
- Says a malformed line ends the stream
- Expects unknown fields to reach the receiving application
- Thinks a line with no colon is simply discarded
- Believes byte order marks are stripped wherever they appear
- Treats client tolerance as proof the feed is correct