Can a Server-Sent Events stream be delivered in an encoding other than UTF-8, and how would a client know?
answer
- one encoding, no negotiation
- nothing in the grammar announces it
- always encoded and decoded the same way
- spelled UTF-8, and only that
- bad bytes become replacement characters
basics
~20 sNo. An event stream is always encoded and decoded as UTF-8, and the grammar offers no way to announce or select another encoding. A client never inspects anything to decide - it decodes as UTF-8 unconditionally.
solid answer
~40 sThe encoding is fixed at both ends and is not negotiable. An emitter must produce `UTF-8`, and a conforming client decodes the body as `UTF-8` whatever else the response says; there is no field in the grammar and no parameter a client consults to pick something else. That removes a whole class of ambiguity - no sniffing, no per-stream encoding state, no mismatch between what the emitter meant and what the client guessed. It also means a mis-encoded payload is not an error: bytes that are not valid `UTF-8` are replaced during decoding, so the block still dispatches and the fault text simply arrives corrupted, with nothing anywhere reporting a failure.
go deeper
Remember one fact: an event stream is always UTF-8. There is nothing to configure, nothing to negotiate and nothing to check before you start reading the body.
Explain why the grammar fixes it: parsing is defined over characters, so the bytes must be characters before a colon or a blank line means anything. Note that binary payloads must be encoded as text.
Recognise the quiet failure - structurally perfect events with mangled text, no error reported - and the incremental-decoding bug where a multi-byte character straddles a read and corrupts differently on each run.
Appreciate the trade the grammar makes: fixing one encoding removes sniffing, negotiation and per-stream state, and buys a wire you can read by eye, at the cost of pushing every binary payload through a text encoding.
## One encoding, fixed at both ends Most text formats carried over HTTP let the sender declare a character encoding and make the receiver honour it. Event streams do not. The rule is short: the body is encoded as `UTF-8`, and it is decoded as `UTF-8`. There is no field in the grammar that names an encoding, nothing in a block that can change it, and no step where a client chooses one. An emitter's only duty is to produce `UTF-8`; a client's only duty is to assume it. This is a deliberate simplification rather than an oversight. The grammar is defined over characters - a colon separates a field name from a value, a line feed ends a line, a blank line dispatches - so the parser cannot begin until the bytes are characters. Allowing the encoding to vary would mean either deciding it before parsing starts, or changing it mid-stream, and both are worse than fixing it. ## What that rules out - **No negotiation.** A client cannot ask for another encoding and a server cannot offer one. - **No declaration inside the stream.** No field, comment or first-block convention selects an encoding. - **No sniffing.** The client does not inspect the opening bytes to work out what it is reading. - **No per-stream state.** The decoder is the same for every event stream, which is one reason a conforming client is small. ## What happens when an emitter gets it wrong This is the part worth knowing, because the failure is quiet. If the emitting side writes text in some other encoding - typically by handing the response writer bytes that were already encoded elsewhere - the client still decodes as `UTF-8`. Byte sequences that are not valid `UTF-8` are substituted with a replacement character rather than aborting anything. The consequences, in order: 1. The block still parses, because the structural characters - colon, line endings, the blank line - are ASCII and survive most mangling. 2. The event still dispatches, with the correct type and the right number of payload lines. 3. The payload arrives corrupted: accented names, non-Latin fault text and symbols come through as replacement characters. 4. Nothing reports an error. There is no decode failure to catch and no signal on the stream, so this reaches users rather than logs. The diagnosis is therefore a content diagnosis, not a protocol one: structurally perfect events whose text is mangled in exactly the places where it stopped being ASCII. ## Characters, not bytes Two practical consequences follow from the grammar being character-oriented: - **A payload's byte length is not its character length.** A car label or a fault message containing non-ASCII text takes more bytes than characters, and nothing in the block declares a length anyway - lines and the blank line do all the framing, so there is no length field to get wrong. - **A multi-byte character can straddle a network read.** A client reading the body incrementally must decode incrementally too, holding a partial sequence until the rest arrives, rather than decoding each chunk independently. Decoding chunk by chunk produces replacement characters at chunk boundaries that vary run to run - a genuinely confusing bug, because the same stream fails differently each time. ## Carrying things that are not text Since the payload is text in a fixed encoding, anything binary must be turned into text before it goes into a `data` value - a base64 encoding of a captured sensor reading, for instance. Two structural characters make this non-optional rather than stylistic: a line feed inside the value would end the line, and the value cannot contain a blank line at all, because a blank line dispatches. The upside of the whole arrangement is the one the grammar keeps buying: a stream you can read with your eyes. An event stream is text in one known encoding, so a fault report can be inspected byte for byte on the wire and a hand-written line will parse exactly as a generated one does.
- What happens to bytes in the body that are not valid UTF-8?They are substituted with a replacement character as the body is decoded, rather than failing the stream. The block still parses and still dispatches, so the event arrives with corrupted text and no error anywhere - which is why this shows up as a user report, not a log line.
- How do you carry binary data, such as a captured sensor reading, in a payload?Encode it as text first, for example with base64. The payload is decoded as UTF-8 text and the grammar reserves line feeds and the blank line for structure, so raw bytes cannot go into a `data` value at all.
- Does the emitter need to announce the encoding anywhere?No - there is nothing to announce and nothing that would read it. Both ends are fixed on UTF-8, so the emitter's entire obligation is to produce it, and a client that received something else has no mechanism to find out.
saying these in an interview costs you the question
- Thinks a charset parameter can select another encoding for the stream
- Believes the client sniffs the opening bytes to detect the encoding
- Assumes a payload may carry raw bytes in any encoding
- Thinks mis-encoded text raises an error on the stream
- Counts payload bytes as if every character were one byte