skip to content

In a WebSocket connection, what does the protocol guarantee about a text message that it does not guarantee about a binary one?

level: juniorimportance: must knowfreq 62%

answer

  1. one of the two types carries a promise
  2. about encoding, not about size
  3. the other type is opaque octets
  4. valid UTF-8 or the connection fails
  5. close status 1007 on bad text

basics

~20 s

A WebSocket text message must carry valid UTF-8, and an endpoint receiving invalid UTF-8 in one fails the connection with close status 1007. A binary message is an opaque sequence of octets the protocol never inspects.

solid answer

~40 s

A WebSocket delivers whole messages, and every message is declared either text or binary. That declaration has one enforceable consequence: a text message's payload is character data encoded as UTF-8, and an endpoint that receives a text message whose bytes are not valid UTF-8 fails the connection with close status `1007` rather than delivering or repairing it. A binary message is any sequence of octets, and the protocol validates nothing about it. Everything else is an application choice. Binary avoids the `base64` or hexadecimal step that already-binary data would need inside a text message; text stays readable in a capture. Sending a JSON document as binary saves nothing - it is the same bytes with a different declared type.

code

pseudocode · 11 lines
pseudocode
on message(m):
    if m is a text message:
        # payload is already valid UTF-8 characters
        command = parse json(m.text)
        handle command(command)
    else:
        # opaque octets; nothing has been validated for you
        if length of m.bytes < 8:
            reject as malformed
        else:
            decode telemetry record(m.bytes)

go deeper

for a junior

Recall the pair: text means the payload is valid UTF-8 character data, binary means opaque octets. Know that a text message with bad encoding costs the connection, not just the message.

for a middle

Explain where the check lands. The encoding rule applies to the reassembled message, so a fragment may legally split a multi-byte sequence, and the failure surfaces as a close with status 1007 rather than as a parse error.

for a senior

Show the operational consequence. One mis-encoded text message drops a connection and everything queued on it, so a service forwarding untrusted third-party strings should validate before sending or carry that payload as a binary message.

for a principal

The tradeoff is who owns validation. Text buys a free encoding check and a readable capture at the cost of an inflation step for binary data; binary buys raw bytes at the cost of a decoder before anyone can diagnose anything. Decide per payload, not per product.

## Two message types, and one difference the protocol enforces A WebSocket connection does not hand the application a byte stream. It hands it **messages**, and every message is declared as either a **text message** or a **binary message**. That declaration is the only typing the protocol has, and it carries exactly one enforceable consequence: the payload of a text message is character data encoded as **UTF-8**, while the payload of a binary message is an opaque sequence of octets the protocol never inspects. The obligation on text is real rather than advisory. An endpoint that receives a text message whose bytes are not valid UTF-8 does not pass it to the application and does not quietly substitute replacement characters. It fails the connection, closing with status **1007** - data inconsistent with the type of the message. A malformed byte in a text message is therefore not a parse error the application can catch and skip. It ends the connection, and everything else in flight on that connection goes with it. ## Validity is a property of the reassembled message A single message may arrive as several fragments, and a fragment boundary can fall anywhere - including in the middle of a multi-byte UTF-8 sequence. The encoding rule applies to the **complete** message, not to each piece of it. A receiver that validates incrementally has to carry a partial sequence across the boundary and resume it when the next fragment arrives; one that validates each fragment independently will reject perfectly legal messages whose only crime is being split at an awkward byte. The practical consequence is that a text message is not fully trustworthy until it is complete. If you hand fragments to a consumer as they arrive, the encoding failure that condemns the whole message can surface on the last fragment, after you have already acted on the first. ## What binary actually buys, and what it does not - **It removes an encoding step for data that is already octets.** An image tile, a compressed blob or a packed sensor record placed inside a text message must first become characters - `base64` inflates it by about a third, hexadecimal doubles it. Sent as a binary message it travels as-is. - **It saves nothing on a document that is already text.** A JSON envelope sent as a binary message is byte-for-byte what it was as a text message. Only the declared type changed, and the receiver now has to decode the octets itself. - **It costs readability.** A text message is legible in a capture and in a log; a binary one needs the matching decoder before anyone can see what went wrong. - **It moves all validation to you.** The protocol checks a text message's encoding. It checks nothing whatsoever about a binary payload, so a truncated or corrupt structure inside one is your problem to detect. - **It is not faster by itself.** The per-message overhead the protocol adds is the same for both types, and the type declaration does not change how a message is scheduled or delivered. ## The two side by side | | Text message | Binary message | |---|---|---| | Payload obligation | valid UTF-8, enforced by the receiver | none; any octets | | Violation | connection failed with close status `1007` | no such failure exists | | Who checks the content | protocol checks encoding, application checks meaning | the application, entirely | | Readable in a capture | yes | only through a decoder | | Natural fit | a JSON or line-shaped envelope | data that is already bytes | ## Choosing, in a real console A crane cabin console usually wants both on the same connection, and that is legal: the type travels with each message, so the console can send short JSON commands as text messages and receive camera tiles or packed telemetry as binary messages over the same socket. Pick the type from what the payload actually is: 1. If the payload is characters a human might read during an incident, send a text message and get the encoding check for free. 2. If the payload is already octets, send a binary message rather than paying a third or double the bytes to dress it up as characters. 3. If you are choosing binary purely to make a textual document smaller, stop - it will not be smaller, and you have given up readability and gained a decode step.

  • A peer sends a 40 KB JSON document as a binary message rather than a text message. What changes?
    The byte count does not change at all - the same octets travel either way. What changes is the declared type: the receiver no longer gets the UTF-8 guarantee, has to decode the octets to characters itself, and the connection will no longer be failed with `1007` if that document is mis-encoded. You have given up a free check and gained nothing.
  • Why is a mis-encoded text message not just an application-level parse error?
    Because the encoding rule belongs to the protocol, not the application. A receiving endpoint that finds invalid UTF-8 in a text message fails the connection with close status `1007` instead of delivering it, so one bad message takes the whole socket down along with anything else in flight on it. An application-level parser never sees it.
  • Can a fragment of a text message be checked for valid UTF-8 on its own?
    Not safely. Fragments may split at any byte, including inside a multi-byte UTF-8 sequence, so a fragment that looks invalid in isolation may be part of a perfectly legal message. The rule applies to the reassembled message, so an incremental validator has to carry a partial sequence across the boundary.

A text message is a crate with a declared contents label the dockside inspector actually checks; a binary message is a sealed container that passes through unopened. Only one of them can be refused at the gate for what is inside it.

saying these in an interview costs you the question

  • Claims a binary message is always smaller or faster than text
  • Thinks a text message may carry arbitrary bytes
  • Expects the receiver to repair invalid UTF-8 silently
  • Believes the protocol validates a binary payload's structure
  • Sends JSON as a binary message expecting it to shrink
  • Says the message type is only a hint to the application