skip to content

questions

5

Why does framing by a delimiter byte force the writer to escape or re-encode the payload?

level: middleimportance: must knowfreq 54%

answer

  1. the delimiter is not available to payloads
  2. reserving a byte costs the alphabet
  3. escape it, or restrict the alphabet
  4. stuffing doubles a hostile payload
  5. the prefix needs the size up front

basics

~20 s

A delimiter means 'the message ends here', so any occurrence of that byte inside the payload would end the frame early. The writer must escape it, or restrict the payload to an alphabet that excludes it.

solid answer

~40 s

Delimiter framing reserves a byte value to mean end of message, which takes that value away from the payload. Over arbitrary binary data every value occurs, so the writer must either **escape** the delimiter (and the escape byte itself) before sending — byte stuffing — or re-encode the payload into a restricted alphabet that cannot contain it. Both cost size and a pass over the bytes: stuffing can double a hostile payload in the worst case, and a restricted-alphabet re-encoding of three bytes into four characters adds about a third. A length prefix has no such problem, because the count is outside the payload and the payload stays opaque. The prefix pays elsewhere: the writer must know the size before it sends the first byte.

code

pseudocode · 24 lines
pseudocode
// writer: stuff, then terminate
function frame(payload):
    out = empty
    for each b in payload:
        if b == DELIM or b == ESC:
            append ESC to out
        append b to out
    append DELIM to out
    return out

// reader: track escape state while scanning
escaped = false
message = empty
for each b in incoming_bytes:
    if escaped:
        append b to message          // literal DELIM or ESC
        escaped = false
    else if b == ESC:
        escaped = true
    else if b == DELIM:
        deliver(message)
        message = empty
    else:
        append b to message

go deeper

for a junior

Hold on to the core fact: a byte reserved to mean end of message can no longer appear freely inside a message, so something has to be done to payloads that contain it.

for a middle

Explain both escapes and their costs — stuffing with the escape byte itself escaped, or re-encoding into a restricted alphabet — and state the matching cost of a length prefix: the size must be known before the first byte is sent.

for a senior

Argue a choice for a concrete stream: payload shape, message sizes, whether an operator needs to find boundaries in a capture, and how you would carry a body whose length is not known in advance.

for a principal

Treat the framing as the contract hardest to change: the prefix width caps message size forever, and switching families later means every producer and consumer cuts over together.

## Two ways to say where a message ends A framing convention has one job: tell the reader where the current message stops. There are two families. - **Length prefix** — a fixed-width count precedes the body. The reader consumes the count, then exactly that many bytes. The payload is never inspected. - **Delimiter** — a reserved byte or sequence terminates the body. The reader scans forward until it finds the delimiter. A **fixed header** (a marker, a format identifier, flags and a length) is the length-prefix family with extra fields, and is what most binary protocols actually use. ## Why the delimiter reaches into the payload Reserving a byte value for a boundary removes it from the payload's usable alphabet. If the payload can be arbitrary bytes, every value — including the reserved one — will eventually appear, and the reader would cut the message short at the first one. Two escapes exist: 1. **Byte stuffing.** Before sending, the writer scans the payload and prefixes every occurrence of the delimiter, and of the escape byte itself, with the escape byte. The reader reverses it: an escape byte means *take the next byte literally*, and only an unescaped delimiter ends the frame. Correctness depends on escaping the escape — otherwise a payload ending in a literal escape byte swallows the delimiter. 2. **Restricted alphabet.** The writer re-encodes the payload into a character set that excludes the delimiter. Encoding three bytes into four characters from a 64-symbol alphabet is the familiar instance and costs about **33%** in size, plus the encode and decode passes. Stuffing's worst case is a payload consisting entirely of delimiter or escape bytes: every byte becomes two, so the frame is about **twice** the payload plus the terminator. That worst case is reachable by an adversary, which matters when sizing anything downstream. ## What the length prefix costs instead The prefix keeps the payload opaque — no scan, no transformation, any byte sequence passes through untouched — but it moves the cost to the writer: - **The size must be known before the first byte goes out.** A body being produced incrementally must be buffered in full, or split into a sequence of length-prefixed chunks ending with a zero-length chunk to mean 'done'. - **The prefix width caps the message.** An unsigned 16-bit prefix tops out at **65,535** bytes; an unsigned 32-bit prefix at **4,294,967,295** bytes, just under 4 GiB. Widening later is a breaking change to the framing, not to the schema. - **The boundary is not findable by inspection.** Any bytes at a given offset read as a plausible length, so a reader that loses the boundary cannot find it again by scanning. ## Side by side | Property | Length prefix | Delimiter | |---|---|---| | Writer must know the size up front | Yes | No | | Payload must be transformed | No | Yes — escaped or re-encoded | | Per-message overhead | Fixed prefix width | One byte, plus every escape | | Worst-case size growth | None | About 2x (stuffing) or ~33% (restricted alphabet) | | Reader work per byte | None — skip by count | A comparison on every byte | | Boundary recoverable by rescanning | No | Yes — scan to the next delimiter | | Natural fit | Arbitrary binary payloads | Payloads already in a restricted alphabet | ## Choosing one - Binary payloads, large messages, throughput that matters: **length prefix**, because skipping by count beats comparing every byte and nothing touches the payload. - Payloads already restricted to a text alphabet, where a human or a simple tool should be able to find boundaries in a captured stream: **delimiter**, and its rescannability is a real operational asset. - Bodies of unknown length: **chunked length-prefixed framing** — a run of sized chunks with a zero-length terminator — which keeps the payload opaque without requiring the total up front. - Long-lived binary protocols in practice: a **fixed header** with a marker, a format identifier and a length, which buys the prefix's opacity and, through the marker, some of the delimiter's recoverability. ## The trap to name The escape rule must be applied to the escape byte itself and only to unescaped delimiters. A reader that unescapes the whole buffer first and then splits on delimiters, rather than tracking escape state as it scans, will split inside a payload that contained an escaped delimiter — a bug that lies dormant until the first payload with that byte in it.

  • The writer must send a body whose total size it does not know in advance. How do you keep a length-prefixed framing?
    Split the body into a sequence of length-prefixed chunks and end it with a zero-length chunk. Each chunk carries a size the writer does know, the reader concatenates until it sees the terminator, and the payload is still never inspected. The cost is one prefix per chunk and a reassembly step.
  • Why is a delimiter-framed stream easier to recover after a reader loses the message boundary?
    Because the boundary is a distinguishable byte value that can be searched for. A reader that finds itself confused can scan forward to the next unescaped delimiter and resume at a real boundary, losing one message. With a pure length prefix there is nothing to search for: arbitrary bytes read as a plausible length.
  • What breaks if the escape byte itself is not escaped?
    A payload whose last byte is the escape byte turns the following delimiter into a literal, so the frame never terminates and the next message is absorbed into this one. The rule is symmetric by necessity: both reserved values — the delimiter and the escape — must be escaped inside the payload.

saying these in an interview costs you the question

  • Thinks escaping is only needed for text payloads
  • Escapes the delimiter but not the escape byte
  • Unescapes the whole buffer before splitting on delimiters
  • Claims a length prefix costs nothing on the writer
  • Assumes the chosen delimiter simply never appears in payloads
  • Believes widening a length prefix is a compatible change
open as a page

Framing turns a byte stream into messages: on a connection whose reads never align with message boundaries, what must the reader's framing loop do?

level: middleimportance: must knowfreq 66%

basics

~20 s

A stream connection carries bytes, not messages, so the reader keeps a buffer across reads: append each chunk, test whether a complete message is present from its length prefix or delimiter, consume exactly that message, and retain the remainder.

open as a page

What belongs in a message envelope that an intermediary must read without decoding the payload it wraps?

level: seniorimportance: should knowfreq 49%

basics

~20 s

Whatever a hop needs in order to route, dispatch, drop or trace a message it will never decode: a message identifier, a content type and schema identifier, correlation and causation metadata, a timestamp, the payload length and an integrity check.

open as a page

Designing a shared envelope for a long-lived message stream, how do you decide which facts belong outside the encoded payload?

level: principalimportance: should knowfreq 41%

basics

~20 s

Promote a fact to the envelope only when a participant that cannot decode the payload still needs it, and when it describes the transfer rather than the business event. Everything else stays in the body, with one authority per fact.

open as a page

A framed stream starts yielding garbage messages hours into a connection — how do you detect and recover from framing desynchronisation?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

Detect it with per-frame invariants — a fixed marker at a known offset, a checksum, an implausible declared length — and recover by rescanning to the next marker, or by dropping the connection, since framing state is per connection.

open as a page