skip to content

A framed stream starts yielding garbage messages hours into a connection — how do you detect and recover from framing desynchronisation?

level: seniorimportance: nice to knowfreq 29%

answer

  1. the boundary, not the bytes, is wrong
  2. one connection fails, others do not
  3. any four bytes look like a length
  4. marker and checksum catch the shift
  5. reconnect resets the framing state

basics

~20 s

Detect it with per-frame invariants — a fixed marker at a known offset, a checksum, an implausible declared length — and recover by rescanning to the next marker, or by dropping the connection, since framing state is per connection.

solid answer

~50 s

Desynchronisation means the reader's idea of where a message starts is offset from the writer's, so every subsequent slice is cut from the wrong place. It usually comes from a writer whose declared length disagrees with the bytes it wrote, two writers interleaving on one connection, a missed escape, or a half-written frame after a crash. A pure length-prefixed stream rarely recovers on its own, because whatever bytes sit at the offset read as a plausible length and the error propagates. The defences are per-frame invariants: a **fixed marker** at the start of every frame, a **checksum** over the payload, and a known **frame type**. Any of them failing means stop — then either rescan forward to the next marker and resume, or, most often, tear down the connection and reconnect, since framing state lives entirely in that connection.

go deeper

for a junior

Take away the core idea: if a reader is wrong about where one message ends, everything after it is cut in the wrong place, not just the message that failed.

for a middle

Explain the invariants that make a shift detectable — a fixed marker at a known offset, a declared length, a checksum over the payload — and why a bare length prefix offers nothing to search for.

for a senior

Describe the diagnosis you would actually run: the per-connection signature, whether a reconnect clears it, what you capture before discarding, and why connection teardown beats an in-place resync.

for a principal

Decide what every frame in the estate must carry and what a framing error costs: a few bytes per message against silent corruption, and whether framing failures page someone or are quietly retried.

## What desynchronisation actually is A framed reader holds one crucial piece of state: the offset at which the next message begins. Desynchronisation is that offset being wrong. The consequences are not local — once the boundary shifts by k bytes, every later message is sliced from the wrong place, so the connection produces an unbounded run of decode failures, or, far worse, values that decode successfully and are wrong. The symptom in production is distinctive: a connection that worked for hours begins failing continuously, restarting the client fixes it immediately, and other connections to the same peer are unaffected. That per-connection signature is the tell. ## How a reader loses the boundary - **The writer's declared length disagrees with the bytes it wrote** — the most common cause, usually from a body computed twice, or one measured in characters and written in bytes. - **Two writers share one connection without serialising the write path**, so their frames interleave and neither is intact. - **A missed escape in delimiter framing** — an unescaped reserved byte ends a frame early, and the remainder becomes the head of a phantom message. - **A crash mid-frame**, leaving a partial frame that the next reader treats as a whole one. - **A reader that consumes a different count than it declared** — an off-by-one in the slice is enough. ## Why a length-prefixed stream does not self-heal A length prefix has no distinguishable form. Read four arbitrary bytes and you get a number; it looks exactly like a length. The reader confidently skips that many bytes, lands somewhere else arbitrary, and repeats. There is nothing to search for, which is the single structural weakness of pure length framing and the reason production protocols add something recognisable. ## Detecting it | Signal | What it usually means | |---|---| | Marker absent at the expected offset | The boundary is wrong right now | | Checksum fails over the declared payload | Wrong bytes, wrong boundary, or corruption | | Declared length wildly implausible | The count was read from payload bytes | | Unknown frame type on a stable protocol | The type byte came from the middle of a body | | Decode failures clustered on one connection | Desync rather than a bad producer version | The invariants matter more than any single one of them: a marker tells you the boundary is wrong, a checksum tells you the bytes are wrong, and together they catch both the shifted case and the corrupted case. **None of them can be checked if the frame carries no redundancy at all** — which is the argument for spending a few bytes per frame on it. ## Recovering 1. **Stop immediately.** Never hand a frame that failed its invariant to a decoder or to business logic; a garbage message that parses is worse than one that fails. 2. **Resynchronise if the format allows it.** With a fixed marker, scan forward byte by byte for the next occurrence, verify the frame that follows it against its length and checksum, and resume there. Accept that a marker value can also occur inside a payload, so the confirmation step is what makes the resync trustworthy. 3. **Otherwise drop the connection.** Framing state is per connection, so closing and reconnecting resets the offset to a known good zero. This is the honest answer for most systems, and it is why framing errors are usually modelled as fatal to the connection rather than to the message. 4. **Preserve evidence.** Log the offset, the declared length, the surrounding bytes and the producer identity before you discard, or the cause is unrecoverable after the fact. ## Preventing it - Give every frame a **marker, a frame type and a checksum**, so any shift is caught at the next frame rather than fifty frames later. - **Assert on the writer** that the bytes emitted equal the declared length, and fail the frame rather than send a lie. - **Serialise writes** to one connection, or give each writer its own. - **Test with adversarial chunking**: feed the same stream split at every byte offset, and include a truncated final frame, and assert the reader reports rather than delivers. - Fail **loudly** — a framing error is a correctness incident, not a warning to be sampled. ## Why it is a differentiator Most engineers have never seen this because a hardened protocol implementation hides it. The ones who have can name the per-connection signature, know why a bare length prefix cannot be searched, and reach for connection teardown rather than a clever partial recovery. That combination is what the question is really testing.

  • Why prefer tearing down the connection over trying to recover the boundary in place?
    Because framing state lives entirely in the connection, so a reconnect restores a known good offset with no guesswork, while an in-place resync can land on a marker value that happened to occur inside a payload. Resync is worth it only when reconnecting is expensive and the frame carries enough redundancy — a marker plus a length plus a checksum — to confirm the new boundary before trusting it.
  • A frame fails its checksum. Why is delivering it anyway with a warning the wrong call?
    Because the failure means you do not know what these bytes are: they may be a corrupted message or a slice from the wrong offset. Anything built on them is fabricated, and a warning that is sampled or ignored turns a detectable framing incident into silent data corruption downstream.
  • How do you tell desynchronisation apart from a producer that has started emitting a bad payload?
    Look at the blast radius and the invariants. A bad producer fails specific messages across many connections while frame markers and lengths still check out; desync fails everything after a point on one connection and breaks the frame-level invariants. A reconnect fixing it instantly confirms desync.

saying these in an interview costs you the question

  • Claims a length-prefixed stream can always resynchronise
  • Delivers a frame that failed its checksum with a warning
  • Treats the symptom as payload corruption, not a boundary shift
  • Scans for a marker and resumes without confirming the frame
  • Lets several threads write frames to one connection
  • Restarts the consumer without capturing the surrounding bytes