skip to content

A sensor's JSON telemetry record is re-encoded in MessagePack or CBOR: what changes, and what stays the same?

level: middleimportance: must knowfreq 62%

answer

  1. same model, new notation
  2. tags and counts, not punctuation
  3. numbers stop being decimal text
  4. byte strings need no Base64
  5. keys still travel per record

basics

~20 s

The data model survives; only the syntax is replaced. Maps, arrays, strings, numbers, booleans and null become tagged bytes instead of punctuation and decimal text, so records shrink and decode cheaper — but every key name still travels on every record.

solid answer

~50 s

MessagePack and CBOR keep the document model of a text encoding — maps with keys, arrays, strings, numbers, booleans, null — and throw away the *syntax* that expressed it. Braces, quotes, colons and commas are replaced by a **type tag** in front of each value, and lengths or element counts replace closing delimiters. Numbers stop being decimal text and become native integers or floating-point values, strings carry a length instead of needing escape sequences, and a **byte-string type** appears that text encodings lack, so raw bytes no longer need Base64. What does *not* change is just as important: the record is still self-describing, so the field names are encoded in full on every record, and nothing about the bytes gives you a contract, validation or an evolution guarantee. You get a smaller, faster-to-parse version of the same document, not a schema.

go deeper

for a junior

Recall the one-line shape: same document model, binary notation. Maps, arrays, strings, numbers, booleans and null all survive; braces and quotes do not.

for a middle

Explain the mechanics: a type tag in front of every value, lengths and element counts instead of delimiters, native numbers, and a byte-string type that text encodings lack.

for a senior

Be honest about the ceiling in front of an interviewer: the saving is bounded by repeated key names, and batching plus compression closes much of the gap you were selling.

for a principal

Frame the choice as what you are buying — representation, not contract — and say which operational problem that leaves untouched before anyone approves a fleet-wide re-encoding.

## What "schemaless binary" means A **schemaless binary encoding** keeps the data model of a text document encoding and replaces only the notation. The model is the familiar one: maps (key/value), arrays, text strings, numbers, booleans and null. The notation is new: instead of writing `{`, `"`, `:`, `,` and decimal digits, the encoder writes a **type tag** — usually a single byte, often with a small length or a small value packed into the same byte — followed by the payload. `MessagePack`, `CBOR` and `BSON` are the members of this family usually named in an interview. "Schemaless" is the market's word; the accurate word is **self-describing**. Nothing in the bytes points at an external declaration. Hand a decoder the bytes and it can reconstruct the whole document — every key, every value and every type — with no other input. That property is the entire point of the family, and it is also the source of its ceiling. ## What changes on the wire - **Punctuation vanishes.** Structure is carried by the tag plus a count: a map header says "this is a map of *n* pairs", an array header says "*n* elements". There is no closing brace to look for. - **Numbers stop being text.** `-1234567` is written as an integer of an appropriate width rather than eight characters, and a fractional value is written as an IEEE 754 float. Parsing a number becomes a load rather than a decimal scan. - **Strings carry a length.** The bytes stay UTF-8, but nothing inside them has to be escaped, because the decoder is told how many bytes to take rather than hunting for a closing quote. - **Raw bytes get their own type.** A **byte string** is distinct from a text string. Binary payloads — a sensor's packed readings, a signature, a compressed blob — travel as themselves instead of being Base64-wrapped, which costs four output bytes per three input bytes, about a third more. - **Some members add a document length.** One member of the family was designed as a database's document representation, and its documents begin with a total byte length, which lets a reader step over a whole nested document in one jump. ## What stays the same | Aspect | Text document encoding | Schemaless binary encoding | |---|---|---| | Data model | maps, arrays, strings, numbers, booleans, null | the same, plus a byte-string type | | Field names | in every record | in every record | | External declaration needed to decode | none | none | | Value types known from the bytes alone | by syntax | by inline type tag | | Human-readable on the link | yes | no — needs a decoder | | Contract or validation | none inherent | none inherent | The first two rows are the ones candidates skip. Because the model is preserved, the migration is mostly mechanical — the same object graph goes in and comes out. Because the field names are preserved, the size win is real but **bounded**: on a small record whose keys are as long as its values, the keys are the bulk of the payload and they are unchanged. ## Why a metered link makes this attractive Picture battery-powered field sensors uplinking telemetry over a metered radio link, with no way to push a new build to the whole fleet at once. Every byte is airtime, and airtime is battery. Dropping quotes, colons, commas and decimal text off each record shaves a genuine fraction of the frame, and decoding gets cheaper on the gateway because there are no digits to parse and no escapes to unwind. Crucially, none of this requires the fleet to agree on anything new: a device that was emitting the same document yesterday can emit the binary form today, and a reader that understands the family can read it without knowing which firmware wrote it. ## Where the honesty lives Two caveats belong in any good answer. First, the saving is usually a fraction, not an order of magnitude — if you want the field names gone you are asking for a schema-driven encoding, which is a different family with a different operational cost. Second, if you are already batching records and running a general-purpose compressor over them, repeated key names compress extremely well, so much of the raw gap closes; the binary form's remaining advantages are then decode cost, typed numbers and native byte strings rather than size alone. ## How to say it in an interview "It is the same document with a cheaper skin." Then name the three things that genuinely improve — typed numbers, length-prefixed strings, native byte strings — and the one thing that does not: the keys still ride along, on every single record, forever.

  • If the field names still travel, where does the size saving actually come from?
    From the syntax and the value representations: no quotes, colons, commas or braces; a number written as a fixed-width or short integer rather than its decimal digits; a string given a length instead of escapes; and raw bytes carried directly instead of Base64, which otherwise costs four bytes for every three. On a record with long values the saving is noticeable; on a record that is mostly short keys it is modest.
  • Does moving to one of these encodings give you any evolution guarantee?
    No. The bytes describe themselves, so a decoder can always read a record it has never seen, but nothing tells it what a new field means or that a removed field used to exist. Whatever your readers did with unknown or missing keys in the text form, they still have to do — the encoding change buys representation, not contract.
  • Why does a general-purpose compressor weaken the size argument?
    Because the largest remaining redundancy in a self-describing record is the key names, repeated identically in every record of a batch, and that is exactly what a dictionary-based compressor such as DEFLATE or Zstandard removes best. Over a compressed batch the binary form's edge shrinks toward its decode-cost and type-fidelity advantages rather than its byte count.

It is the difference between reading a form out loud and handing over the filled-in form: the questions are still printed on every copy, you have just stopped spelling out the punctuation.

saying these in an interview costs you the question

  • Claiming the encoding removes field names from the record
  • Saying these encodings need a schema before a reader can decode
  • Promising an order-of-magnitude size drop from the swap alone
  • Treating the change as an evolution or compatibility guarantee
  • Assuming numbers are still decimal text underneath the tags