skip to content

How do MessagePack and CBOR carry a timestamp or a big integer that their base data model has no type for?

level: seniorimportance: nice to knowfreq 26%

answer

  1. base model is deliberately small
  2. tag number plus opaque payload
  3. length declared, so always skippable
  4. registered tags versus application range
  5. unknown tag: surface, skip or reject

basics

~20 s

Through an extension mechanism: a numbered tag in front of an opaque payload. The tag tells a decoder how to interpret the bytes, a decoder that knows the number produces a typed value, and one that does not can keep the pair intact rather than failing.

solid answer

~50 s

The base model of these encodings is the document model — maps, arrays, strings, numbers, booleans, null — which has no timestamp, no arbitrary-precision integer and, historically, no way to say "these bytes mean something specific". Both formats solve it the same way: an **extension** is a **tag number** plus a payload of bytes. Some tag numbers are registered and widely implemented, so a timestamp or a big integer decodes into a real typed value on most readers; others are left for applications to claim privately. The crucial property is what a decoder does with a tag number it does not recognise: because the payload's length is declared, it can surface the value as an opaque `(tag, bytes)` pair and re-emit it unchanged, rather than rejecting the record. Implementations differ on whether they do that or raise, which is exactly what to check before relying on a private tag across a fleet you cannot upgrade in step.

code

pseudocode · 10 lines
pseudocode
function decode_extension(bytes, pos):
    (tagNumber, payloadStart, payloadLen) = read_extension_header(bytes, pos)
    payload = bytes[payloadStart .. payloadStart + payloadLen]

    if known_tags contains tagNumber:
        value = known_tags[tagNumber].interpret(payload)   // typed value
    else:
        value = OpaqueExtension(tagNumber, copy of payload) // preserved, not an error

    return (value, payloadStart + payloadLen)              // end is known either way

go deeper

for a junior

Know that the base model stops at maps, arrays, strings, numbers, booleans and null, and that anything richer arrives under a numbered extension tag.

for a middle

Explain the shape — tag number, declared length, opaque payload — and why the declared length is what lets an unaware decoder keep going.

for a senior

Show the operational care: check what every decoder in the path does with an unknown tag before a device emits one, and prefer a relay that preserves it.

for a principal

Treat tag-number assignment as a permanent contract with no tooling behind it, and decide who owns the register before teams start claiming numbers.

## Why an extension mechanism is needed at all The data model these encodings inherited from a text document format is deliberately small: maps, arrays, text strings, numbers, booleans and null, plus the byte string the binary family adds. Real telemetry wants more than that: - a **point in time**, with an agreed epoch and resolution, rather than a number whose units live in a comment; - an **integer too large for 64 bits**, or an exact decimal that a floating-point value would round; - a **domain value** — an identifier, a coordinate pair, a unit-bearing reading — that a consumer should reconstruct as a type rather than a raw map. Without a mechanism, each team invents a convention ("the field named `ts` is epoch milliseconds"), and that convention lives nowhere the bytes can express. ## The mechanism: a tag number plus opaque bytes Both formats extend the model in the same shape. An extension value carries: 1. a **tag number**, which identifies the interpretation; 2. a **length**, so any decoder knows where the payload ends; 3. the **payload bytes**, whose meaning is defined entirely by the tag number. Some tag numbers are **registered** — published interpretations that most general-purpose decoders implement, including ones for points in time and for arbitrary-precision integers. Others are explicitly left for **application use**, which is where a private convention can live as a first-class value instead of as a comment. Because the length is declared, the mechanism is **structurally transparent**: a decoder can always find the end of an extension value, even with no idea what it means. That is what makes the next section possible. ## The behaviour that actually matters: an unknown tag The interesting question is not "how do I encode a timestamp" — it is what happens when a record reaches a reader that does not know the tag number, which on a long-lived fleet is the normal case. Three behaviours exist in the wild: - **Surface it opaquely.** The decoder returns the tag number and the raw payload bytes as a value. The record round-trips unchanged, and a relay can forward it without understanding it. This is what you want from a gateway. - **Skip it.** The decoder walks past the value using the declared length and drops it. Structurally safe, but the record no longer round-trips — a relay would silently strip the field. - **Reject.** The decoder treats an unknown tag as a decode error and fails the record. Implementations genuinely differ here, so if you plan to use a private tag on a fleet with mixed firmware, **test the behaviour of every decoder in the path before the first device ships it**, and prefer a relay that preserves unknown tags. | Decoder behaviour on an unknown tag | Record round-trips | Field survives a relay | Risk | |---|---|---|---| | Surface as tag plus bytes | yes | yes | consumer must handle an opaque value | | Skip using the declared length | yes, minus the field | no | silent data loss | | Reject the record | no | no | one new tag can drop all traffic | ## Using it well on a fleet you cannot upgrade in step 1. **Prefer a registered tag** where one exists for what you are carrying — a point in time, an arbitrary-precision integer. It costs nothing and buys you decoders you did not write. 2. **Claim a private tag from the application range**, not an arbitrary number, and write down what the payload means with the same seriousness you would give a schema. 3. **Do not change a tag's meaning.** The tag number *is* the contract; redefining it is the same defect as reusing a field identifier for a different meaning, and there is no declaration anywhere to catch it. 4. **Keep the payload self-contained.** A payload that only makes sense alongside another field in the same record is a convention with extra steps. 5. **Have a fallback.** If a consumer in the path rejects unknown tags and cannot be changed, carry the value in the base model as well during the transition, and drop the duplicate once the path is clean. ## The framing for an interview Extensions are how a self-describing encoding admits its model is incomplete without giving up self-description: the tag number travels *with* the value, so the record still explains itself, just at one remove. That is also the limit — a tag number tells a decoder how to read the bytes, never what the value means to your domain. Meaning stays where it always was, in the agreement between the people who wrote the producer and the consumer.

  • Why is the declared payload length the part that makes this mechanism safe?
    Because it separates structure from meaning. A decoder that has never heard of a tag number still knows exactly where the value ends, so it can keep parsing the rest of the record and can copy the value forward untouched. Without a declared length an unknown tag would leave the decoder unable to resynchronise, and the whole record would be lost.
  • What is the failure mode of reusing a private tag number for a new meaning?
    Older readers keep interpreting the payload under the old definition and produce wrong values without any error, because nothing on the wire or in any declaration can detect the change. It is the same defect as reusing a retired field identifier, with no tooling to catch it — so treat a tag number as permanently assigned once a device has shipped it.

saying these in an interview costs you the question

  • Assuming the base model already has a timestamp type
  • Believing every decoder rejects tag numbers it does not know
  • Picking an arbitrary tag number outside the application range
  • Redefining a tag number that shipped devices already emit
  • Expecting a tag number to convey domain meaning to a consumer