skip to content

Data serialization

15 roadmaps133 questionsupdated

Turning in-memory values into bytes and back: what the wire carries, how it changes without breaking readers, and what decoding costs. It sits under every API and queue, so interviewers probe it.

on this pageshow

guide

overview

~1 min

Serialization turns a value that lives inside one running process into bytes that another process, a file or a queue can hold, and turns those bytes back into a value later. Nearly every API call, queued message, shared cache entry and archived event passes through it, which is why interviewers use it to test whether you understand what actually crosses a boundary. They want to hear which encoding you would choose and why, what happens when the contract changes while old and new code are both running, how a missing field differs from an empty one, and what the read side costs in bytes, processor time and risk. The subject splits into six sections. [The wire form model](/topics/found-serialization-fundamentals) explains why encoding is needed at all, the two splits (text or binary, self-describing or schema-driven) and the byte-level primitives underneath. [Format families](/topics/found-serialization-formats) places whole encodings on that map and matches them to workloads. [Schema evolution](/topics/found-serialization-schema-evolution) is about change over time: compatibility directions, safe and breaking edits, versioning, registries. [Wire-contract design](/topics/found-serialization-wire-contracts) covers the edge cases behind real outages: null against absent, fields a reader does not know, where one message ends, numbers and timestamps that lose fidelity, byte-stable encodings. [Throughput and footprint](/topics/found-serialization-performance) measures the cost, and [the decode-path attack surface](/topics/found-serialization-security) treats the parser as the first code that touches hostile input. Start with the wire form model, since every later section uses its vocabulary. Formats and schema evolution come next, then contracts, then performance and security. Juniors are usually asked what separates two encodings; seniors and principals are asked how a contract survives years of independent deploys, so expect the same idea at several depths.

primer

A few ideas carry the whole subject. Hold them and most questions below read as applications rather than facts. - **Encoding is translation, not copying.** A value in memory is shaped for the process that owns it: its addresses, padding and layout mean nothing anywhere else. The bytes on the wire must stand on their own, and the reader builds a new value from them. Success is judged by what the reader can observe, not by whether two memory images match. - **Two axes place almost any encoding.** One is readability: characters a person can inspect against packed bytes a machine reads quickly. The other is where the structure lives: carried inside the bytes beside every value, or held in a schema both sides must agree on in advance. A third axis, record-shaped against column-shaped, matters once the reader scans many records for a few fields. - **Field identity is the real contract.** Whether fields are matched by name, by number or by position decides which edits are harmless and which silently break someone. Most schema-change questions reduce to asking what identifies a field on the wire. - **Readers and writers rarely upgrade together.** In most systems with more than one deployable, old and new versions of a contract run side by side for hours, days or years. Compatibility is therefore directional, and every design choice is judged by what happens when mismatched versions meet. - **Presence carries meaning.** Absent, explicitly empty, set to a default, and not understood by this reader are four different states. Encodings differ in which of them they can express, and a contract that collapses two of them eventually corrupts data. - **The read side is where cost and risk collect.** The writer knows its own data; the reader receives bytes it did not produce and cannot trust. Decoding usually does more work, and it runs before any application-level validation has had a chance to reject anything. - **The dangerous failures are quiet.** Precision that rounds away, fields that vanish on re-emit, a rename that stops a value arriving: none of these raises an error. Interviewers reward naming the silent case.

Marshalling
Converting an in-memory value into a flat byte sequence for transfer or storage; unmarshalling rebuilds a value from those bytes on the other side.
Self-describing encoding
An encoding whose bytes carry field names or type markers alongside values, so a reader with no prior agreement can still parse the structure.
Schema-driven encoding
An encoding that sends values with minimal structural markers, relying on a schema both sides hold to say which bytes mean which field.
Interface definition language
A neutral notation for declaring messages and their fields, from which code for each consuming language is generated; the declaration, not any codebase, owns the contract.
Field tag
A small number standing in for a field's name in the bytes; the stable identity a tag-numbered schema uses to match fields across versions.
Varint
A variable-length integer encoding that spends fewer bytes on small values, using a marker bit in each byte to say whether more bytes follow.
Backward compatibility
A reader built against the newer schema can still decode data produced under an older one.
Forward compatibility
A reader built against an older schema can still decode data produced under a newer one, typically by skipping what it does not recognise.
Unknown field
A field present in the bytes that the reader's schema version does not declare; what the reader does with it decides whether additions are safe.
Schema registry
A central service storing versioned schemas under named subjects and checking each new version against a configured compatibility rule before accepting it.
Framing
The rule that marks where one message ends in a continuous byte stream, usually a length prefix or a reserved delimiter.
Canonical encoding
A variant of an encoding that removes the writer's free choices, so equal values always produce identical bytes; needed for hashing and signatures.
Zero-copy access
Reading fields directly from the received buffer by computed offsets, instead of decoding the whole message into separate objects first.
Columnar layout
Storing all values of one field together rather than whole records together, so a scan reads only the fields it needs and compresses well.
Parser differential
Two decoders interpreting the same bytes differently, for example on duplicate keys or number range; exploitable when one checks and another acts.

The sections lean on each other in a fairly fixed order. **The wire form model is the vocabulary.** Every later section assumes you can say whether an encoding is text or binary, whether it carries its own structure, how integers and lengths are laid out as bytes, and what happens to shared or cyclic references. **Formats are points on that map.** The [format families](/topics/found-serialization-formats) section is less about memorising encodings than about placing them: human-readable interchange, compact binary without a schema, schema-driven binary, columnar files for analytics, and runtime-native object snapshots. The [choosing](/topics/found-serialization-formats-choosing) subsection turns the map into a decision from workload properties. **Schema evolution and wire contracts are two views of one problem.** Evolution asks how a contract may change once deployed code depends on it; contract design asks what the contract promises at any single moment. They meet at field identity, defaults and unknown fields: a reader's handling of a field it has never seen is what makes an added field safe or unsafe, and the way a schema is sourced, from an interface definition or from application types, decides whether a refactor can change the wire without anyone noticing. **Performance is the price of the earlier choices.** Payload size follows from whether names and structure travel with the data; decode cost follows from how much the reader must discover; in-place access is only possible when the layout was designed for it. Compression sits beside the encoding and overlaps with it rather than stacking. **Security is the decode path seen from the attacker's side.** Length prefixes, nesting and decompression become resource-exhaustion vectors, and any encoding that tolerates ambiguity lets two decoders in one request path see different values. That section reuses almost everything before it, which is why it comes last.

  1. Marshalling and Boundaries →

    Why a value cannot simply be copied out of memory; the question every other section assumes you have already answered.

  2. Self-Describing vs Schema-Driven →

    The split between structure in the bytes and structure in a shared schema drives format choice, size and evolution alike.

  3. Format Families and Trade-offs →

    Places whole encoding families on the map so you can defend a choice for a given workload instead of reciting syntax.

  4. Backward and Forward Guarantees →

    Backward and forward compatibility are the terms every later change, versioning and registry question is phrased in.

  5. Nullability and Optionality →

    Absent, null and default are where real contracts break; learn to tell them apart before designing partial updates or defaults.

  6. Resource Exhaustion Limits →

    The first decode-path security topic: why limits must sit inside the parser, framed with the performance ideas you now know.

  • Treating absent, null and empty as the same thing, then discovering a partial update cleared a field the client never meant to touch.

  • Sending money or large identifiers as binary floating-point numbers; the value changes on decode and nothing reports an error.

  • Renaming a field in application code when the serializer derives wire names from identifiers, turning an internal refactor into a breaking contract change.

  • Reusing a deleted field's number or name in a tag-numbered schema, so old readers misinterpret new data under an identity they already trust.

  • Claiming a change is compatible without saying in which direction; the answer decides whether readers or writers must deploy first.

  • Letting an intermediary decode and re-emit messages with an older schema, which quietly strips fields it does not know.

  • Assuming each socket read returns exactly one message, instead of buffering until the framing rule says a whole message is present.

  • Checking body size after parsing, or only checking bytes, and leaving nesting depth, declared lengths and decompressed size unbounded.

  • Validating one decoder's view of a payload and then forwarding the original bytes to a different decoder downstream.

  • Using a runtime-native object snapshot for anything stored or shared, where a class rename or move breaks every existing reader.

The same handful of choices recur across the sections. Naming the one you are making, and the workload fact that decides it, usually earns more credit than the choice itself. - **Readability against compactness.** Text can be read, diffed and hand-edited, and a parser exists everywhere; binary is smaller and cheaper to decode. Who reads the bytes, people or only machines, and how much volume flows usually settle it. - **Self-description against a shared schema.** Carrying structure in every message costs bytes but lets any reader parse without prior agreement. A shared schema saves those bytes and enables code generation and compatibility checks, at the price of distributing and governing the schema. - **Contract-first against code-first.** Declaring the contract in a neutral definition makes it reviewable and stable; deriving it from application types is faster to start but lets every refactor edit the wire. - **Strict against tolerant readers.** Rejecting unknown or malformed input catches mistakes early and closes security gaps; tolerating it allows independent upgrades. Many systems want tolerance for unknown fields and strictness for ambiguity. - **Decode everything against read in place.** A full decode gives convenient objects and pays proportional to message size; in-place layouts make partial reads cheap but cost wire size and constrain how the buffer's lifetime is managed. - **Encoding change against simpler levers.** Migrating formats is expensive; pruning unread fields, batching and compression are often cheaper wins.

Several shapes appear under different names across the sections; recognising them helps you place an unfamiliar question quickly. - **Separate identity from label.** Numbered tags, reserved identifiers and registry subjects all keep what identifies a field or schema stable while its human-facing name is free to change. - **Ship an incompatible change as compatible steps.** Expand the contract, write both forms, move readers, backfill, stop writing the old form, then remove it. The same sequence appears in schema evolution, versioned endpoints and storage migrations. - **Make presence explicit.** Wrapper values, presence flags and lists of updated field paths all exist because a bare value cannot say whether the sender meant to set it. - **Count while reading.** Size, depth, declared length and decompressed output are each bounded by a counter the decoder checks as it goes, never by a measurement taken after the work is done. - **Interpret input a single time.** Normalising or canonically re-encoding it at one entry point removes both repeated decode cost and the room for two decoders to disagree. - **Carry a version or schema reference with the data.** Whether it is a field, a header or an identifier pointing into a registry, the reader needs to know which contract produced the bytes before it can resolve them against its own.

explore

report an issue with this guide →

questions

133 · 6 sections

What distinguishes a text encoding from a binary encoding on the wire, and what does each choice cost?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A text encoding writes values as readable characters you can grep, diff and hand-edit; a binary encoding packs them as raw sized bytes that are typically smaller and cheaper to parse, but unreadable without a decoder.

open as a page

Why can a program not write a value's in-memory bytes straight to a file for another process to read later?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An in-memory value is laid out for one running process: it holds machine addresses valid only there, plus alignment holes whose contents are undefined. Another process shares none of that context, so the value must be re-expressed in a flat, self-standing form.

open as a page

Why must a binary wire format pin the byte order of its fixed-width integer fields instead of leaving it to each writer?

level: middleimportance: must knowfreq 66%
basics
~20 s

Because a multi-byte integer can be laid out most-significant byte first or least-significant byte first, and the two readings disagree: the bytes 00 00 00 01 mean 1 one way and 16,777,216 the other. Only the format can settle which.

open as a page

How does a base-128 varint use the top bit of each byte to encode an integer whose width the reader does not know in advance?

level: middleimportance: must knowfreq 62%
basics
~20 s

A base-128 varint carries seven payload bits per byte and spends the eighth as a continuation flag: set means another byte belongs to this number, clear means this is the last one. The reader stops at the first byte whose top bit is clear.

open as a page

How does a serializer that emits references by identifier keep a cyclic object graph from driving its encoder into unbounded recursion?

level: middleimportance: must knowfreq 58%
basics
~20 s

It keeps an identity-keyed map of nodes already written. A node's first visit assigns an identifier, recorded before recursing, and writes the body once; every later encounter writes only a reference, so no body is written twice.

open as a page

Why does the JSON/XML/YAML family dominate partner-facing API interchange despite producing large, verbose documents?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Ubiquity and human readability win. A parser is already on every platform, so a partner integrates with no shared artifact; a person can read, paste and hand-edit a document; and the text diffs in review. Verbosity is the accepted price.

open as a page

A sensor's JSON telemetry record is re-encoded in MessagePack or CBOR: what changes, and what stays the same?

level: middleimportance: must knowfreq 62%
basics
~20 s

The data model survives; only the syntax is replaced. Maps, arrays, strings, numbers, booleans and null become tagged bytes instead of punctuation and decimal text, so records shrink and decode cheaper — but every key name still travels on every record.

open as a page

When choosing a wire encoding for a new service-to-service integration, which properties of the workload actually decide it?

level: middleimportance: must knowfreq 62%
basics
~20 s

Six workload properties carry most of the decision: payload shape and volume, how fast the contract will change, who the consumers are and what tooling they run, whether humans read the bytes, decoder safety posture, and migration cost.

open as a page

Why do per-column encodings such as run-length, dictionary and delta pay off in a columnar file but not a row-major one?

level: middleimportance: must knowfreq 56%
basics
~20 s

Values inside one column share a type and a value domain, so runs repeat, distinct values are few and neighbouring numbers sit close together — exactly the redundancy run-length, dictionary and delta encoding exploit. Row-major order interleaves unrelated fields and destroys it.

open as a page

Why does a column-oriented file layout beat a row-oriented one for a query that reads three of a table's 200 columns?

level: middleimportance: must knowfreq 68%
basics
~20 s

A column-oriented layout stores each column's values contiguously, so a scan reads only the three columns the query names instead of every byte of every row; a row-major file interleaves all 200 fields, so the whole record comes along.

open as a page

A wire schema gains a new version: what does backward compatibility guarantee about readers and data, and how does forward compatibility differ?

level: middleimportance: must knowfreq 78%
basics
~20 s

Backward compatibility means a reader on the new schema can decode data written under the old one. Forward compatibility is the mirror image: a reader still on the old schema can decode data written under the new one.

open as a page

When a central schema registry rejects a new version of a record schema, what has it actually compared?

level: middleimportance: must knowfreq 62%
basics
~20 s

The registry compares the submitted schema document against the stored versions of one named lineage, in the direction that lineage's compatibility mode requires. It is a document-to-document check made at registration time, not an inspection of deployed readers or code.

open as a page

In a shared schema, why is adding an optional field with a default safe for old readers, while promoting a field to required is not?

level: middleimportance: must knowfreq 66%
basics
~20 s

Adding an optional field only asks an old reader to skip bytes it never needed, and its default gives new readers a value when old writers omit it. Making a field required invalidates every message already written without it.

open as a page

Why can renaming a field be a non-event in one schema family and a breaking change in another, given the same wire bytes?

level: middleimportance: must knowfreq 70%
basics
~20 s

Identity decides. Where a reader matches fields by number, a rename changes only the label. Where it matches by name, a rename is a delete plus an add, so readers on the old schema quietly stop finding the value.

open as a page

A service must mark which version of its payload shape a message carries. Where can that version identifier live, and what does each placement cost?

level: middleimportance: must knowfreq 68%
basics
~20 s

Three placements are common: a version field inside the payload, a negotiated media type, and a version segment in the address. Each moves the cost elsewhere - routing and caching, client effort, or resource identity - and none removes it.

open as a page

In a JSON request body, what is the difference between an absent field, a null field, and an empty-string field?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Absent means the sender said nothing about the field. Null means the sender said it has no value. Empty means it has a value that happens to be zero-length. Three different statements, and many encodings collapse them.

open as a page

A content-addressed store keys artefacts by a digest of their encoded bytes, so why can two services encoding the same value produce different ids?

level: middleimportance: must knowfreq 62%
basics
~20 s

Most encodings give the encoder freedom — member order, whitespace, number spelling, escapes, omitting defaults — so one value has many valid byte sequences. A canonical encoding fixes every one of those choices so the digest depends only on the value.

open as a page

What happens to unknown fields when an older reader decodes a newer producer's message and re-emits it?

level: middleimportance: must knowfreq 66%
basics
~20 s

They are usually dropped. Decoding keeps only the fields the reader's own schema version names, and the output is rebuilt from that smaller in-memory value, so the newer producer's data disappears with no error anywhere. Preserving it requires a deliberate side buffer.

open as a page

Why does framing by a delimiter byte force the writer to escape or re-encode the payload?

level: middleimportance: must knowfreq 54%
basics
~20 s

A delimiter means 'the message ends here', so any occurrence of that byte inside the payload would end the frame early. The writer must escape it, or restrict the payload to an alphabet that excludes it.

open as a page

Framing turns a byte stream into messages: on a connection whose reads never align with message boundaries, what must the reader's framing loop do?

level: middleimportance: must knowfreq 66%
basics
~20 s

A stream connection carries bytes, not messages, so the reader keeps a buffer across reads: append each chunk, test whether a complete message is present from its length prefix or delimiter, consume exactly that message, and retain the remainder.

open as a page

Why is a small JSON log record several times larger than the same values in a schema-driven binary encoding?

level: middleimportance: must knowfreq 68%
basics
~20 s

Most of a short JSON record is structure, not data: field names repeated on every record, quotes and punctuation, and numbers spelled out as digit characters. A schema-driven binary encoding sends small field tags and raw integer bytes instead.

open as a page

Why is the decode side of a message pipeline usually more expensive per byte than the encode side?

level: middleimportance: must knowfreq 58%
basics
~20 s

The writer already knows which fields exist and appends them into one buffer. The reader starts from opaque bytes and must discover boundaries and types, validate what it cannot trust, convert, and allocate an object per value. Discovery, validation and allocation have no write-side counterpart.

open as a page

When does a streaming parser that emits events beat materialising the whole decoded document in memory?

level: middleimportance: must knowfreq 66%
basics
~20 s

Streaming wins when the document is large relative to memory, or unbounded, and the reader consumes it in one forward pass — peak memory then tracks the largest single value, not the payload. It loses when the work needs random access or back-references.

open as a page

What does in-place field access mean: reading straight out of a received buffer instead of decoding the message first?

level: middleimportance: must knowfreq 55%
basics
~20 s

In-place access means the received bytes are the data structure: each field is located by offset arithmetic and read on demand, so no decode pass builds a parallel tree of objects. The buffer must stay alive and unchanged while those values are used.

open as a page

A gzip-compressed ingestion link moves JSON batches; why does switching to a compact binary encoding save far less than the raw sizes suggest?

level: seniorimportance: must knowfreq 58%
basics
~20 s

Because the compressor has already removed most of what the compact encoding removes. Repeated field names collapse into short back-references, so the two levers overlap instead of stacking, and the dense binary payload has far less redundancy left to squeeze.

open as a page

A JSON object carries the same member name twice: why might a filtering proxy and the backend behind it read different values?

level: middleimportance: must knowfreq 55%
basics
~20 s

The grammar allows a repeated member name but the data model holds one value per name, so each decoder invents its own resolution: last wins, first wins, reject, or keep both. Two decoders can pick differently.

open as a page

Why does a byte-size cap on a request body fail to protect a recursive-descent decoder from deeply nested input?

level: middleimportance: must knowfreq 56%
basics
~20 s

A byte is cheap and a stack frame is not: a megabyte of opening brackets is about a million nesting levels, and a decoder that recurses per level exhausts its stack long before it exhausts that input. Depth needs its own ceiling.

open as a page

On a public upload endpoint, why must a decoder's size and nesting ceilings be enforced during the parse rather than after it?

level: middleimportance: must knowfreq 62%
basics
~20 s

By the time a parse finishes, the memory, CPU and stack the sender asked for have already been spent, so a later check only reports the damage. A ceiling has to be a counter the parser itself tests as it reads, aborting mid-stream.

open as a page

Why is validating a decoded copy of a request body unsafe when the original bytes are forwarded unchanged?

level: seniorimportance: must knowfreq 50%
basics
~20 s

Validation applies to one decoder's interpretation of the bytes, and that interpretation is thrown away when the raw bytes are forwarded. The next decoder derives its own meaning, so the payload that is used was never the payload that was checked.

open as a page

A binary decoder allocates a buffer from the four-byte length prefix it just read. What does that enable?

level: middleimportance: should knowfreq 48%
basics
~20 s

A few bytes declaring four gigabytes make the server reserve four gigabytes, and the sender never has to produce them. A declared length is a claim about the stream, not authorisation to reserve memory on the sender's behalf.

open as a page