skip to content

questions

15

Why is a small JSON log record several times larger than the same values in a schema-driven binary encoding?

level: middleimportance: must knowfreq 68%

answer

  1. structure, not data
  2. the names ship on every record
  3. quotes, colons, commas, braces
  4. numbers spelled as digit characters
  5. a schema turns a name into a tag byte

basics

~20 s

Most of a short JSON record is structure, not data: field names repeated on every record, quotes and punctuation, and numbers spelled out as digit characters. A schema-driven binary encoding sends small field tags and raw integer bytes instead.

solid answer

~40 s

Take a five-field record: `{"timestamp":1789243200,"level":"INFO","service":"checkout","latency_ms":37,"status":200}`. That is 89 bytes, and only 27 of them are values — the other 62 are field names, the quotes around them, the quotes around string values, and the colons, commas and braces a parser needs to find boundaries. Those names ship again on every record even though both ends already know them. A schema-driven binary encoding knows the field order and types up front, so three things change: each name becomes a one-byte numeric tag, each integer travels as its own bytes rather than as digit characters, and length prefixes replace delimiters. With `level` declared as an enumerated code, the same record lands near 23 bytes. The saving is concentrated in short records of short values; a record dominated by one long text field barely moves.

code

json · 1 line
json
{"timestamp":1789243200,"level":"INFO","service":"checkout","latency_ms":37,"status":200}

go deeper

for a junior

Be able to look at a text record and point out which characters are field names, quotes and punctuation rather than values. That split is the whole answer at this level.

for a middle

Name the three levers precisely: a schema removes the repeated names, integers travel as bytes rather than digit characters, and tags or length prefixes replace delimiters that a parser would otherwise scan for.

for a senior

Bring the numbers from a real record shape and say which levers actually pay on that data — a payload of long free text barely moves, one of short numeric fields moves by a large factor.

for a principal

Frame it as a bill rather than a benchmark. The per-record gap only matters multiplied by daily volume, and a wire-format change lands on every producer and consumer team at once.

## Where the bytes in a text record actually go Serialization size arguments usually start as a slogan — "binary is smaller" — and the slogan hides the mechanism. The useful move is to take one record and account for every byte in it. Here is a log line with five short fields, written as a **self-describing text encoding** (one that carries its own field names inside the payload): ```json {"timestamp":1789243200,"level":"INFO","service":"checkout","latency_ms":37,"status":200} ``` That is **89 bytes** of ASCII, which UTF-8 stores one byte per character. Broken down: | part of the record | bytes | what it carries | |---|---|---| | field names | 37 | `timestamp`, `level`, `service`, `latency_ms`, `status` | | quotes around the names | 10 | two per name | | quotes around string values | 4 | two per string value | | colons, commas, braces | 11 | five colons, four commas, two braces | | **the values themselves** | **27** | the only bytes the reader did not already know | So **62 of 89 bytes — about 70% — are structure**, and every one of them repeats identically on the next record and the one after that. That is the observation the rest of this subject is built on: in short records, the payload is mostly a description of the payload. ## The three levers a schema pulls When both ends hold a **schema** — an agreed list of fields, each with a number and a type — the encoder can stop describing and start assuming. Three distinct levers come into play, and it is worth naming them separately because they pay off differently on different data: 1. **Field-name elimination.** A field arrives as a small numeric **tag** rather than a quoted name. In a tag scheme that packs the field number together with a three-bit type code, field numbers 1–15 fit in a single tag byte and 16 upward need two — which is why the fields present on every record are given the low numbers. 2. **Numbers as numbers.** `1789243200` is ten digit characters in text. As a **variable-length integer** it is five bytes; as a fixed 32-bit field, four. The digits were never the value, only a spelling of it. 3. **Delimiters replaced by structure.** Quotes, colons and commas exist so a parser can find where one field ends. A tag plus a length prefix (or a fixed width) does the same job in one byte instead of several, and does it without scanning for a character. ## The same record against a schema | field | binary cost | why | |---|---|---| | `timestamp` = 1789243200 | 6 bytes | 1 tag byte + 5 varint bytes (a 31-bit value) | | `level` = an enumerated code | 2 bytes | 1 tag byte + 1 varint byte | | `service` = `checkout` | 10 bytes | 1 tag byte + 1 length byte + 8 characters | | `latency_ms` = 37 | 2 bytes | 1 tag byte + 1 varint byte (37 < 128) | | `status` = 200 | 3 bytes | 1 tag byte + 2 varint bytes (200 ≥ 128) | | **total** | **23 bytes** | against 89 | Roughly a quarter of the size — **for this record shape, and with `level` declared as an enumerated code rather than sent as text**. Both qualifications matter; a ratio quoted without the record it came from is not a fact about encodings, it is a fact about someone else's data. ## Which payloads move, and which do not - **Many short records with numeric, enumerated and boolean fields** — the best case. Structure dominates, and structure is what disappears. - **Long field names over short values** — also strongly improved, because the name is the bulk. - **Records whose bulk is free text** — barely improved. A 2 KB message body is 2 KB either way; only the structural fraction shrinks. - **Records carrying binary blobs** — a text encoding must armour them, typically in Base64, which costs four characters per three bytes, a flat one-third expansion. A binary encoding carries the bytes as they are, so here the saving is the armour, not the structure. - **Records where one string field is repeated across records** (a service name, a host) — worth turning into an enumerated code or an interned identifier, which is a schema decision rather than an encoding one. ## What the comparison does not settle Two cautions belong next to every number above. First, raw size is not the bill if the link compresses: a general-purpose compressor removes repeated names too, so the two savings overlap rather than stack, and the honest comparison is between compressed byte counts of the same records. Second, a **self-describing binary encoding** — one that carries keys inside the payload — takes the punctuation and the digit spelling but keeps the repeated names, so it lands between the two poles rather than at the compact end. Whether the remaining gap is worth a migration is a question about volume and disruption, not about one record.

  • Which record shapes barely shrink when a pipeline moves from a text encoding to a compact binary one?
    Payloads whose bulk is one long free-text value, or an already-binary blob. Only the structural bytes and the numeric spellings shrink, so a 2 KB message body travels at nearly the same size either way — though a text encoding must armour raw bytes, typically in Base64, which costs a flat third extra. The gap is largest for many small records of short numeric and enumerated fields.
  • Why does a self-describing binary encoding not reach the size of a schema-driven one?
    Because it still carries the keys, or at least type markers, inside every record. It removes the quoting, the delimiters and the digit spelling, but the field names repeat exactly as they did in text. That puts it between the two poles: clearly smaller than a text encoding, clearly larger than one where both ends already hold the field list.
  • How much does simply shortening field names buy on a text encoding?
    On raw bytes, a lot in a record like the one above — the names and their quotes are 47 of 89 bytes. But it is still a contract change for every consumer, it costs readability, and on a compressed link most of that saving has already been taken, because repeated names are exactly what a compressor collapses. It is rarely the lever with the best ratio of saving to disruption.

Every carton in a shipment is labelled with the full text of the packing list, item names spelled out, even though the warehouse at both ends already holds the same numbered catalogue. Switching to catalogue numbers does not change what is in the cartons, only what is written on them.

saying these in an interview costs you the question

  • Thinks the gap is mostly whitespace, so minifying closes it
  • Believes a number costs the same as text or as raw bytes
  • Says binary encodings are smaller because they compress the data
  • Assumes a reader needs the field names, so they cannot be dropped
  • Forgets the names repeat on every record rather than once per stream
  • Quotes one encoding ratio as if it held for every record shape
open as a page

Why is the decode side of a message pipeline usually more expensive per byte than the encode side?

level: middleimportance: must knowfreq 58%

basics

~20 s

The writer already knows which fields exist and appends them into one buffer. The reader starts from opaque bytes and must discover boundaries and types, validate what it cannot trust, convert, and allocate an object per value. Discovery, validation and allocation have no write-side counterpart.

open as a page

When does a streaming parser that emits events beat materialising the whole decoded document in memory?

level: middleimportance: must knowfreq 66%

basics

~20 s

Streaming wins when the document is large relative to memory, or unbounded, and the reader consumes it in one forward pass — peak memory then tracks the largest single value, not the payload. It loses when the work needs random access or back-references.

open as a page

What does in-place field access mean: reading straight out of a received buffer instead of decoding the message first?

level: middleimportance: must knowfreq 55%

basics

~20 s

In-place access means the received bytes are the data structure: each field is located by offset arithmetic and read on demand, so no decode pass builds a parallel tree of objects. The buffer must stay alive and unchanged while those values are used.

open as a page

A gzip-compressed ingestion link moves JSON batches; why does switching to a compact binary encoding save far less than the raw sizes suggest?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Because the compressor has already removed most of what the compact encoding removes. Repeated field names collapse into short back-references, so the two levers overlap instead of stacking, and the dense binary payload has far less redundancy left to squeeze.

open as a page

A consumer reads two fields out of each multi-megabyte message. Why does an offset-addressed in-place layout cut its cost so sharply?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Because the cost model changes from pay-everything-up-front to pay-per-field-read. A decode pass is proportional to message size; reading by stored offset touches only the two fields, so the other several megabytes are never interpreted, allocated or copied.

open as a page

When does encoding an integer field as a varint make a payload larger rather than smaller?

level: middleimportance: should knowfreq 45%

basics

~20 s

When the values are large or negative. A varint spends seven bits per byte, so a 32-bit value at or above 2^28 costs five bytes against four fixed, and a full-width 64-bit value costs ten against eight.

open as a page

Why is a message laid out for in-place field access usually larger on the wire than a tightly packed encoding?

level: middleimportance: should knowfreq 36%

basics

~20 s

Computable field positions have to be bought. Fields sit at fixed, aligned offsets, so padding fills the gaps; each object carries a table of offsets; and scalars are stored at full natural width rather than packed by magnitude.

open as a page

A service's tail latency tracks collector pauses while it decodes many small messages per request; how would you confirm decoding is the source?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Separate churn from a leak first: a high allocation rate with a flat post-collection live set means short-lived garbage, not growth. Then attribute allocation by sampling site, expect decode frames, and confirm by removing the decode from the path and watching the rate fall.

open as a page

Two processes share a mapped region where one writes frames and the other reads fields in place. What lifetime hazard does that create?

level: seniorimportance: should knowfreq 31%

basics

~20 s

Every value the reader holds is a view into the shared bytes, not a copy. The moment the writer reuses that part of the region, or the mapping goes away, those views silently refer to different data — and each field is read at its own instant, so a frame can be seen half-overwritten.

open as a page

Your ingestion pipeline's per-byte bill is the top cost line; how do you choose which payload-size lever to spend a quarter on?

level: principalimportance: should knowfreq 38%

basics

~20 s

Measure first: compressed bytes per field per record, multiplied by volume, for both the moved and the stored bill. Then rank levers by saving against disruption — pruning unread fields and batching usually beat a wire-format migration.

open as a page

How would you tune decoding differently for a throughput-bound batch job versus a tail-latency-bound request path?

level: principalimportance: should knowfreq 34%

basics

~20 s

A batch job optimises average cost per byte and can absorb pauses, so amortise: large buffers, big batches, parallel readers. A request path optimises the worst percentile, so eliminate bimodal costs — resizes, pauses, warm-up, shared-pool contention — even at a worse mean.

open as a page

Why does compressing a batch of many small log records beat compressing and sending each record on its own?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Two fixed costs are paid once per message rather than once per record: the envelope around the message and the compressed stream's own framing. A compressor also has no earlier bytes to reference until it has seen some.

open as a page

A reader needs three of a message's forty fields — when does decoding the remaining fields lazily actually pay off?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

It pays when skipping is cheap — length-prefixed fields can be stepped over arithmetically — and the skipped fields are expensive to materialise. It costs when the encoding forces a byte scan anyway, when the retained bytes outlive the request, or when everything is eventually read.

open as a page

A team wants to move a hot internal path to an in-place encoding for latency. What would you require them to prove first?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Prove the ceiling before paying the price: decode must be a measured share of latency, consumers must read only part of each message, the hop must not be bandwidth-bound, and someone must own the rule that keeps buffers alive under their views.

open as a page