Why is a small JSON log record several times larger than the same values in a schema-driven binary encoding?
answer
- structure, not data
- the names ship on every record
- quotes, colons, commas, braces
- numbers spelled as digit characters
- a schema turns a name into a tag byte
basics
~20 sMost of a short JSON record is structure, not data: field names repeated on every record, quotes and punctuation, and numbers spelled out as digit characters. A schema-driven binary encoding sends small field tags and raw integer bytes instead.
solid answer
~40 sTake a five-field record: `{"timestamp":1789243200,"level":"INFO","service":"checkout","latency_ms":37,"status":200}`. That is 89 bytes, and only 27 of them are values — the other 62 are field names, the quotes around them, the quotes around string values, and the colons, commas and braces a parser needs to find boundaries. Those names ship again on every record even though both ends already know them. A schema-driven binary encoding knows the field order and types up front, so three things change: each name becomes a one-byte numeric tag, each integer travels as its own bytes rather than as digit characters, and length prefixes replace delimiters. With `level` declared as an enumerated code, the same record lands near 23 bytes. The saving is concentrated in short records of short values; a record dominated by one long text field barely moves.
code
json · 1 line{"timestamp":1789243200,"level":"INFO","service":"checkout","latency_ms":37,"status":200}go deeper
Be able to look at a text record and point out which characters are field names, quotes and punctuation rather than values. That split is the whole answer at this level.
Name the three levers precisely: a schema removes the repeated names, integers travel as bytes rather than digit characters, and tags or length prefixes replace delimiters that a parser would otherwise scan for.
Bring the numbers from a real record shape and say which levers actually pay on that data — a payload of long free text barely moves, one of short numeric fields moves by a large factor.
Frame it as a bill rather than a benchmark. The per-record gap only matters multiplied by daily volume, and a wire-format change lands on every producer and consumer team at once.
## Where the bytes in a text record actually go Serialization size arguments usually start as a slogan — "binary is smaller" — and the slogan hides the mechanism. The useful move is to take one record and account for every byte in it. Here is a log line with five short fields, written as a **self-describing text encoding** (one that carries its own field names inside the payload): ```json {"timestamp":1789243200,"level":"INFO","service":"checkout","latency_ms":37,"status":200} ``` That is **89 bytes** of ASCII, which UTF-8 stores one byte per character. Broken down: | part of the record | bytes | what it carries | |---|---|---| | field names | 37 | `timestamp`, `level`, `service`, `latency_ms`, `status` | | quotes around the names | 10 | two per name | | quotes around string values | 4 | two per string value | | colons, commas, braces | 11 | five colons, four commas, two braces | | **the values themselves** | **27** | the only bytes the reader did not already know | So **62 of 89 bytes — about 70% — are structure**, and every one of them repeats identically on the next record and the one after that. That is the observation the rest of this subject is built on: in short records, the payload is mostly a description of the payload. ## The three levers a schema pulls When both ends hold a **schema** — an agreed list of fields, each with a number and a type — the encoder can stop describing and start assuming. Three distinct levers come into play, and it is worth naming them separately because they pay off differently on different data: 1. **Field-name elimination.** A field arrives as a small numeric **tag** rather than a quoted name. In a tag scheme that packs the field number together with a three-bit type code, field numbers 1–15 fit in a single tag byte and 16 upward need two — which is why the fields present on every record are given the low numbers. 2. **Numbers as numbers.** `1789243200` is ten digit characters in text. As a **variable-length integer** it is five bytes; as a fixed 32-bit field, four. The digits were never the value, only a spelling of it. 3. **Delimiters replaced by structure.** Quotes, colons and commas exist so a parser can find where one field ends. A tag plus a length prefix (or a fixed width) does the same job in one byte instead of several, and does it without scanning for a character. ## The same record against a schema | field | binary cost | why | |---|---|---| | `timestamp` = 1789243200 | 6 bytes | 1 tag byte + 5 varint bytes (a 31-bit value) | | `level` = an enumerated code | 2 bytes | 1 tag byte + 1 varint byte | | `service` = `checkout` | 10 bytes | 1 tag byte + 1 length byte + 8 characters | | `latency_ms` = 37 | 2 bytes | 1 tag byte + 1 varint byte (37 < 128) | | `status` = 200 | 3 bytes | 1 tag byte + 2 varint bytes (200 ≥ 128) | | **total** | **23 bytes** | against 89 | Roughly a quarter of the size — **for this record shape, and with `level` declared as an enumerated code rather than sent as text**. Both qualifications matter; a ratio quoted without the record it came from is not a fact about encodings, it is a fact about someone else's data. ## Which payloads move, and which do not - **Many short records with numeric, enumerated and boolean fields** — the best case. Structure dominates, and structure is what disappears. - **Long field names over short values** — also strongly improved, because the name is the bulk. - **Records whose bulk is free text** — barely improved. A 2 KB message body is 2 KB either way; only the structural fraction shrinks. - **Records carrying binary blobs** — a text encoding must armour them, typically in Base64, which costs four characters per three bytes, a flat one-third expansion. A binary encoding carries the bytes as they are, so here the saving is the armour, not the structure. - **Records where one string field is repeated across records** (a service name, a host) — worth turning into an enumerated code or an interned identifier, which is a schema decision rather than an encoding one. ## What the comparison does not settle Two cautions belong next to every number above. First, raw size is not the bill if the link compresses: a general-purpose compressor removes repeated names too, so the two savings overlap rather than stack, and the honest comparison is between compressed byte counts of the same records. Second, a **self-describing binary encoding** — one that carries keys inside the payload — takes the punctuation and the digit spelling but keeps the repeated names, so it lands between the two poles rather than at the compact end. Whether the remaining gap is worth a migration is a question about volume and disruption, not about one record.
- Which record shapes barely shrink when a pipeline moves from a text encoding to a compact binary one?Payloads whose bulk is one long free-text value, or an already-binary blob. Only the structural bytes and the numeric spellings shrink, so a 2 KB message body travels at nearly the same size either way — though a text encoding must armour raw bytes, typically in Base64, which costs a flat third extra. The gap is largest for many small records of short numeric and enumerated fields.
- Why does a self-describing binary encoding not reach the size of a schema-driven one?Because it still carries the keys, or at least type markers, inside every record. It removes the quoting, the delimiters and the digit spelling, but the field names repeat exactly as they did in text. That puts it between the two poles: clearly smaller than a text encoding, clearly larger than one where both ends already hold the field list.
- How much does simply shortening field names buy on a text encoding?On raw bytes, a lot in a record like the one above — the names and their quotes are 47 of 89 bytes. But it is still a contract change for every consumer, it costs readability, and on a compressed link most of that saving has already been taken, because repeated names are exactly what a compressor collapses. It is rarely the lever with the best ratio of saving to disruption.
Every carton in a shipment is labelled with the full text of the packing list, item names spelled out, even though the warehouse at both ends already holds the same numbered catalogue. Switching to catalogue numbers does not change what is in the cartons, only what is written on them.
saying these in an interview costs you the question
- Thinks the gap is mostly whitespace, so minifying closes it
- Believes a number costs the same as text or as raw bytes
- Says binary encodings are smaller because they compress the data
- Assumes a reader needs the field names, so they cannot be dropped
- Forgets the names repeat on every record rather than once per stream
- Quotes one encoding ratio as if it held for every record shape