A gzip-compressed ingestion link moves JSON batches; why does switching to a compact binary encoding save far less than the raw sizes suggest?
answer
- the two levers overlap
- repeated names become back-references
- a dense payload has less left to squeeze
- compare compressed sizes, not raw
- the gap narrows sharply after compression
basics
~20 sBecause the compressor has already removed most of what the compact encoding removes. Repeated field names collapse into short back-references, so the two levers overlap instead of stacking, and the dense binary payload has far less redundancy left to squeeze.
solid answer
~40 sA general-purpose compressor replaces a byte run it has seen before with a short reference back to the earlier occurrence. In a batch of similar records, the field names, quotes and punctuation are exactly that: after the first record, each repeat costs a handful of bits rather than its full length. A compact binary encoding removes the same bytes, permanently — so the two are competing for the same redundancy, not adding up. Meanwhile the compact payload compresses in a **worse ratio**, because tags and variable-length integers leave little repetition behind. The result is that a four-to-one raw gap typically collapses to a much narrower compressed gap, and which encoding wins after compression depends on the record shape. The only honest comparison is compressed byte counts over a sample of your own records.
go deeper
Know that wire payloads are often compressed, so a comparison of uncompressed sizes does not tell you what the link will actually carry or what it will cost.
Explain the mechanism: a compressor emits a short reference to an earlier identical run of bytes, and repeated field names across a batch are exactly such runs.
Insist on compressed byte counts over a sample of real records before committing to a wire-format migration, and report the raw and compressed figures separately rather than blending them.
Decide what the migration is really buying. If the compressed saving turns out small, the case has to rest on schema enforcement, decode cost or tooling, and should be argued openly on those terms.
## What a general-purpose compressor actually removes A general-purpose compressor works on bytes, not on your schema. It keeps a window of recently seen input, and when the next stretch of bytes has appeared before, it emits a **back-reference** — a distance and a length — instead of the bytes themselves. What remains is then coded more compactly according to how often each symbol occurs. That mechanism is aimed squarely at **repetition**, and a batch of similar records is one of the most repetitive things a system produces. The second record's `"service":"` is byte-for-byte the first record's, and so is the third's. After the opening record, the structural skeleton of every later record costs a handful of bits. ## Why the two levers overlap Now put the two size levers side by side on the same batch: | what it removes | a compact binary encoding | a general-purpose compressor | |---|---|---| | repeated field names | removed outright, replaced by tags | collapsed into back-references | | quotes, colons, commas | removed outright | collapsed into back-references | | digits spelled as characters | replaced by integer bytes | partly, where digit runs repeat | | distinct values, ids, free text | untouched | little to remove — nothing to point back at | The first two rows are the bulk of a short text record — around 70% of a five-field log line — and **both levers claim them**. Apply either one and those bytes are largely gone; apply both and the second one finds the work already done. This is why the arithmetic people reach for is wrong: you cannot take a four-to-one raw ratio, multiply it by the link's compression ratio, and report the product. The savings intersect. There is a second effect pulling the same way. The compact payload is **denser**: tag bytes and variable-length integers vary from record to record and look close to random to a compressor, so its own compression ratio is typically much worse than the text payload's. The text batch has a long way to fall; the binary batch does not. Both effects narrow the compressed gap, and for name-heavy records with few distinct values the gap can close nearly all the way — while for value-dense records a real gap survives. Which side wins after compression is a property of your records, not of the encodings. ## What survives compression What the compressor cannot remove is **variety**: distinct identifiers, free-text messages, timestamps that differ every record, measurement values that never repeat. There is no earlier occurrence to point at, so those bytes cost close to what they cost raw. A useful mental split for a batch is therefore: - **skeleton** — names, punctuation, framing; near-free after compression, whichever encoding you use; - **variety** — the data itself; roughly as expensive either way. How far a compressor can push the second category before it hits a floor is the subject of information theory, and it is not decided by your choice of encoding. ## How to measure it properly 1. **Capture a real sample.** Several thousand consecutive records from production, not a handcrafted example — the repetition across records is the whole effect, and one record cannot show it. 2. **Encode the same sample both ways**, batch it the way the pipeline actually batches, and compress each batch the way the link actually compresses. 3. **Report four numbers**: raw text, compressed text, raw binary, compressed binary. The pair that matters is the compressed one, because that is what the invoice counts. 4. **Re-run it on a second sample from a different producer.** Record shapes vary across a fleet, and a migration is justified on the fleet, not on the best case. ## Where the levers do still stack The overlap is not total, and three cases keep a real compressed saving on the table: - **Small or unbatched messages**, where the compressor never accumulates enough history to amortise the names — it sees one record, finds nothing to reference, and may even add bytes. - **Value-dense records**, where digit spelling and per-value punctuation are a large share of the payload and the names are a small one. - **Payloads carrying armoured binary**, where a text encoding pays a flat one-third expansion for the armour before compression ever runs. And the decision was never only about bytes: schema enforcement, decode cost and tooling belong in the same discussion. The discipline this question is really testing is that a headline raw ratio is not a saving until it has been measured downstream of the compressor.
- Which parts of a compressed batch still cost real bytes?The variety: distinct identifiers, free text, and values that differ from record to record. A compressor can only replace a run it has already seen, and those runs are unique, so they cost close to their raw size whatever the encoding. Structure is cheap after the first record; distinct data is not.
- When does a compact binary encoding still cut the compressed bill significantly?When records are value-dense rather than name-heavy, so the bytes the compressor cannot reach are the ones the encoding shrinks; when messages are small or unbatched, so the compressor never builds enough history to amortise the names; and when a text payload has to armour binary data, paying a flat one-third expansion before compression runs at all.
- Two teams report different savings for the same migration. What most likely differs?Their record shapes and their batching. A producer emitting many short numeric records with long names sees the raw gap almost vanish after compression, while one emitting value-dense records or unbatched single messages keeps most of it. Compare the sampled inputs before concluding that either measurement is wrong.
Two people are sent to tidy the same room an hour apart. Whoever goes first takes most of the mess, and the second finds far less to do than their usual rate would predict — their effect does not add to the first's, it overlaps with it.
saying these in an interview costs you the question
- Multiplies the raw saving by the compression ratio as if independent
- Says a dense binary payload compresses in the same ratio as text
- Claims compression makes the encoding choice irrelevant for every payload
- Quotes an uncompressed benchmark for a link that already compresses
- Forgets a compressor can only reference bytes it has already seen
- Measures one record instead of a batch of real ones