skip to content

When would you choose ZSTD over LZ4 for a columnar table's blocks, and what does it cost?

level: seniorimportance: should knowfreq 48%

answer

  1. Trading bytes moved against CPU cycles
  2. One codec has a level knob, one does not
  3. The two directions do not scale together
  4. Ask what the scan is actually bound by

basics

~20 s

Choose ZSTD when scans are bound by bytes moved — remote object storage, network, or a storage bill — and LZ4 when scans are already CPU-bound on locally cached data. ZSTD buys a better ratio and charges compression CPU, mostly at write time.

solid answer

~50 s

Both are general byte codecs layered over the column encodings; they differ in where they sit on the speed-versus-ratio curve. LZ4 and Snappy optimise for throughput in both directions and give a modest ratio. ZSTD gives a better ratio with **tunable levels**, and its cost is asymmetric: raising the level makes compression substantially more expensive while decompression cost changes far less. So the decision follows the bottleneck. If a scan is dominated by fetching bytes from object storage or across the network, fewer bytes wins and ZSTD is the right default. If the working set is hot on local NVMe or in page cache and the scan is already CPU-bound, the extra decompression can make queries slower even though the table is smaller. Ingest matters too: heavy levels eat CPU on the write path and can cap load throughput. Decide by measuring both stored bytes and CPU time per query on your own data — never by ratio alone.

go deeper

for a junior

Know that both are general compressors over the encoded bytes, that one favours speed and the other favours ratio, and that compression always trades bytes for CPU.

for a middle

Explain the asymmetry — level mainly costs compression time, decompression stays fast — and connect the choice to whether a scan is I/O-bound or CPU-bound.

for a senior

Show the measurement plan: per-column bytes, cold and warm query CPU, ingest throughput, and a tiered policy that treats hot and cold data differently.

for a principal

Own the trade as a spend decision: storage and egress against compute, per temperature tier, with a rule for when the policy gets revisited and who pays for the reorganisation.

## The two codecs on one axis After the per-column encodings have done their work, the residual bytes go through a general-purpose compressor. The realistic candidates are the fast family — **LZ4** and **Snappy** — and the tunable family, principally **ZSTD**. They are not different in kind; they occupy different points on the same speed-versus-ratio curve. - **LZ4 / Snappy:** designed so that compression and decompression are fast enough to be almost free relative to I/O. Ratio is real but modest. - **ZSTD:** a wider search for redundancy, exposed as a **level** knob. Higher levels find more, cost much more CPU to compress, and — importantly — do **not** cost proportionally more to decompress. Decompression throughput stays high across levels; sometimes a higher level even reads faster because there are fewer bytes to touch. That asymmetry is the single most useful fact here: **compression cost is level-sensitive, decompression cost mostly is not.** Data written once and read many times is exactly the shape that rewards a higher level. ## Decide by locating the bottleneck Compression trades bytes for CPU cycles, so the question is only ever: which of the two is scarce for this workload? **Bytes are scarce — prefer the higher ratio.** - Compute reads from remote object storage, so every byte crosses the network and often shows up in latency. - Data is replicated, backed up or copied between regions; the ratio multiplies across every copy. - Storage volume itself is the dominant line item, or the workload scans far more than it caches. - Cold data that is retained for a long time and scanned rarely — the write-side cost is paid once, the storage saving accrues forever. **CPU is scarce — prefer the fast codec.** - The hot working set already sits on local NVMe or in page cache, so decompression is the scan's inner loop rather than a way to avoid I/O. - High-concurrency dashboards where CPU is the contended resource and added milliseconds of decompression multiply across many concurrent queries. - Ingest-heavy pipelines where writers are already CPU-saturated: a heavy level directly caps load throughput and lengthens the write path. ## The middle path: differentiate It is rarely one setting for everything. Sensible policies split by temperature and by column: - **Hot recent data** in a fast codec, **cold historical data** re-compressed at a higher level during a background reorganisation. The cold data is scanned rarely, so its decompression cost matters little and its storage cost matters most. - **Per-column choices** where the engine allows them: a huge, rarely projected text column can carry a heavy codec, while the narrow keys and timestamps that every query touches stay fast. Since columnar scans read only the projected columns, this targets the cost precisely. ## Diminishing returns ZSTD's level scale is not linear in value. Moving off the default costs meaningfully more CPU for progressively smaller ratio gains, and at the top of the range you can spend a large multiple of the compression time for a few percent of size. The practical approach is to test a small number of levels on representative data and take the knee of the curve, not the maximum. Remember too that the codec is the **second** layer. If a column is already dictionary- and run-length-encoded down to near nothing, no codec choice will move the needle; the remaining headroom lives in the encodings and in the sort order that feeds them. Chasing codec levels on a table with a bad sort order is optimising the wrong layer. ## How to actually decide Run the comparison rather than arguing it: 1. Load a representative sample under each candidate setting. 2. Record **compressed bytes per column**, not just table totals — that shows where the win comes from. 3. Replay a realistic query mix and record **wall-clock and CPU time per query**, both cold (uncached) and warm (cached). The cold run shows the I/O saving; the warm run exposes the decompression tax. 4. Record **ingest throughput and writer CPU** under each setting. 5. Choose per temperature tier, and re-check after any major change in data shape. The answer that impresses is not "ZSTD, it compresses better" — it is naming which resource is scarce, showing that you measured both directions, and recognising that the write-once/read-many asymmetry is what makes a higher level defensible at all.

  • Why is a high compression level defensible for cold data but not for hot data?
    Compression cost is paid once at write, decompression on every read. Cold data is written once and scanned rarely, so a heavy level buys durable storage savings for a one-off CPU bill. Hot data is scanned constantly, so any per-read decompression tax is multiplied by query volume and concurrency.
  • A table is already dictionary and run-length encoded down to a fraction of its raw size. What does switching codecs gain?
    Little. The encodings have already removed the structural redundancy, so the codec is working on a nearly incompressible residue. The remaining headroom is in the encoding layer and the sort order that feeds it — changing the load order will usually beat any codec level change on such a table.
  • How would you measure the decompression tax rather than guess at it?
    Replay a representative query mix twice under each setting: cold, with caches dropped, and warm, with the working set resident. The cold run shows the I/O saving from fewer bytes; the warm run isolates decompression CPU, since there is no I/O left to save. Compare CPU time per query, not just wall clock.

saying these in an interview costs you the question

  • Picks the codec purely on compression ratio
  • Assumes higher levels slow reads as much as writes
  • Ignores ingest CPU when choosing a heavy level
  • Applies one codec setting to hot and cold data alike
  • Tunes the codec when the sort order is the real problem

context