In a columnar table, how does a per-column encoding differ from a block codec like ZSTD?
answer
- Two layers, and their order is fixed
- One layer knows the data type
- Codes stay queryable, compressed bytes do not
- Dictionary and run length, then LZ4 or ZSTD
basics
~20 sAn encoding rewrites one column's values in a type-aware way — dictionary codes, run lengths, deltas — and can often be queried directly. A codec such as ZSTD then compresses those encoded bytes opaquely, and the block must be decompressed before anything can read it.
solid answer
~50 sA columnar engine compresses in two stages. First it applies a **lightweight, type-aware encoding** to each column chunk: dictionary codes for repeated strings, run lengths for long stretches of the same value, deltas or a per-block minimum plus bit-packing for integers. These encodings understand the column's type, are cheap to decode, and leave values addressable — the engine can often evaluate a predicate against dictionary codes without ever materialising a string. Second, it runs a **general-purpose byte codec** — LZ4, Snappy, ZSTD — over the already-encoded bytes. The codec knows nothing about types, columns or rows; it finds byte-level redundancy and produces an opaque blob that must be decompressed as a whole before any value inside it can be read. The order matters: encoding removes structural redundancy first so the codec has less to chew on. Most of the ratio comes from the encoding; the codec mops up the remainder.
code
text · 4 linescolumn values : US US DE US FR DE DE
dictionary : 0=US 1=DE 2=FR
codes : 0 0 1 0 2 1 1 -- 2 bits per row
then ZSTD/LZ4 over the code stream (opaque bytes)go deeper
Be ready to name the two layers and give one example of each: a dictionary or run-length encoding on a column, and LZ4 or ZSTD over the resulting bytes.
Explain why the encoding runs first, why it is chosen per column and per chunk, and why encoded values can still be filtered while compressed bytes cannot.
Show that you reason about where the bottleneck lands: compression trades I/O for CPU, and after a good scheme analytical scans are often CPU-bound rather than disk-bound.
Own the platform-level position: which layer you standardise (codec) versus which you let the engine decide per column, and how that choice interacts with storage bills, ingest throughput and scan latency.
## Two layers, applied in a fixed order When an analytical engine writes a block of a column — a column chunk, stripe segment or part column, depending on the system — it compresses that block in two independent stages. 1. **Encoding (lightweight, type-aware).** The engine looks at the actual values in this chunk and picks a representation that exploits their structure: a dictionary for repeated strings, run lengths for repeated neighbours, deltas or a shared minimum plus narrow bit-packed offsets for integers. The output is still a structured stream of values — just a cheaper one. 2. **Compression (general-purpose byte codec).** The encoded bytes are then handed to a stock compressor such as LZ4, Snappy or ZSTD, which looks only for repeated byte sequences. Its output is an opaque blob. On read the layers unwind in reverse: decompress the block, then decode (or, better, work directly on the encoded form). ## What each layer knows The encoding layer knows the column's type and its value distribution. That knowledge is what makes it both effective and *transparent*: a dictionary-encoded string column is a small dictionary plus a stream of small integer codes, and an engine can compare, hash, group and filter those integers without reconstructing a single string. Filtering `country = 'DE'` becomes "look up the code for 'DE' in this chunk's dictionary, then scan for that integer" — and if the literal is not in the dictionary at all, the entire chunk is skipped without reading the code stream. The codec layer knows nothing. It sees bytes. That opacity is exactly why it cannot help query execution: to read one value out of a ZSTD-compressed block you must decompress the block. Codecs therefore buy storage and I/O, never query-time shortcuts. ## Why encode before compressing The two layers are complementary rather than redundant. Encoding removes *semantic* redundancy that a byte-level compressor would only partly find — for example, converting 64-bit ids into 5-bit deltas shrinks the stream in a way LZ4 could never infer. It also makes the residual stream more regular, so the codec's job gets easier. Running the codec first would destroy the structure the encoding needs and produce a strictly worse result. In practice a well-chosen encoding contributes most of the total ratio, and the general codec typically adds a further, smaller win on top. ## Choice is per column, per chunk Encodings are chosen **per column**, and usually per chunk, because the right choice is data-dependent: a `status` column with four values wants a dictionary; an `event_id` column wants delta plus bit-packing; a column of random hashes wants neither and is stored plain, leaving the codec to do what little it can. Many engines sample the chunk and pick automatically; some let you declare the encoding in DDL. Either way, one table can carry a different encoding for every column, and the same column can be encoded differently in two chunks whose data differs. The general codec is normally a table- or engine-level setting instead, because it applies uniformly to bytes and its trade-off — CPU versus ratio — is a system-wide operational choice rather than a per-column data property. ## What this buys and what it costs The payoff is more than disk. Compressed data means fewer bytes read from local disk or object storage, fewer bytes over the network in a shared-storage architecture, and more of the working set fitting in cache. The cost is CPU: every scan pays decompression, and after a good compression scheme many analytical scans stop being I/O-bound and become CPU-bound. That is the trade-off that separates the two layers again — encodings decode cheaply and sometimes for free, whereas a heavy codec level shifts the bottleneck onto the CPU. ## How to talk about it A crisp framing for an interview: *encodings are how the engine represents the data; codecs are how it packs the representation.* The first survives into query execution, the second does not. If someone claims that "dictionary encoding is just gzip for columns", that is the misconception to correct — one produces integers you can compute on, the other produces bytes you must unpack first.
- Why is the general codec usually a table-level setting while encodings are chosen per column?The codec sees only bytes, so its trade-off is uniform: CPU spent versus bytes stored. Encodings depend on what is actually in each column — cardinality, sortedness, numeric range — so the best choice differs per column and often per chunk. Engines therefore sample each chunk and pick an encoding, but take the codec from configuration.
- If a column is stored plain because no encoding fits, is the general codec wasted too?No, but expect much less. A codec still finds byte-level repetition, and even high-entropy data usually gives back something. What you lose is the large structural win: with random 128-bit values there is no dictionary, no run and no small delta to exploit, so the total ratio is modest no matter which codec you pick.
- Where does compression actually change query cost, as opposed to storage cost?Every scan reads fewer bytes from disk or object storage and moves fewer bytes over the network, which matters most when compute is separated from storage. Against that, each block costs decompression CPU. Well-compressed analytical scans frequently become CPU-bound rather than I/O-bound, which is why the codec choice is a performance decision, not just a storage one.
saying these in an interview costs you the question
- Treats dictionary encoding and ZSTD as the same thing
- Thinks the general codec runs before the per-column encoding
- Believes a query can filter ZSTD-compressed bytes directly
- Assumes a higher codec level is always the better setting
- Thinks one encoding is chosen for the whole table