skip to content

How does an on-disk columnar file format differ in purpose from an in-memory columnar layout that engines share?

level: middleimportance: should knowfreq 38%

answer

  1. same shape, two different jobs
  2. bytes at rest versus values under compute
  3. encoded blocks versus fixed-width buffers
  4. validity bitmap keeps slot width constant
  5. an interchange layout is not an archive

basics

~20 s

Both group values by column, but for opposite goals. An on-disk format such as Parquet or ORC minimises bytes at rest with encoding, compression and chunk statistics. An in-memory layout such as Arrow keeps fixed-width buffers that compute can address directly.

solid answer

~50 s

The shared idea is columnar grouping; the optimisation targets differ. A file format optimises **bytes at rest and bytes fetched**: values are dictionary, run-length or delta encoded, blocks are compressed, and a footer carries per-chunk ranges so readers can skip. All of that removes fixed-width positional access, so a value cannot be reached without decoding. An in-memory interchange layout optimises **the representation compute actually runs on**: fixed-width contiguous value buffers with a separate validity bitmap for nulls, typically uncompressed so the `i`-th value is an offset calculation and processing kernels can run without a decode step. It is also a shared shape rather than a private one, so different engines and processes agree on the buffers instead of re-encoding between them. A pipeline normally uses both: the file for durability and size, the memory layout for the compute in between.

go deeper

for a junior

Recall that columnar appears in two places for two reasons: a file on storage built to be small and selectively readable, and an in-memory arrangement built so computation can work on values directly.

for a middle

Explain why encoding and compression rule out positional access, and why an interchange layout keeps fixed-width slots plus a validity bitmap so the element at a given index is an offset calculation.

for a senior

Bring the pipeline consequences: expanded memory footprint against encoded file size, where conversion costs land, and why a shared buffer shape removes conversions between components that never needed them.

for a principal

Frame it as a boundary standard. Decide which representation each stage owns, what agreeing on a public in-memory shape buys across teams, and what you accept in memory headroom to get it.

## Same grouping, different optimisation target Columnar is a grouping strategy, not a single design. Two quite different artefacts adopt it for different reasons, and confusing them is one of the more common mistakes in a data interview. An **on-disk analytical format** — Parquet and ORC are the standard examples — exists to make a dataset small at rest and cheap to fetch selectively. Every design decision follows from that: encode each column against its own domain, compress the encoded blocks, cut the data into chunks, and write a footer that lets a reader fetch only some chunks and only some columns. An **in-memory columnar layout** — Arrow is the standard example — exists to be the representation that computation runs over and that separate engines and processes can agree on without converting. Its design decisions follow from that instead: keep each column as a contiguous buffer of fixed-width slots, keep nulls in a separate validity bitmap so the slot width never varies, and leave the values uncompressed so the `i`-th element is a pointer arithmetic step rather than a decode. ## Where the designs diverge | Concern | On-disk analytical format | In-memory interchange layout | |---|---|---| | Optimised for | Bytes at rest and bytes fetched | Direct access and cross-engine sharing | | Value encoding | Dictionary, run-length, delta, then a codec | Fixed-width slots, typically uncompressed | | Reaching element `i` | Decode the block that contains it | Offset arithmetic into a buffer | | Nulls | Recorded per block, often a count plus flags | A separate validity bitmap, one bit per slot | | Metadata | Footer with schema, offsets and per-chunk ranges | Schema plus buffer layout description | | Lifetime | Durable, versioned, read for years | The lifetime of a process or a transfer | | Size | Much smaller than the values it holds | Comparable to or larger than the raw values | The row of that table that carries the most weight is the third one. Encoding is exactly what makes a file small, and it is exactly what makes positional access impossible: in a run-length encoded block there is no fixed offset at which the millionth value lives, and a dictionary-encoded block holds codes rather than values. A compute layer that wants to sum a column as a dense array therefore cannot work on the file's bytes directly, no matter how they are cached. ## The interchange half of the story The second purpose is easy to miss because it is organisational rather than mechanical. Historically, every engine held columns in its own private shape, so moving a result from one to another meant encoding it into some format and decoding it back — a cost paid for no computational reason, twice, on data that never left the machine. A **standard** in-memory layout removes that: if two components agree on the buffer shape, handing data across is handing over the buffers. That is why such a layout is specified publicly at all rather than being one engine's internal detail, and why it defines things a private representation would never bother to define, such as exact bitmap semantics and buffer alignment. Reading values in place out of a shared buffer is its own subject, and this leaf stops at the goal rather than the technique. ## Practical consequences to get right - **Neither replaces the other.** A pipeline reads the on-disk format, decodes into the in-memory layout, computes, and writes the on-disk format back out. Each is doing a job the other is bad at. - **The interchange layout is not an archive.** It is designed for a version of the data that is alive right now. Storing years of history in an uncompressed fixed-width form pays enormous storage costs for no read benefit, and durability guarantees are not what it was designed around. - **The file format is not a compute format.** Trying to run kernels straight over encoded blocks means reimplementing decode inside every operation; the exception is a filter evaluated against dictionary codes, which works precisely because it needs no positional access. - **Compression can appear on both sides, differently.** Ecosystems differ here: some transports compress interchange buffers on the way across a boundary and decompress them on arrival, which keeps the compute layout uncompressed while shrinking the transfer. That is a transport choice and does not change the layout compute runs on. - **Sizes are not comparable.** A dataset that is a few hundred gigabytes as encoded files can be several times larger once expanded into fixed-width buffers, which is why memory limits are computed against the expanded size. The short version an interviewer wants: columnar on disk is an answer to "how few bytes must I fetch and store?"; columnar in memory is an answer to "what shape should the values be in while they are being computed on, and can everyone agree on it?".

  • Why does the in-memory layout keep nulls in a separate bitmap instead of a sentinel value?
    Because a sentinel would steal a legal value from the domain and force a comparison on every element. One bit per slot in a side buffer keeps the value slots fixed-width, so element `i` is still pure offset arithmetic, and lets a kernel check whole words of the bitmap at once when deciding whether a range has any nulls at all.
  • Why write the on-disk format at all if everything is already in memory?
    Durability, size and selectivity. The file survives the process, is several times smaller thanks to encoding and compression, and carries per-chunk statistics so a later query can fetch a fraction of it. An uncompressed fixed-width buffer has none of those properties and is not designed to be read years later.

saying these in an interview costs you the question

  • Says the two are the same format with different file extensions
  • Believes the in-memory layout is the file's bytes loaded unchanged
  • Claims the in-memory form is smaller because it is columnar
  • Treats the interchange layout as a durable archive format
  • Thinks compute kernels can run over encoded blocks unchanged