skip to content

Why does writing one event at a time into a columnar file layout destroy the advantages the layout exists for?

level: seniorimportance: nice to knowfreq 24%

answer

  1. a column block needs many rows
  2. the writer buffers before it can encode
  3. per-file metadata is paid per file
  4. one-row files are mostly footer
  5. compaction restores the layout's economics

basics

~20 s

Every advantage of the layout comes from many rows being written together. One event per file gives column blocks of a single value, no runs for encoders to collapse, per-chunk statistics that rule nothing out, and a footer that can be larger than the data.

solid answer

~40 s

A columnar writer cannot emit anything until it has buffered a chunk of rows, because a column block is built from many rows' values at once. Write one event per file and each column block holds a single value: run-length has no run, a dictionary is as large as the data it describes, and deltas have no neighbour. The per-file metadata is paid in full regardless — a 200-column schema needs a description and statistics per column, easily tens of kilobytes against 1.6 KB of actual values. Scans then pay per-file overhead thousands of times over, since each file costs an open and a footer read before any data is fetched. The standard shape is therefore buffer-then-flush at the writer, and a compaction step that rewrites many small files into few large ones.

go deeper

for a junior

Recall that a columnar file is written in chunks of many rows at once, so a per-record write pattern produces tiny files whose metadata outweighs their data.

for a middle

Explain the buffer-encode-flush sequence and why encoding is a batch transform: run-length, dictionary and delta all need many values before they can save anything.

for a senior

Bring the read-side consequences — per-file requests, useless per-chunk statistics, planning cost over huge file counts — and describe buffering plus compaction as the standard remedy.

for a principal

Own the freshness contract. Decide the flush interval and chunk size against the writer's memory budget and query selectivity, and say which store serves the window the columnar copy has not yet caught up to.

## Where the layout's advantages come from Every benefit of a columnar file is a **statistical** benefit that needs volume: - **Selective reads** work because a column's values for many rows are contiguous, so one request fetches many of them. - **Encoding** works because a block contains enough values for runs, repeated entries and small differences to appear. - **Skipping** works because a chunk covers enough rows that ruling it out saves a meaningful fetch. - **Metadata amortises** because the cost of describing a column is paid once for many rows. A one-row file breaks all four at once, and it breaks them by construction rather than by configuration. ## The write path, step by step 1. The writer accepts rows and **buffers** them, accumulating values per column in memory. 2. When the buffer reaches the configured chunk size, it **encodes each column** over the buffered values, compresses the encoded blocks, and appends them to the file. 3. It records that chunk's per-column offsets and statistics, then starts a new buffer. 4. On close it writes the **footer**: schema, every chunk's column offsets and lengths, and every chunk's statistics. Step 2 is the one that explains the question. Encoding is a transform over a batch of values, so nothing can be written before a batch exists. A writer asked to flush per row will produce a chunk per row, which is a legal file and a useless one. ## The arithmetic of a one-row file Take the 200-column schema again, with roughly eight bytes per value, so a row carries about 1,600 bytes of actual data. The footer must describe 200 columns for the single chunk: type and location, encoding used, compressed and uncompressed lengths, minimum, maximum, null count. At a conservative hundred bytes per column that is about 20 KB of metadata against 1.6 KB of data — roughly 93% of the file is description. Write a million events that way and you have a million files, about 20 GB of which is footer. The read side is worse than the size suggests: - Each file costs at least one request to locate and read the footer before any column block can be fetched, so per-file overhead is paid a million times. - Statistics per chunk cannot help: with one row per chunk they are exact but there are as many of them as there are rows, so evaluating them costs what scanning would have cost. - Encoded runs never form, so the data is stored at close to its raw size and the codec has almost nothing to work with. - Listing and planning over a million objects becomes its own bottleneck, before a byte of data is read. ## What the write path actually does instead - **Buffer and flush.** Accumulate rows in memory and write a chunk when a row count or byte threshold is reached, or when a time-based flush fires so that data does not sit unwritten indefinitely. - **Accept a latency floor.** Freshness is bounded by the flush interval; a pipeline that needs lower latency serves recent data from somewhere else and lets the columnar copy catch up. - **Compact.** A background step rewrites many small files into few large ones, restoring long encoded runs, coarse-but-useful statistics and amortised metadata. Compaction is not housekeeping for disk space; it is what restores the layout's read economics. - **Sort while compacting if it is worth it.** Since the data is being rewritten anyway, clustering it by the column most queries filter on is close to free at that moment and expensive at any other. ## The trade-off chunk size encodes Chunk size is the knob that sits under all of this, and it trades in three directions at once: larger chunks give encoders longer homogeneous runs and amortise metadata better, but they hold more rows in the writer's memory before a flush and make skipping coarser, because ruling out a chunk is all-or-nothing. Smaller chunks give tighter value ranges and finer skipping, at the cost of more footer entries and weaker encoding. There is no universally right number — it follows from the row width, the writer's memory budget, and how selective the typical filter is. The insight worth stating plainly in an interview: a columnar layout is a **batch** format. Asking it to behave like a per-record log is not a tuning problem, and the fix is never a format option — it is putting a buffer, or a different store, in front of it.

  • What does a compaction step actually restore?
    Long homogeneous runs, so encodings and the codec work again; amortised metadata, so the footer is a small fraction of the file; and a small number of files, so scan planning and per-file requests stop dominating. Since the data is being rewritten anyway, it is also the cheapest moment to sort it by the column most queries filter on.
  • How does a pipeline serve fresh data if the columnar write path is batched?
    By separating recent data from historical data. Very recent records are served from a store built for small frequent writes, while the columnar copy trails behind by roughly one flush interval and owns everything older. Queries that need both read from each and combine, and the boundary moves forward as chunks are written and compacted.

saying these in an interview costs you the question

  • Thinks rows can be appended in place to an existing columnar file
  • Assumes many small files cost about the same as one large file
  • Believes a writer can encode a column without buffering many rows
  • Ignores per-file footer and request cost when sizing files
  • Says compaction is only about reclaiming storage space