A read asks for 3 of a file's 200 columns: what does that avoid on a columnar layout, and what on delimited text?
answer
- two different savings, not one
- separators must be found either way
- bytes never read, against bytes read once
- conversion and keeping, on text
basics
~20 sOn a declared columnar layout the unread columns are never read from disk at all. On delimited text every byte of every line still crosses the parser to find the separators; only converting and keeping the unwanted columns is avoided.
solid answer
~50 sBoth savings are real and they are not the same saving. On **a self-describing binary columnar file** — values stored column by column in their declared types, with the declaration written into the file — each column's bytes sit apart and the file records where, so naming the columns you want before the read means the other 197 are never fetched. On **re-parsed delimited text** — one record per line, every value stored as characters — the fields are found by scanning each line for its separators, so every byte is still read and split whatever you asked for; what you skip is converting those fields into values and holding them. The practical difference: on the columnar shape the win scales roughly with the share of columns you dropped, while on text it is bounded by the share of the read's cost that is conversion rather than scanning. Predicting one from the other is how teams end up disappointed.
go deeper
Know that you can name the columns you want before the read rather than dropping them afterwards, and that doing so is cheaper than loading everything and trimming.
Separate the two savings out loud: bytes never read from disk, against bytes read and split but not converted or kept. Say which shape gives which, and why the separators force the difference.
Refuse the column-ratio speedup estimate and say how you would measure the real one, then treat a permanently wide-read pattern as evidence about the stored shape rather than about the read call.
Decide whether the read pattern justifies committing a dataset to a layout that stores columns apart, against the interoperability that gives up, and who pays to maintain the second copy if you keep both.
## What asking for a column subset means Most readers let you name the columns you want **before** the read builds the rest, rather than loading everything and dropping columns afterwards. Those two are not the same operation, and the difference between them is worth having clear before anything else: dropping afterwards means the work was already done and then thrown away, while naming up front means some part of the work is never done at all. *Which* part is never done depends entirely on the shape of the file, and that is the substance of this question. ## On delimited text: the parse still crosses everything In a file of one record per line with separators between fields, there is no way to find field 57 of a line except by walking the line and counting separators. Nothing in the file records where any field starts. So: - **Every byte is still read from disk.** The file is consumed front to back regardless of how many columns you asked for. - **Every line is still split.** Finding your three fields means finding all the separators that precede and follow them. - **What is genuinely saved** is converting the other 197 fields from characters into values, and allocating and keeping them. That saving is not trivial — converting characters into typed values is often the larger share of a text read's time, and not keeping 197 columns is nearly all of its memory — but it has a ceiling, and the ceiling is the cost of the scan itself. ## On a declared columnar layout: the bytes are never read Here the values of one column are stored together, apart from the others, and the file's own declaration records where each column's bytes are. A read for three columns can therefore go to those three and leave the rest untouched. The unread columns are not read and discarded; they are never read. The saving is a saving in **input**, not merely in **conversion**, and that is why it can be an order of magnitude on a wide file rather than a modest percentage. ## The two savings side by side | What happens to the other 197 columns | Delimited text | Declared columnar layout | |---|---|---| | Bytes read from disk | all of them | none of them | | Line scanned for separators | yes, every line | not applicable | | Characters converted to values | skipped | skipped | | Values allocated and kept | skipped | skipped | | Roughly how the win scales | with the conversion share of the read | with the share of columns dropped | ## When the read returns a plan instead of the data Not every design hands you the data when you ask for it. Some return a description of the read, and materialise only the columns and rows that survive the operations you then express against it. The distinction above is unchanged by that — on a columnar layout the unwanted columns are still never read, on delimited text every line is still crossed — but you no longer write the subset yourself, because the pruning is folded into what the engine decides to execute. The thing worth knowing either way is which of the two savings your file shape can actually give you. ## Estimating the win honestly The common failure is arithmetic: "three of two hundred, so sixty times faster". That ratio is a statement about a layout that stores columns separately, and applying it to text produces a prediction nobody can hit. A more honest approach: 1. **Establish the shape first.** Ask whether unread columns can be skipped at all, or only unconverted. 2. **Measure the split.** For a text read, time the same read with and without the subset; the gap is your conversion-and-keep share, and it is the whole of what a subset can buy you there. 3. **Do it anyway on text.** A bounded win is still a win, and the memory effect is usually larger and steadier than the time effect. 4. **If the wide-read pattern is permanent, change the shape.** A pipeline that repeatedly reads three columns of two hundred is telling you something about which shape the dataset should be stored in, not about which flag to pass. The habit this question is really testing is whether you can say *what a saving is made of* rather than *that there is one*. Two operations that look identical at the call site — asking for a few columns of many — do different amounts of work for different reasons, and a candidate who says only "it's faster because it reads less" is right about one shape and wrong about the other.
- Why is a column subset on delimited text still worth asking for?Because converting characters into typed values is usually the larger part of a text read's time, and not building 197 columns is nearly all of its memory. You cannot avoid the scan, but you can avoid almost everything that happens after it.
- What changes if the read call returns a plan rather than the data?The saving is the same in kind but you no longer write the subset yourself: the pruning is folded into what gets executed when you ask for the answer. A columnar layout still never reads the unwanted columns; delimited text still crosses every line.
- Does column order in a delimited file change the cost of reading three columns?Not meaningfully. Separators must be counted from the start of the line to find any field, so the last field is not appreciably dearer than the first and the whole line is scanned either way.
A printed address book against a filing cabinet with one drawer per field. To collect everyone's postcode from the book you still turn every page and read past every name, even though you write down only the postcodes. In the cabinet you open the postcode drawer and touch nothing else. Both save you writing; only one saves you reading.
saying these in an interview costs you the question
- Says naming fewer columns cuts the bytes read on any shape.
- Assumes the saving on text is zero, so never asks for one.
- Treats the subset as a filter applied after loading everything.
- Predicts the speedup from the ratio of columns alone.
- Believes no shape can skip reading bytes for unwanted columns.