skip to content

A pass in pieces over a 180-column file needs four of them: what does naming those four at the read change, and does the file's shape matter?

level: seniorimportance: should knowfreq 49%

answer

  1. refuse it at the read
  2. two savings, not one
  3. depends how the file stores columns
  4. parsing still touches every line

basics

~20 s

Naming them at the read avoids converting and keeping 176 columns on every turn. Whether it also avoids reading their bytes depends on the shape: a columnar layout never fetches them, while delimited text still reads and splits every line.

solid answer

~50 s

There are two savings here and they are not the same thing. On re-parsed delimited text, every byte of every line still crosses the parser, because the separators have to be counted to locate the four fields at all; what you avoid is converting 176 fields per row into typed values and holding them. That is usually most of the cost, so it is well worth doing, but it is not reading less. On a self-describing binary columnar file each column's bytes are stored apart and the file's own declaration says where they begin, so the unread columns are never fetched and you genuinely do read less. Either way the saving is per turn, multiplied by the number of turns. Dropping the columns inside the loop body instead saves nothing at all: the reader has already converted them and the peak has already been paid.

go deeper

for a junior

Know that you can name the columns you want at the read, instead of reading everything and dropping the rest afterwards, and that the first is cheaper.

for a middle

Explain that the saving has two parts, bytes read and conversion plus retention, and that a text shape gives you only the second while a columnar shape gives both.

for a senior

Show that you quantify against what the reader built rather than the file's byte count, and that you know a drop inside the loop body has already paid the cost it appears to avoid.

for a principal

The leverage is upstream. The width every consumer is forced to read on every turn is a property of what the producer writes, and narrowing it once pays in every pass that follows it.

## Two savings that get confused for one When somebody says that naming four columns of 180 at the read makes the pass cheaper, they are usually collapsing two different savings: 1. **Bytes not read.** The reader never fetches the storage that holds the unwanted columns. 2. **Work not done and memory not held.** The reader never turns those fields into typed values and never keeps them in the batch it hands you. Which of the two you get is decided by how the file stores its columns, and a candidate who asserts the first one flatly for every shape has learned the habit from one shape and generalised it. ## On re-parsed delimited text Here a record is a line and the fields are positions within it, found by counting separators from the start. To hand you field 7 the parser must walk the line at least as far as field 7, and in practice the whole line is scanned. So: - the bytes are read from the file regardless; - the line is split regardless; - what is avoided is the **conversion** of 176 fields per row into typed values, and the **allocation and retention** of those 176 columns in the batch. That is still a large saving, because conversion and materialisation usually dominate a text read, and because the batch you carry through the loop is now a fraction of the width. But describing it as reading less is wrong, and it leads to the wrong prediction: the pass will not get faster in proportion to the columns you dropped. ## On a self-describing binary columnar file Here each column's values are stored together, separately from the others, and the declaration written into the file says where each one begins. The reader can therefore fetch the four it was asked for and never touch the storage of the other 176. This is the case where naming the columns really does mean reading less, and the reduction shows up in bytes moved as well as in time and occupancy. ## Rows are the weaker lever Refusing rows at the read is a genuinely different proposition from refusing columns: - On a text shape, the line has to be read and parsed far enough to evaluate the condition before the reader can know whether to keep the record. What you save is materialising and retaining the row, not reading it. - Where the read call returns a description of the read rather than the data — a deferred design, in which the plan is built first and executed when you ask for an answer — a condition stated as part of the read can be pushed into it, and more of the work is genuinely skipped. Either way, refusing rows at the read still beats filtering afterwards, for the same reason columns do: the batch that survives into the loop body is smaller. ## What you get, by shape and by timing | what you do | re-parsed delimited text | self-describing binary columnar | |---|---|---| | name the four columns at the read | every line still read and split; conversion and retention of 176 columns avoided | the other columns' bytes are never fetched | | drop the 176 after the read | nothing avoided; the conversion and the peak are already paid | nothing avoided; the bytes are already read | | state a row condition at the read | lines still read and parsed; rows never materialised | avoided at the read on deferred designs; otherwise materialisation only | ## Why the loop cares more than a one-shot read does In a single read the saving is paid once. In a pass in pieces it is paid on every turn, so it multiplies by the number of turns — and it changes the other setting you have. A narrower batch means either the same rows per turn at lower occupancy, or more rows per turn at the same occupancy, which reduces the number of turns and therefore the fixed cost you pay per turn. Pruning and sizing are not independent knobs; prune first, then size. The habit worth breaking is reading everything and dropping the unwanted columns inside the loop body. It feels equivalent and is not: by the time your code sees the batch, the reader has already parsed, converted and allocated all 180 columns, and the peak you were trying to bound has already occurred. Your drop only affects what survives to the next statement. ## Measure the right thing Judge the saving by what the reader built and how long it took, not by the file's size on disk. The two shapes give completely different answers from the same byte count, and on a compressed columnar file the bytes on disk predict neither the time nor the occupancy of the batch you end up holding.

  • Does a row condition applied at the read save as much as a column subset does?
    Usually less on a text shape: the line must still be read and parsed far enough to evaluate the condition, so what you avoid is materialising and keeping the row. Where the read returns a plan rather than data, the condition can be pushed into the read and more is skipped.
  • How would you measure whether the pruning actually helped?
    Against what the reader built, not against the file's size on disk. Time the pass and watch occupancy with and without the column subset; the two shapes give very different answers from the same byte count.
  • Why does pruning change the rows-per-turn setting you chose earlier?
    Because a narrower batch occupies less at the same row count. After pruning you can raise the rows per turn at the same peak, which cuts the number of turns and so the fixed cost paid per turn. Prune first, then size.

saying these in an interview costs you the question

  • Says naming four columns of 180 always means fewer bytes are read.
  • Drops the unwanted columns in the loop body and calls it the same thing.
  • Assumes a row condition at the read avoids reading those lines.
  • Thinks pruning changes nothing because the file is opened anyway.
  • Predicts the speed-up from the file's size on disk rather than measuring.