How do you choose how many rows the reader hands back per turn, and what does too small or too large cost?
answer
- two costs pulling opposite ways
- fixed cost per turn amortises
- peak scales with the batch
- the curve is flat in the middle
basics
~20 sBalance a fixed cost paid once per turn, which is divided over the rows in the batch, against a peak that scales with the batch. Too small is slow, too large stops bounding anything; measure, do not quote a number.
solid answer
~50 sTwo costs pull against each other. Each turn pays a fixed cost — the reader's own per-turn setup, entering the loop body, any per-turn finalisation such as flushing output — that does not scale with the rows in the batch, so it is divided by them; shrink the batch far enough and that cost dominates and the pass can run several times slower. Peak occupancy meanwhile scales with the batch plus whatever the loop body allocates from it, so growing the batch gradually gives back the bound you were buying. Between the two the curve is flat over a wide range, which is the useful fact: this is a plateau, not a knob with a peak. Prune the columns and rows first, time two or three widely spaced values on your own file, take the smallest that sits on the plateau, and never carry the number to another file.
go deeper
Know that the rows per turn is a choice you make, not a property of the reader, and that both very small and very large values have real costs.
Explain the two forces: a fixed cost per turn divided over the rows in the batch, and a peak that scales with the batch. Be able to name what that fixed cost consists of.
Demonstrate that you prune before you tune, leave headroom for what the loop body allocates, and time a couple of widely spaced values on the real file rather than reasoning to a number.
Treat the value as a property of a file and a machine rather than of the code. Hard-coded once, it goes stale silently as columns change, so it belongs where an operator can set it.
## Two forces, pulling opposite ways The rows-per-turn setting is the only number in a pass in pieces, and it is squeezed from both sides. From below, there is a **fixed cost per turn** that does not depend on how many rows the turn contains. It is paid once whether the batch holds a hundred rows or a million, so it is effectively divided by the batch size. Halve the batch and you double the number of times you pay it. From above, there is **peak occupancy**, which does scale with the batch — the batch itself, plus everything the loop body builds out of it while the batch is still alive. That is the quantity the whole exercise exists to bound, so every row you add to a turn gives a little of the bound back. Between the two the pass is insensitive. That is the single most useful thing to know here: the cost curve has a wide flat middle, so you are not hunting for an optimum, you are avoiding two cliffs. ## What the fixed cost is actually made of - the reader's own per-turn work: positioning, locating the record boundary the turn starts on, setting up the batch it is about to hand back; - entering and leaving the loop body, including re-entering an interpreted host language if that is where your body is written; - any per-turn finalisation — combining into a running result, flushing a partial output, opening or closing something; - per-turn decisions the reader repeats, such as a type guess made afresh for each batch. None of these scale with the rows in the batch, which is precisely why a very small batch is a bad trade: you have multiplied the number of times you pay all of them in exchange for a bound you already had. ## Symptoms of each cliff | setting | what you observe | why | |---|---|---| | far too small | the pass is several times slower than one large read, and any per-turn output fragments into many tiny pieces | the fixed cost per turn now dominates, and it is paid far more often | | about right | steady throughput, memory flat across the run, output in sensible units | the fixed cost is amortised and the bound still holds | | far too large | memory climbs until the run fails, often late and in the loop body rather than the read | the batch plus what the body builds from it no longer leaves working room | The far-too-small case has a second, less obvious cost: where the reader decides types per batch, a smaller batch means each decision rests on less evidence, so turns disagree with each other more often than they would at a larger setting. ## Rows are a proxy, not a cost Counting rows per turn is convenient because it is the unit the reader deals in, but a row is not a fixed quantity. Four narrow numeric columns and eighty columns of long text give wildly different batches at the same row count, and the count is only a stand-in for what the batch occupies once it is built — which is not proportional to the bytes the same records occupied on disk. Two consequences follow. A count chosen for one file is not transferable to another, and a count chosen before you pruned the columns is measured against the wrong thing. Where a reader lets you bound a turn by input bytes rather than rows, that is a steadier proxy on data whose rows vary a lot in width, and it still cuts on record boundaries. ## A procedure that finishes 1. **Refuse the columns and rows you do not need at the read.** This changes what a batch costs, so it has to come first. 2. **Pick a starting count that leaves the loop body obvious room** alongside the batch. You are not trying to fill memory. 3. **Time two or three counts an order of magnitude apart** on the real file, on the real machine, with the real loop body. 4. **Take the smallest count that sits on the plateau.** Headroom is worth more than the last few percent of throughput, because the cost of being wrong on the large side is a failed run. 5. **Record why**, and treat the number as configuration rather than as a constant in the code. ## Why the number does not travel A rows-per-turn value that works is a joint property of the file's column widths, the columns you kept, what the loop body allocates, and the machine. Change any of them and the value is stale, silently: the pass gets slower, or the headroom evaporates, and nothing announces it. That is also why a number copied from a tutorial is worth nothing — it encodes somebody else's four variables, and quoting one as though it were a property of the technique is a reliable sign that the candidate has never had to measure.
- Why is a count of rows only a proxy for what you are trying to bound?Rows are not a fixed cost. Four narrow numeric columns and eighty columns of long text produce very different batches at the same row count, so a count chosen for one file is stale for the next. Where a reader can bound a turn by input bytes, that is steadier.
- What should you do before tuning the number at all?Refuse the columns and rows you do not need at the read. Pruning changes what a batch costs per turn, so any count chosen beforehand was measured against a batch you are no longer building.
saying these in an interview costs you the question
- Quotes a row count from a tutorial as the right piece size.
- Assumes a smaller piece is always safer and never slower.
- Treats rows as a fixed cost regardless of how wide they are.
- Tunes the number without timing anything on the real file.
- Ignores what the loop body allocates when sizing the batch.