A function you hand the library does one multiplication per record yet dominates runtime over ten million rows. What is paid at each record?
answer
- fixed cost per crossing
- unwrap in, wrap out, place the result
- the ratio worsens as the body shrinks
- multiply it by the row count
- shipped operators cross once
basics
~20 sThe boundary, not the arithmetic. Each record pays a call into your code and back, a conversion of the stored value into something your code can hold, a conversion of what you return, and a place in the collected result.
solid answer
~50 sThe multiplication is not the cost; the crossing is. On a surface that calls a caller-supplied body once per record, every record pays a call into your code and back out, plus a conversion of the raw stored value into a **boxed value** - a value wrapped as a full language object with its own header - so your code can touch it, plus the reverse conversion of what you return, plus placing that return into the result being assembled. None of that shrinks because your body is a single multiplication, so the ratio of boundary cost to useful work is at its worst precisely when the body is trivial. Expressing the same rule with operations the library ships removes the crossings entirely: one dispatch over the packed values instead of the row count of them.
go deeper
Recall that a step can be slow because of how often your code is entered, not because of what it computes. Ten million small crossings cost more than ten million multiplications.
Explain the four per-record costs - the call, the conversion in, the conversion out, the placement - and why a one-line body makes the useful fraction of the time smaller rather than larger.
Show that you measure before rewriting: the hand-in version against the shipped-operator version on the same data, so the decision rests on a ratio rather than on a rule of thumb.
Weigh the rewrite honestly. The gain is real only where the row count is large and the body is small, and a rewrite that nobody on the team can read has costs that do not appear in the timing.
## Price the boundary, not the body The instinct when a step is slow is to look at what the step computes. Here that instinct points at a multiplication, finds nothing, and stalls. The cost is one layer out: on a surface that calls a **body you hand in** — a function you wrote and passed to the library — once per record, the library and your program exchange control the row count of times, and each exchange has a fixed price that is independent of what the body does. Per record, that price is: - **the call itself** — setting up a call frame into your code and returning from it; - **the conversion inward** — the value sits in the column as raw bytes of one shared representation; your code cannot touch raw bytes, so it is materialised as a **boxed value**, a full language object with a header and a reference to it; - **the conversion outward** — whatever your body returns is a language object too, and it has to be read back as a raw value of some representation; - **the placement** — that returned value has to be put into the result being assembled, and the result's own representation is decided by what came back. ## Why a small body makes the ratio worse, not better The boundary cost is fixed per record. The body's cost is whatever you wrote. So the fraction of time spent usefully is roughly *body / (body + boundary)*: | Body's own work per record | Share of time that is useful | Practical reading | |---|---|---| | one multiplication | tiny | the step is a boundary-crossing benchmark with a multiplication attached | | a handful of arithmetic steps | still small | rewriting as whole-column operations is almost free money | | a genuinely expensive computation | dominant | the boundary is noise; leave it and optimise the body | Candidates often reason the opposite way — *the body is trivial, so the step must be cheap*. The trivial body is the worst case for this contract, because the fixed cost has nothing to amortise against. ## What the whole-column form removes Expressing the same rule with operations the library ships — **a whole-column operator**, an operation that takes the whole column and loops over the values inside compiled code so your program issues one call instead of the row count — removes every one of the four costs above. There is one dispatch, the values are read as raw bytes in place, the result buffer is allocated once, and your program's code is never entered during the pass. That is what **column-wise** work means: saying the operation once for a whole column rather than once per row. It is worth being precise about what it does *not* remove. A loop still exists; it lives inside the library's compiled pass and it never goes away. What went away is the crossing, the conversion and the per-record dispatch. ## The two numbers to bring to the conversation 1. **The row count.** The whole effect is a fixed cost multiplied by the number of records. At a few thousand records the crossings are invisible; at ten million they are the job. 2. **The measured gap.** Time the hand-in version and the shipped-operator version over the same data. A ratio you measured beats an argument about mechanisms, and it also protects you from the opposite error — rewriting a step that was never the bottleneck. ## Where designs differ This cost model belongs to the per-record contract, not to callbacks as a category: - On surfaces that hand a body of the same shape **the whole column**, there is one crossing in total and the model above simply does not apply. - Some designs inspect a hand-in body and translate it into their own whole-column operations instead of invoking it per value, so identical written code has two different cost profiles in two tools. - Storage matters too. Where a column is already held as **boxed values** rather than packed bytes, the inward conversion is cheaper because there is nothing to unpack — but then the shipped whole-column form is slower as well, so the *gap* between the two narrows rather than the hand-in form becoming good. ## A by-product worth noticing The library cannot see into your body, so it cannot know what the body will return. It finds out by collecting the returns. Where those returns are not uniformly one narrow representation, the result lands in the general boxed form — and every downstream operation on that column then pays per value too. A slow step can therefore make its successors slow, which is why the timing sometimes points at the step *after* the one that caused it.
- Does the same reasoning hold if the body does something genuinely expensive per record?No, and that is the useful boundary. When the body's own work dwarfs the crossing, the crossing is noise and rewriting it as whole-column steps buys almost nothing. Optimise the body, or reduce how many records reach it, instead.
- The step got faster after the rewrite, but the step after it got slower. How?Look at what the hand-in step used to return. Because the library cannot see into a body, the result's representation is inferred from what came back, and a non-uniform set of returns lands the column in the general boxed form. Downstream operations then pay per value.
- Why does the same body cost noticeably less over ten thousand rows?Because the cost is a fixed price multiplied by the row count. At ten thousand crossings the total is below the noise of everything else in the step, so the choice is free and should be made for readability instead.
Each record is passed through a service window: a form is filled in on the way out, another on the way back, and both are filled in whether the parcel is a feather or a piano. Ten million small parcels is ten million pairs of forms, and the paperwork, not the parcels, is the day's work.
saying these in an interview costs you the question
- Says the step must be cheap because the body is one line.
- Blames the arithmetic rather than the boundary crossing.
- Thinks the crossing cost shrinks when the body is simplified.
- Believes the cost is the same at ten thousand rows and ten million.
- Claims callbacks as a category are slow, whatever they are handed.
- Expects a faster body to fix a step dominated by crossings.