skip to content

A colleague replaced a per-row loop by handing that same function body to a column surface. What decides whether the loop is really gone?

level: juniorimportance: must knowfreq 68%

answer

  1. ask what the body is handed
  2. one record, or the whole column
  3. the surface name is not evidence
  4. count how many times it is called
  5. one crossing, or n crossings

basics

~20 s

What the surface hands the body. Called once per record, it is the same loop with a crossing into your code at every record. Called once with the whole column, it is one hand-off and the work stays inside compiled code.

solid answer

~40 s

Handing in a function says nothing about how it runs; the contract of the surface you handed it to says everything. Some surfaces call a caller-supplied body once per value or once per record, so the number of calls is the row count and you have written the record loop again, plus a crossing into your own code and a conversion at each record. Other surfaces hand a body of the same shape the entire column in one call, and the work then happens inside whole-column operations. Across this family those two contracts frequently wear similar surface names, so the name is not evidence. Establish it: count how many times the body is invoked, or look at how many values arrive in its argument.

go deeper

for a junior

Recall that passing a function in is not the same as removing the loop. The question to ask is what the library hands that function: one record, or the whole column.

for a middle

Explain the two contracts and what each costs per record: a crossing into your code and a conversion, paid the row count of times, on top of whatever the body itself does.

for a senior

Demonstrate that you establish the contract by measurement rather than by name - a counter in the body, or the size of the argument that arrives - before you argue about performance.

for a principal

Frame the portability problem: the same written code has different execution models in different tools, so a team standard phrased around a surface name will be wrong the first time the tool changes.

## The claim under the claim `I removed the loop` is a statement about what is **written**, not about what **runs**. A **body you hand in** — a function you write and pass to the library for the library to call — has no cost of its own that you can quote. Its cost comes from the **contract of the surface** you passed it to: what that surface hands the body on each call, and therefore how many calls there are. Three contracts exist across the tools in this family, and they are routinely exposed under similar or identical surface names: - the body is called **once per value**; - the body is called **once per record**, receiving one row's fields together; - the body is called **once**, receiving **the whole column** — a single argument holding every value. Under the first two the call count is the row count. Under the third the call count is one, and whatever the body does to that argument it does with whole-column operations. ## Same word, different argument | What the body receives | Times it is called | Where the per-value work happens | Cost shape | |---|---|---|---| | one value | once per row | in your code | one crossing and one conversion per row | | one record | once per row | in your code | one crossing per row, with more assembled per call | | the whole column | once | inside the operations the body itself uses | one crossing in total | This is why the flat sentence *a function you hand in is called once per record, so it is always slow* cannot travel. It is true of the per-value and per-record contracts and false of the whole-column contract, and a candidate who states it unconditionally is describing one tool and calling it the class. ## How to establish which contract you have 1. **Count the calls.** Increment a counter inside the body and look at it afterwards. One means the whole column arrived; the row count means you are in a loop. 2. **Look at what arrived.** Ask the argument how many values it holds. One value or one record is the loop contract; many values is the whole-column contract. 3. **Compare against the shipped form.** Write the same rule with the library's own whole-column operations over the same data. If the hand-in version lands far behind, it is crossing per record; if it lands beside it, it is not. 4. **Read the contract for that exact surface**, not for its neighbour. Surfaces sitting on the same object, with names one word apart, often differ on this precise point. ## What the per-record contract adds The body's own work is added to the boundary cost, not substituted for it. Per record you pay the call into your code and back out, the conversion of the stored value into something your code can hold and the conversion of what you return, and at the end the collection of every returned value into a column. The smaller the body, the worse the ratio of boundary to useful work — a one-line body is the worst case, not the best. ## The whole-column contract is not automatically fast Handing the body the whole column removes the per-record crossing. It does not stop the body being slow. If the body receives a whole column and then iterates it in your own program, the dispatch per value is back — it has simply moved inside your function where nothing announces it. The honest statement is: **one crossing instead of n, and then the body's own execution is whatever the body's own code is.** ## When the per-record form is the right answer anyway - Exploratory work on a few thousand rows, where the total is a second either way. - Logic that no shipped operation expresses and that you cannot restate as whole-column steps. - A step whose body is genuinely expensive, where the boundary is a small fraction of the work. In all three the professional move is the same: say the row count and the measured time out loud. `Callbacks are slow` and `callbacks are fine` are both positions rather than answers. ## What varies between designs Some designs expose the two contracts under near-identical names and leave the difference to the documentation. Some inspect the body and translate it into their own whole-column operations rather than calling it per value, so the same written code has two different execution models in two tools. And a design with **deferred evaluation** — where writing the expression only records what is to be done and nothing runs until the answer is asked for — can plan around shipped operations but cannot see inside your body, so a hand-in body remains the part of the pipeline it must take at face value. Establish the contract before you price anything.

  • How would you check which contract a surface has without reading any documentation?
    Put a counter in the body and run it over a small table of known size. One invocation means the whole column arrived in a single call; an invocation per row means the loop contract. Inspecting how many values the argument holds answers the same question in one run.
  • If the body is handed the whole column, what cost is left?
    One crossing in total, which is negligible. What remains is the body's own execution: if it uses whole-column operations, it runs as those operations run; if it loops the values inside your program, the per-value dispatch is back, only now it is hidden inside your function.
  • Why is 'never hand a function to the library' a bad heuristic?
    It bans the callback rather than the contract. On surfaces that hand the body the whole column it forbids the clearest expression of a rule for no gain, and it gives no guidance at all on the case that actually costs money - a per-record body over tens of millions of rows.

saying these in an interview costs you the question

  • Says handing a function to the library is always slow.
  • Assumes any column-shaped call runs as one compiled pass.
  • Takes the surface's name as proof of what the body receives.
  • Claims the loop is gone because no loop is written.
  • Prices the body's arithmetic and ignores the per-record crossing.
  • Thinks a whole-column hand-off makes the body's own internals fast.