skip to content

A per-record body applies one rate above a threshold and another below it. How do you express that branch as whole-column work?

level: middleimportance: should knowfreq 50%

answer

  1. branch becomes three columns
  2. condition, both sides, then pick
  3. the chooser cannot short-circuit
  4. two full-length results held at once
  5. cheap sides win, expensive sides do not

basics

~20 s

Compute both alternatives over the whole column, build the comparison as a condition of one value per position, then choose per position. The branch becomes data instead of control flow, and no caller-supplied body is invoked at all.

solid answer

~40 s

Turn the branch into three whole-column steps and a selection. Compute the high-rate result for every position, compute the low-rate result for every position, compute the comparison once to get a condition column holding one true-or-false per position, then **compute both sides and choose per position** - an expression that picks from the two results position by position. Nothing crosses into your code, so the per-record boundary cost is gone. It is not free: both alternatives are evaluated everywhere, so you do the work of the unselected side too, and each alternative is a **full-length temporary**, a freshly allocated whole-length result that is live at the same time as the other. That trade is excellent when both sides are cheap arithmetic and poor when one side is expensive or unsafe.

code

pseudocode · 14 lines
pseudocode
# 1) per-record: the body is invoked once for every record
function fee_for_one(amount):
    if amount > threshold:
        return amount * high_rate
    else:
        return amount * low_rate

fees = call_body_per_record(amount_column, fee_for_one)   # n crossings into your code

# 2) whole-column: both sides computed everywhere, then chosen per position
high_side = amount_column * high_rate      # full-length result
low_side  = amount_column * low_rate       # full-length result
picked    = amount_column > threshold      # one true-or-false per position
fees      = choose_per_position(picked, high_side, low_side)

go deeper

for a junior

Recall the three pieces: compute each alternative for the whole column, compute the comparison as a column of true and false, then pick per position. No function is handed to the library at all.

for a middle

Explain what the rewrite costs: both sides evaluated at every position and two whole-length results live at once, in exchange for removing every per-record crossing.

for a senior

Show judgment about when not to do it - an expensive or unsafe side selected by a small fraction of rows - and name the middle road of running the expression on only the qualifying positions.

for a principal

Consider the readable limit. Beyond a few cases the nested form is harder to review than the branch it replaced, and a transform nobody can check is a correctness risk that no timing shows.

## Control flow becomes data A branch inside a **body you hand in** — a function you write and pass to the library to call — is control flow: for each record, one side runs and the other does not. The whole-column rewrite deletes the control flow and replaces it with three columns and a selection: 1. **The condition.** The comparison is itself a whole-column operation: it walks the values and yields a column holding one true-or-false per position. 2. **The two alternatives.** Each is computed over every position, producing a complete result. 3. **The choice.** One operation walks the condition and the two results together and emits, at each position, the value from whichever side the condition names. Your program issues a handful of calls instead of the row count of them, and it is never entered during the pass. ## The chooser does not branch This is the part people get wrong, and it is worth stating flatly: **the chooser has no short-circuit.** It cannot skip the side it will not use, because by the time it runs, both sides are already complete columns sitting in memory. Evaluation happens first; selection happens second. A branch says *do one of these*; this construction says *do both, then keep one per position*. ## What the rewrite costs | | Per-record branch | Compute both, then choose | |---|---|---| | Crossings into your code | one per record | none | | Work done | one side per record | both sides at every position | | Peak memory | one record in flight | two full-length results plus a condition | | Behaviour of the unused side | never runs | runs on every value, including ones the branch excluded | So the rewrite trades **wasted computation and peak memory** for **removed boundary cost**. When each side is a multiplication, the waste is a second multiplication over the column and the exchange is overwhelmingly favourable. When one side is a costly computation selected for two percent of the rows, you have made the step roughly fifty times more expensive than it needed to be. ## More than two outcomes Three cases do not need a callback either. Two shapes are available across this family: - **Nest the choices.** Choose between the third alternative and the result of an inner choice; each nesting level adds another condition and another full-length result. - **Use a first-match form** where the tool offers one: a list of condition-and-result pairs evaluated per position, taking the first condition that holds, with a fallback for positions no condition matched. This reads far better than deep nesting once there are more than three cases. Watch the fallback either way. A position matched by no condition has to get *something*, and the default in several designs is the **absent-value marker** — whatever that tool stores where there is no value — which then travels silently into everything downstream. ## Where the operands come from matters The three operands of the choice have to line up, and how they line up depends on what they are: - Where all three carry row labels, they are aligned on those labels before the choice, and a position present on only one side yields absence rather than a value. - Where they are bare positional buffers, they match strictly by position and their lengths must agree; a mismatch is an error rather than a fill. - A mixture of the two is where quiet wrongness lives: a condition that was derived from a differently ordered or differently filtered column can be correct in length and wrong in meaning. The practical guard is to build all three from the same column in the same expression, rather than assembling them from variables created at different points in the program. ## When to keep the branch The rewrite is not doctrine. Keep a per-record body when one alternative is expensive and rarely selected, when one alternative is unsafe to evaluate on the rows the branch was protecting it from, or when the branch is deep enough that the whole-column form stops being readable. In the first two cases there is a middle road that is usually better than either extreme: select the qualifying positions first, run the whole-column expression on that shorter column, and place the results back — two passes over subsets, no crossings, and the unsafe or expensive side never sees the rows it should not.

  • How do you extend this to four outcomes without deep nesting?
    Use a first-match form where the tool has one: a list of condition-and-result pairs, evaluated per position, taking the first condition that holds. Give it an explicit fallback, because positions matched by no condition otherwise receive the absent-value marker and carry it downstream.
  • The rewrite doubled peak memory. Is that expected?
    Under eager evaluation, yes: each alternative is a complete result and both are live when the choice runs, alongside the condition. Under deferred evaluation, where the expression is recorded and run once, the intermediates need never be materialised, so the same code has a different memory profile.
  • One alternative is costly and selected for two percent of rows. What do you do?
    Do not compute it everywhere. Select the qualifying positions, evaluate the costly expression on that much shorter column as a whole-column operation, and place the results back into a result initialised with the cheap side. No crossings, and the costly work is proportional to the rows that need it.

saying these in an interview costs you the question

  • Assumes the chooser skips the side it will not use.
  • Says the rewrite is free because no extra code was written.
  • Ignores that both alternatives are computed at every position.
  • Forgets the fallback for positions no condition matched.
  • Builds the condition from a differently ordered column.
  • Treats the rewrite as mandatory regardless of cost per side.