Your team's review rule says never hand your own function to the data library. As the lead, what is wrong with it and what replaces it?
answer
- ban the contract, not the callback
- cost scales with the row count
- surface names are not portable
- rules must be visible in a diff
- a blunt rule buys worse rewrites
basics
~20 sIt bans the wrong thing. Cost comes from what the surface hands the body and from the row count, not from callbacks as a category. Replace it with a rule about the contract and the scale, plus a documented escape hatch.
solid answer
~50 sThe rule treats a callback as inherently expensive, which is a property of some surfaces and not of others: where a body is handed the whole column, the rule forbids the clearest expression of a rule for no gain. It is also unportable, because the same surface name means different contracts in different tools, so a team that moves tools carries a rule that no longer describes anything. A better standard names the contract and the scale: a body invoked once per record is a record loop and must be justified by row count and measurement; the shipped whole-column operations and a compute-both-sides-then-choose form are tried first; an escape hatch is allowed for logic nothing expresses, with a measured number attached. And it must be checkable in review, or it becomes theatre that pushes people into worse rewrites.
go deeper
Recall that the cost depends on what the body receives and how many rows there are, so a rule that mentions neither cannot be right at both ends.
Explain the two contracts and show why a rule phrased around them still makes sense after a tool change, while one phrased around a surface name does not.
Bring the failure modes a blunt rule causes - guarded sides evaluated everywhere, several whole-length results held at once, transforms too nested to review - and propose the measured escape hatch.
Own the trade between engineer time, machine time and unreviewable code, and make the rule checkable in a diff; a standard reviewers cannot apply consistently is worse than none.
## Why the rule feels right and is still wrong The rule encodes a real observation: a **body you hand in** — a function you write and pass to the library to call — is the commonest cause of a slow single-process transform. But it encodes the observation at the wrong level. The cost is not a property of handing in a function. It is a property of **what the surface hands the body**: once per record and you have a record loop with a boundary crossing at each row; once with the whole column and you have one crossing in total. A blanket ban treats those two as the same thing. Three concrete failures follow: - **It forbids the fast form.** On whole-column contracts, a hand-in body is often the most readable way to name a rule, and it costs nothing measurable. - **It does not bind where it matters.** The rule as written says nothing about row count, so it is equally strict on an exploratory step over four thousand rows and on a nightly pass over fifty million. - **It does not survive a tool change.** Across the ecosystems a team might use, surfaces with near-identical names have different contracts. A rule phrased around surfaces is wrong the day the stack changes; a rule phrased around contracts is not. ## What a good rule looks like | Clause | What it asks of the author | Why it is checkable | |---|---|---| | Name the contract | state what the surface hands the body | the author had to look, and the reviewer can ask for the counter | | Attach a scale | give the row count the step runs at | a number, not an opinion | | Prefer shipped operations | try the library's whole-column operations, and a choice over two computed sides for branches, first | the reviewer can see whether they were tried | | Allow an escape hatch | per-record bodies permitted with a measured time and a one-line reason | the exception is visible instead of hidden | | Require a comparison on rewrite | old and new forms run over the same data, compared with a tolerance for floating-point values | the rewrite cannot silently change the answer | The last clause matters more than it looks. A reorganised pass over floating-point values is not bit-identical to a record-at-a-time accumulation, so a rewrite verified by equality will either fail spuriously or be verified by nothing at all. ## The second-order cost of a blunt rule A rule that people cannot satisfy honestly is satisfied dishonestly. The observed outcomes of a blanket ban are worth naming, because they are worse than the thing banned: 1. **Unsafe rewrites.** A branch that guarded an expression becomes a choice over two computed sides, and the guarded side now runs on the inputs it was protected from. 2. **Memory regressions.** A chain of whole-column steps can hold several whole-length results at once where the record form held one record; the step gets faster and the job starts failing for space. 3. **Unreadable transforms.** Deeply nested choices replace a branch a reviewer could read, and correctness review degrades on a transform that nobody wants to trace. 4. **A rule nobody enforces.** Reviewers cannot see execution in a diff, so a rule stated in terms of outcomes rather than visible artefacts is applied inconsistently and then ignored. ## Where to put the effort instead - **Standardise the alternatives, not the prohibition.** Most per-record bodies in a codebase are the same three or four shapes: a branch, a lookup, a text normalisation, a conditional aggregate. Provide the whole-column expression of each as a shared, reviewed helper and the hand-in body stops being the path of least resistance. - **Make the cost visible automatically.** A timing or row-count line in the pipeline's own output is worth more than a rule, because it tells the author which steps are worth the argument. - **Distinguish exploratory from scheduled work.** Interactive analysis is allowed to be slow; the step that runs every night on the full dataset is not. One rule for both is wrong for one of them. ## What you are actually trading The decision is between engineer time, machine time and the risk carried by a transform people cannot check. A blanket ban optimises machine time only, prices it wrongly, and spends the other two. The defensible position is narrower and harder to write down, which is why it is a lead's job: **per-record bodies are a known cost, they are permitted where the cost is measured and small, and they are the first thing to look at when a single-process step is slow.**
- How would you make the rule enforceable in review rather than aspirational?Require two visible artefacts in the change itself: a line stating what the surface hands the body, and the row count the step runs at. Both are checkable in a diff. Anything that depends on the reviewer imagining the execution will be applied inconsistently and then abandoned.
- A team member says the rewrite is always worth it. What is your counter?Two cases where it is not. Below some column length the fixed cost of each whole-column call exceeds the loop it replaced. And a chain of whole-column steps can hold several whole-length results simultaneously where the record form held one record, so a faster step can turn a working job into one that runs out of space.
- Should the rule differ for exploratory notebooks and scheduled jobs?Yes. Interactive work over a sample is dominated by the analyst's time, and a per-record body that is clear and takes two seconds is the right choice. The scheduled pass over the full dataset is dominated by machine time and by whoever is paged when it overruns, so the bar there is measurement.
saying these in an interview costs you the question
- Defends a blanket ban because callbacks are always slow.
- Writes the rule around surface names rather than contracts.
- Ignores row count, applying one standard to every step.
- Assumes any whole-column rewrite is safe and cheaper.
- Offers no escape hatch, so exceptions become invisible.
- Verifies a rewrite by exact equality on floating-point values.