skip to content

After a per-record branch was rewritten to compute both alternatives and choose per position, the step now warns on rows the branch never reached. What happened?

level: seniorimportance: should knowfreq 42%

answer

  1. the branch was also a guard
  2. no short-circuit in a choice
  3. both sides meet every row
  4. warn quietly or fail loudly
  5. make the side total, or subset first

basics

~20 s

The branch was a guard. It stopped one expression from ever seeing the rows it could not handle. The chooser has no such power: both alternatives are evaluated at every position first, so the guarded side now runs on exactly the inputs it was protected from.

solid answer

~40 s

In the per-record form, the comparison decided whether the risky expression ran at all, so it never met a zero denominator, a non-positive input or an out-of-range key. In the whole-column form, both alternatives are computed over every position and only then is one picked per position, so the risky expression meets every row. The discarded values are thrown away afterwards, which is why the final answer can still be correct while warnings appear - and why, where the operation raises instead of yielding the absent-value marker, the entire pass fails rather than a single record. The repair is to make the guarded side total: substitute a harmless stand-in into the offending operand before computing that side, or evaluate that side on only the qualifying positions and place the results back.

go deeper

for a junior

Recall that computing both alternatives means the risky one runs on every row, including the rows the original branch protected it from.

for a middle

Explain the ordering: evaluation of both sides completes before any selection happens, so a discarded value was still computed and any warning it raised was still raised.

for a senior

Diagnose it properly - evaluate each side alone, count the offending positions against the excluded ones - and repair it by making the guarded side total or by evaluating it on a subset, not by silencing the warning.

for a principal

Turn it into a review rule: a branch may be rewritten as a choice only when both alternatives are defined for every value in the column, and the rewrite is validated against the old form on real data.

## The guard was doing two jobs A branch inside a **body you hand in** — a function you wrote and passed to the library to invoke — does two things at once, and people only notice the first: 1. it **selects** which value ends up in the result; 2. it **guards** one expression from inputs it cannot survive. The whole-column rewrite — **compute both sides, then choose per position** — reproduces job one exactly and abandons job two silently. Both alternatives are evaluated in full before anything is chosen, so a division runs where the denominator is zero, a logarithm runs on a non-positive value, a lookup runs on a key that is not present, an index runs past the end. The chooser then discards exactly those results, which is why the delivered numbers often look right while the run is noisy. ## Two failure modes, and they are not equally survivable | How the operation reacts to a bad input | What you see | Risk | |---|---|---| | yields the **absent-value marker** — whatever the tool stores where there is no value | a warning, or nothing at all | silent: the marker may be discarded by the choice, or may leak in through a later step | | raises | the whole pass fails on the first bad position | loud, but it fails the entire column, not one record | The loud version is the lucky one. The dangerous version is the quiet one, because the same input that produced the marker may also have produced an infinity or a value far outside the sensible range, and only the chooser stands between that and the result. Change the condition slightly in a later edit and the poison is no longer discarded. ## Establishing what actually happened 1. Identify **which alternative** warns, by evaluating each side on its own over the full column and watching which one complains. 2. Count the offending positions and compare that count with the count of positions the condition excludes. If they coincide, the branch was a guard and this is exactly the diagnosis. 3. Check whether the delivered result is still correct. It usually is, which explains why this ships: the defect is in the run, not yet in the numbers. ## Three repairs, in order of preference 1. **Make the guarded side total.** Substitute a harmless value into the operand that causes the trouble, compute the side over everything, then choose. Replace zero denominators with one, clamp the input to the domain the operation accepts, point out-of-range keys at a placeholder. Correctness then comes from the choice, and the stand-in values are only ever discarded. State the stand-in explicitly in the code so a reader sees that it is deliberate. 2. **Evaluate the risky side only where it is needed.** Select the qualifying positions, run the expression as a whole-column operation on that shorter column, and place the results back into a result initialised with the safe side. This keeps the boundary cost at zero, does work proportional to the rows that need it, and never exposes the expression to the inputs it cannot handle. It costs one more step and an alignment you must get right. 3. **Keep a per-record body for that one step**, knowingly, and pay the crossings. This is the correct answer when the guarded side is both expensive and genuinely impossible to express safely, and when the row count is small enough that the crossings do not matter. Say the number out loud rather than defending it on principle. ## What this says about the rewrite in general The rewrite from a branch to a choice is a good default and should stay a default. What it is not is a mechanical transformation: the two forms have the same result only when both alternatives are **total** — defined for every input in the column, not merely for the inputs the branch admitted. That is the property to check before rewriting, and it is a better review question than any rule about callbacks. It also means the rewrite should be validated on data, not by reading. Run both forms over the same input and compare, and compare with a tolerance rather than an equality where the values are floating-point, because a reorganised pass can differ in the last digits from the record-at-a-time accumulation it replaced. ## A note on designs that differ Whether a bad input warns, yields the absent-value marker or raises is a property of the operation and the tool, not a universal. Some designs treat a zero denominator on integers as an error and on floating-point values as a special value; some surface a warning once per run rather than once per row, which is why a noisy step can look quiet in a log. Do not assume the behaviour you have seen is the behaviour everywhere — check what your operation does on the bad input before deciding which of the three repairs you need.

  • The numbers came out identical, so why fix anything?
    Because correctness is resting on the chooser discarding poisoned values. A later edit to the condition, or a step that reads the intermediate, exposes them. There is also the wasted computation and, where the operation raises rather than warns, a run that fails entirely on the next bad input.
  • How do you substitute a harmless value without changing the answer?
    Only into positions the choice will discard. Build a safe operand by replacing the offending values - zero denominators with one, out-of-domain inputs with a value inside the domain - compute that side from the safe operand, then choose using the original condition. Every substituted result is thrown away.
  • When is the subset approach better than making the side total?
    When the risky side is also expensive. Making it total still computes it everywhere; evaluating it on the qualifying positions does work proportional to those rows only. The cost is one extra step and an alignment between the subset's positions and the result's.

A branch is a door you only open if you need the room. This construction opens both doors, looks inside both rooms, and then decides which one you wanted. If one room is on fire, opening the door was the problem, and choosing the other room afterwards does not put it out.

saying these in an interview costs you the question

  • Believes the chooser skips evaluating the unselected side.
  • Says the result is fine, so nothing needs fixing.
  • Suppresses the warning rather than removing its cause.
  • Assumes a bad input always raises rather than yielding a marker.
  • Reverts to a per-record body without pricing the alternatives.
  • Rewrites a branch without checking both sides are defined everywhere.