skip to content

A diff of two transform outputs flags 4,000 cells where both sides hold no value at all. Why, and what must the comparison state?

level: seniorimportance: should knowfreq 44%

answer

  1. absence does not equal itself
  2. the answer depends on the representation
  3. false, or no verdict at all
  4. one side absent is a real difference
  5. state the rule, do not inherit it

basics

~20 s

Absence is not a value that equals itself. Depending on the design, comparing two absent cells yields false, or yields absence rather than a verdict — neither is "equal" — so a diff has to state its rule for absent cells rather than inherit one.

solid answer

~40 s

In designs that mark absence with the sentinel borrowed from fractional arithmetic, that marker compares unequal to everything, including another copy of itself, so an element-wise comparison of two identical outputs returns false on every cell where both sides are absent. In designs carrying a separate marker bit alongside the values, or a first-class absent value, a comparison involving absence yields absence rather than true or false — three-valued logic — and a reducer that demands true then treats it as a failure too. Serious table comparators expose "two absent cells count as equal" as an explicit setting, and different designs default it differently. So state the rule. And keep it narrow: a cell absent on **one** side only is a genuine difference and must stay in the report.

go deeper

for a junior

Recall that a cell holding no value does not compare equal to another cell holding no value, so a diff has to be told what to do with them.

for a middle

Explain the two representations and their two different answers — false in one, no verdict at all in the other — and why neither of them is "equal".

for a senior

Separate the three buckets and keep the one-sided absences out of whatever setting makes two absences agree; a rewrite that emptied a column must not pass.

for a principal

Make the absence rule part of the team's stated acceptance terms alongside the key and the tolerance, so a passing diff means the same thing to everyone who reads it later.

## Absence is not a value A cell where no value exists is not zero, not an empty string, and — the point here — **not something that equals itself**. The moment a diff crosses data containing absences, it needs a rule, and the rule cannot be inherited from the comparison operator because the operator's behaviour is a property of the design, not of logic. ## Two designs, two wrong answers | how absence is represented | what comparing two absent cells gives | what a naive diff concludes | |---|---|---| | the sentinel borrowed from fractional arithmetic, whose defining property is comparing unequal to every value including another copy of itself | false | "these cells differ" | | a separate marker bit beside the values, or a first-class absent value, under three-valued logic | absence — neither true nor false | "not proven equal", which a reducer demanding true counts as a failure | Neither design reports **equal**, and that is why 4,000 identical absences show up as 4,000 differences. The learner who has only met one of these designs will write a diff that is wrong on the other, which is exactly why the rule has to be written down rather than assumed. ## Three buckets per cell, not two The useful diff classifies every compared cell into one of three states, and only the first is negotiable: 1. **Both sides absent.** Usually the rewrite reproduced the old behaviour exactly, and this should count as agreement — but say so explicitly. 2. **One side absent, the other holding a value.** A real difference, every time. This is the interesting bucket: it is the signature of a step in the new version that introduced or removed an absence — a condition no absent value can satisfy, a lookup that found nothing, a conversion whose failures were turned into absences. 3. **Both sides holding values.** Compare on the stated tolerance. A setting that makes two absences equal must not be allowed to leak into bucket 2. A comparator that merely "ignores absences" may do exactly that, and then a rewrite that emptied a column passes the diff. ## Say the rule out loud A diff's terms are three sentences, and this is one of them: - the columns the two outputs are aligned on, so row order stops mattering; - how far apart two numbers may be before they count as different; - **what happens where one or both sides hold no value.** Write all three beside the diff. A reader six months later cannot reconstruct any of them from the verdict, and the third is the one nobody thinks to ask about until a column of absences slips through. ## The subtraction trap The informal way to diff two outputs is to subtract one from the other and look for non-zero cells. Absence defeats it quietly in both directions: - Absent minus anything is absent, so **a cell absent on one side only produces an absent result rather than a large number** — and a scan for non-zero cells never sees it. The most interesting bucket is exactly the one that disappears. - Absent minus absent is also absent, so whether that counts as agreement depends entirely on how the result is reduced: counting non-zero cells treats it as agreement, while asking "are all cells zero?" can answer false or answer absence. So the subtraction reports the wrong thing on both of the buckets that involve absence, and it reports them the same way. Use a comparison that takes an explicit absence rule, or build the three buckets yourself after aligning on the key. ## What the report should say - The count of cells in each of the three buckets, per column. - For bucket 2, a sample of actual rows with the key, the value present on one side, and which side it was on — because the direction tells you whether the new version added absences or removed them. - The rule that was applied, restated in the output rather than buried in the code that produced it, so that a later reader knows what "passed" meant.

  • Why is subtracting the two outputs a poor way to find cells that are absent on only one side?
    Because absence propagates through the arithmetic: absent minus a value is absent, not a large number, so a scan for non-zero cells never reports it. The bucket that most needs attention — a value on one side and nothing on the other — is exactly the one the subtraction erases. Build the buckets from an explicit comparison instead.
  • Is "treat two absent cells as equal" always the right setting for a rewrite diff?
    It is the right default when the claim is that the new version reproduces the old behaviour, including where it produced nothing. It is wrong when absences are themselves the thing under review — for instance when the rewrite was meant to eliminate them — and it is always wrong if the setting also swallows cells absent on one side only. Check which of the two the tool does.

saying these in an interview costs you the question

  • Assumes two absent cells compare equal by default everywhere
  • Believes an absent cell is the same as zero for comparison purposes
  • Subtracts the outputs and scans for non-zero cells to find differences
  • Lets an ignore-absences setting also hide one-sided absences
  • Reports the absent-cell count without saying which side was empty