skip to content

A comparison of two transform outputs reports them unequal while every number on both sides agrees. What else can it be objecting to?

level: middleimportance: should knowfreq 46%

answer

  1. one verdict, several axes
  2. values, representation, structure
  3. a widened column still holds the same numbers
  4. labels can make subtraction useless
  5. say which axis before blaming arithmetic

basics

~20 s

A strict comparison answers about three things at once — the values, the representation each column is held in, and the structure (column names and order, and whether each side carries row labels). One verdict hides which of the three moved.

solid answer

~50 s

A table comparison is really three comparisons collapsed into one verdict. The **values** axis asks whether the cells agree. The **representation** axis asks whether each column is physically held the same way — the same numbers held once as whole numbers and once as a wider fractional representation are equal as numbers and unequal to a strict comparator. The **structure** axis asks whether the column names and their order match, and whether each side carries per-row labels. Strict comparators check all three by default and say so with one boolean; others compare values only, and several let you switch each axis off. So before concluding the rewrite changed the numbers, find out which axis the objection is on — the usual answer is that a step in the new version re-represented a column, not that any arithmetic differs.

go deeper

for a junior

Recall that a comparison can object to how a column is stored or what the columns are called, not only to the numbers inside them.

for a middle

Explain the three axes and give a cause for each: an absence that widened a column, a rename, a reset of the per-row labels. Then show how you would isolate the axis before investigating.

for a senior

Demonstrate that comparison semantics are a property of the tool's design, not of arithmetic, and that a widened identifier column is a real defect rather than a cosmetic one.

for a principal

Decide which axes the team's acceptance standard actually requires: values within tolerance may be enough for a report, while a consumer that reads columns by position or matches on an identifier makes structure and representation part of the contract.

## One verdict, three separate questions A comparator that reports two outputs "not equal" is usually answering three questions at once and reporting the disjunction: 1. **Values** — do the cells hold the same numbers, text and absences? 2. **Representation** — is each column physically held the same way on both sides: the width and kind chosen once for the whole column and paid for on every row? 3. **Structure** — are the column names the same, in the same order, and does each side carry the same per-row labels? A rewrite can trip 2 or 3 without touching 1. That is the common case, and a candidate who only knows the values axis reads a structural objection as a numerical regression and goes looking for a bug that does not exist. ## The representation axis A column carries one representation for all its rows, and an operation can change it without changing what the numbers mean: - A step that introduces a cell with no value re-represents the column in designs whose absence marker is the sentinel borrowed from fractional arithmetic — there is no room for that marker in a narrow whole-number column, so the whole column widens. In designs that carry a **validity bit** — a separate marker alongside the values recording whether each cell holds one — the column keeps its narrow width and its exact values. **The same rewrite therefore changes the representation on one design and not on another.** - A round trip through a file, or a reader that inferred from a different sample, can hand back a column held differently. - An operation that mixes two columns of different widths generally produces the wider of the two. On the values axis these outputs agree. On the representation axis they do not, and a strict comparator says "not equal" with no hint that the numbers matched. ## The structure axis Two further objections live here: - **Column names and their order.** A rewrite that renames a column, or emits the same columns in a different order, is reported as a difference by comparators that check names and order, and passes on comparators that match columns by name only. Which of those you are using is a property of the tool, not of the data. - **Per-row labels.** Some tools carry a per-row identity beside the values and quietly line operations up on it. If one output carries those labels and the other does not, or they carry different ones, a structural comparator objects — and an element-wise subtraction of the two outputs does something worse: it matches the operands on those labels first, so a re-labelled output produces an all-absent result that reads as total disagreement even though every value is identical. On a plain rectangle of numbers with no labels at all, the same subtraction matches strictly by position instead. **Whether an element-wise operation lines its operands up by label or by position is a property of the object, not of arithmetic.** ## Reading the objection | the objection is on | typical cause | what it means for the rewrite | |---|---|---| | values | arithmetic, ordering of operations, a changed rule | a real difference to investigate | | representation | a step introduced an absence, a round trip, a widening operation | numbers agree; decide whether the new width is acceptable | | column names or order | a rename, or columns emitted in a different sequence | cosmetic unless a consumer depends on position | | per-row labels | one side reset or re-derived its labels | cosmetic for the diff; fatal for element-wise subtraction | ## What to do about it - **Ask the three questions separately.** Compare values with the structural checks relaxed, and check names, order and representation as their own assertions. Then a report says which axis moved. - **Never conclude "the numbers changed" from a bare verdict.** State which axis the reported difference is on before anything else. - **Decide deliberately whether a representation change is acceptable.** Often it is not, even when the values print the same: a widened identifier column can lose its low digits past the exactly representable range, and a later match on that column then fails to find rows it used to find. - **Assert the representation after any step that can introduce absence**, rather than assuming either behaviour — because the behaviour differs by design, and the assertion costs nothing next to the diff itself.

  • Why can subtracting one output from the other be a misleading way to find differences?
    Because on label-carrying objects the subtraction lines the two operands up on their per-row labels first. Rows in a different order still match correctly, but a label on one side only yields an absent result rather than a difference, and a re-labelled output produces an all-absent result that looks like total disagreement. On unlabelled rectangles the same expression matches strictly by position instead.
  • Is a column that quietly widened a cosmetic difference?
    Not always. The values usually print identically, but a wider fractional representation cannot hold large whole-number identifiers exactly past a point, so the low digits go and a later match on that column stops finding rows. Treat a representation change on a key or identifier column as a real finding and on a measure column as a decision to make explicitly.

saying these in an interview costs you the question

  • Reads any not-equal verdict as proof the arithmetic changed
  • Believes two outputs holding the same numbers must compare equal
  • Thinks column names and order are never part of a comparison
  • Assumes an absent cell never changes how a column is held
  • Subtracts one output from the other without checking what lines them up