skip to content

A rewritten aggregation agrees with the old one to twelve digits but not exactly. Why is exact equality the wrong verdict here?

level: middleimportance: should knowfreq 52%

answer

  1. reordered addition is not the same addition
  2. the rewrite regrouped the arithmetic
  3. exact equality is a tolerance of zero
  4. two terms: fixed units and a fraction
  5. signed differences, not just the maximum

basics

~20 s

Two correct implementations that combine the same fractional numbers in a different order do not produce bit-identical results. Exact equality is a tolerance of zero, chosen by omission; a diff has to state how close counts as the same.

solid answer

~40 s

The rewrite was free to add the same numbers in a different order or in different sized batches — that is often the point of the rewrite — and fractional values held in a fixed width do not give identical results under reordered addition. So the last digits diverge while the result is just as correct. A comparison that demands bit-identity is not neutral about this: it has silently chosen a tolerance of zero. State the tolerance instead, with both of its terms — how far apart two numbers may be in fixed units, and as a fraction of their magnitude — because a column spanning many magnitudes needs both. And state it *before* running the diff: a tolerance widened until the diff passes is not evidence of anything.

go deeper

for a junior

Recall that two correct versions of the same arithmetic can disagree in their final digits, so "not identical" is not the same as "wrong".

for a middle

Explain that exact equality is itself a tolerance of zero, and state the two terms a tolerance has and why a column spanning many magnitudes needs both of them.

for a senior

Show the discipline: the tolerance is decided from what the consumer can absorb and written before the diff runs, and a passing diff is still read for the direction of its differences.

for a principal

Own the standard: what difference is acceptable in a published number, who may widen it, and whether an accepted difference has to be recorded with a reason rather than absorbed silently.

## Two correct implementations, two different last digits When a transform is rewritten, the new version frequently combines the same numbers in a different order: one long running total becomes several partial totals added at the end, a per-row pass becomes one expression evaluated over a whole column in batches, a condition moves earlier so fewer values reach the sum. Fractional numbers held in a fixed width do not behave like the reals here — regrouping the same additions can change the final digits. *Why* that happens is the subject of fixed-width fractional arithmetic and is not what this question is about. **The consequence is what matters: two implementations can both be right and still not be bit-identical, so a comparison between them has to be given a tolerance at all.** Which computations a faster form is actually free to regroup is a separate question about the operation, not about the diff. For the diff, assume it happened unless you can show it did not. ## Exact equality is a tolerance of zero The important move is to see that there is no tolerance-free option: - A comparison that asks "are these identical?" **has** a tolerance. It is zero, and nobody chose it deliberately. - A zero tolerance makes the diff report on the arithmetic's last digits, which nothing downstream depends on, alongside the differences that matter — which nothing then distinguishes. - Once the diff reports thousands of last-digit differences, people stop reading its output. That is the real cost: the instrument gets ignored rather than fixed. So the choice is between a tolerance you stated and a tolerance you inherited. ## The two terms A stated tolerance has two parts, and a serious diff carries both: | term | what it says | where it is the binding one | |---|---|---| | absolute | how far apart two numbers may be **in fixed units** | values near zero, where a fraction of the magnitude is meaninglessly small | | relative | how far apart they may be **as a fraction of their magnitude** | large values, where a fixed floor would be absurdly strict | A column of monetary totals running from a few pence to several million needs both: the absolute term keeps the near-zero rows from failing on noise, and the relative term keeps the large rows from passing on a real error that happens to be smaller than the fixed floor. Choosing between them as a numerical matter is its own subject; for the diff, the rule is simply **state both, and state them before you run it**. ## A tolerance is a commitment, not a mute button The failure mode here is social rather than technical: 1. The diff fails. Someone widens the tolerance. It passes. Nothing was learned, and the number now in the file records how big the disagreement was, not how big it is allowed to be. 2. The correct order is the reverse: decide what difference the consumer of this output can absorb, write that down as the tolerance, then run the diff and treat a failure as a finding. And a tolerance that passes is not the end of the reading. Look at the **signed** differences, not just the largest magnitude: - Scattered, both-signed, last-digit noise is consistent with a reordered computation. - Tiny differences **all in the same direction**, on every row, are consistent with a changed rounding step, a dropped term, or a unit conversion applied once too often — and they can pass a tolerance while being a real defect that grows when the outputs are summed. ## Where a tolerance must not be applied - **Identifiers, keys and codes.** These must match exactly. A tolerance on a key column is meaningless, and a near-match on an identifier is a defect, not an acceptable difference. - **Text and category labels.** Equal or not; there is no near. - **Counts and other whole-number measures**, unless the rewrite deliberately changed how something is counted — in which case the difference is a change in the definition and belongs in the report, not under a tolerance. So a well-formed diff applies exact comparison to the key and the discrete columns, and a stated absolute-plus-relative tolerance only to the fractional measures — and says in its output which columns were compared under which rule.

  • The diff passes at the tolerance you set, but every difference has the same sign. Is that acceptable?
    Treat it as a finding. Noise from regrouped arithmetic is scattered in both directions; a one-directional bias on every row points at a changed rounding step, a dropped term or a conversion applied an extra time. It can pass row by row and still be large once the column is summed, so investigate before accepting it.
  • Why does a diff need both an absolute and a relative term rather than just one?
    Because a real column spans magnitudes. A purely relative term is unusably strict near zero, where tiny absolute noise is a large fraction of the value; a purely absolute term is unusably loose on large values, where a genuine error can hide under the fixed floor. Stating both makes each one binding where it is the sensible measure.
  • Should the same tolerance be applied to every column in the output?
    No. Keys, identifiers, codes and text must compare exactly — a near-match there is a defect, not an accepted difference. Apply the tolerance only to the fractional measures, and have the diff report which columns were compared under which rule so a reader can see what was actually proved.

saying these in an interview costs you the question

  • Treats any last-digit difference as proof the rewrite is wrong
  • Believes a comparison without a tolerance has not chosen one
  • Picks the tolerance after the diff fails, widening it until it passes
  • Uses only a fixed-unit tolerance on a column spanning many magnitudes
  • Applies a numeric tolerance to identifier or key columns
  • Reads only the largest difference and never the direction of the differences