skip to content

Why can two absent keys fall into one group while an equality test between those same two cells reports no match?

level: seniorimportance: should knowfreq 50%

answer

  1. two layers, two different rules
  2. a comparison may decline; a group may not
  3. all absences as one key
  4. check per surface, never infer
  5. the row count is the evidence

basics

~20 s

Two rules live in the same tool. Scalar comparison is defined by the encoding and cannot call two unknowns equal; the collective surfaces — splitting rows by key, removing repeats, placing rows in an order, matching two tables on a key — must decide for every row, so they treat all absences as one value.

solid answer

~50 s

Because the two layers answer different questions and carry different obligations. A scalar comparison is a statement about values, and no design in this family is willing to say that two unrecorded things are the same thing: under a borrowed floating-point pattern the reserved bits compare unordered with everything including themselves, and under a typed marker the comparison declines to commit at all. A collective surface cannot decline. Splitting rows by key, deciding which rows are repeats of each other, placing rows in an order and lining two tables up on a shared key all have to put every row somewhere, so each defines its own identity rule for absence — and the usual choice is that all absences are one and the same key. The practical consequence is blunt: you cannot infer the collective behaviour from the scalar rule, and you check it per surface.

go deeper

for a junior

Know that a cell holding no value can still act as a key: rows with nothing in the key column are often gathered together into one group rather than discarded.

for a middle

Explain why the layers differ — a comparison is allowed to return no usable answer, while a surface that must place every row has to pick an identity rule, and normally makes all absences one.

for a senior

Show the check you run: count the absent keys on each side before any grouping or key match, then confirm the surface's rule on a tiny fixture instead of inferring it from comparison behaviour.

for a principal

The rule worth arguing for is that key columns may not hold absence at all, enforced where data enters, because no downstream surface's default is right for every team and the two defaults fail in opposite directions.

## Two layers, two obligations The surprise here is that one tool holds two different rules about the same thing at the same time, and both are defensible. The **scalar layer** is comparison. Its job is to make a claim about two values, and it is allowed to refuse. Asked whether one unrecorded quantity equals another unrecorded quantity, it answers either "no" — because the reserved bit pattern the floating-point format uses compares unordered with everything, itself included — or "neither true nor false", because a typed absence marker under a third-outcome logic will not assert what it does not know. The **collective layer** is every surface that has to place each row somewhere. It is not allowed to refuse, because the output must account for the input. Four of these surfaces meet absence routinely: - splitting rows by a key and computing per group; - deciding which rows are repeats of each other; - putting rows in an order; - matching two tables on a shared key. Each of them therefore adopts an **identity rule** for absence, and the common choice — because it is the only one that produces a stable, reproducible result — is that all absences are one and the same key value. ## Why the collective surface cannot borrow the scalar rule Suppose a grouping surface used the scalar comparison. Three things break immediately: 1. **No group is well defined.** If absence never equals absence, every absent-key row is its own group of one, and a dataset with 300 absent keys produces 300 singleton groups that no analyst asked for and no report can use. 2. **The result is not reproducible.** Under a third-outcome logic there is no answer at all, so membership would depend on the order rows were examined in. 3. **Ordering becomes undefined.** A value that is neither less than, nor greater than, nor equal to anything gives a comparison-based ordering no information whatsoever, so where the absent rows end up would be a property of the algorithm's comparison sequence rather than something you can reason about. That is exactly why ordering surfaces expose the placement of absent rows as a setting instead of leaving it to the comparison. ## What the two rules look like at each surface | Surface | Under "all absences are one value" | Under "absence never matches absence" | |---|---|---| | Splitting rows by key | Absent-key rows form a single group and appear in the result | Each is its own group, or they are excluded | | Removing repeated rows | Two rows differing only in that both cells are absent count as repeats | They count as distinct and both survive | | Placing rows in an order | Absent rows are collected and placed together, wherever the setting says | No principled position exists | | Matching two tables on a key | Absent-key rows on one side pair with absent-key rows on the other | Those rows pair with nothing | ## The row count is the evidence The matching row is the one that costs money. Where absence is treated as an ordinary key value at the matching surface, every absent-key row on one side pairs with every absent-key row on the other. Three hundred on each side is not 300 result rows and not 600 — it is **ninety thousand**, produced entirely by cells that hold nothing. Under the opposite rule those rows pair with nothing at all and simply fail to appear, which is quieter and just as wrong if you expected them. So the same input, run through two tools that both claim to match on a key, can produce a result that is 90,000 rows too large or several hundred rows too small. Neither tool is broken; they made different, documented choices about one thing the comparison operator refuses to decide. ## How to check it in five minutes - **Count the absent keys on every side before you run anything.** One number per key column per table. If both are zero, none of this can bite you. - **Run the surface on a tiny fixture** with two absent keys on each side and read the output row count. That single number tells you which rule is in force, and it is faster than reading documentation you may misread. - **Assert the expected row count** after any key match, so a change of tool or version that flips the rule fails loudly instead of producing a plausible number. - **Do not generalise from one surface to another.** The same tool can treat absences as one key when splitting rows and as unmatchable when matching tables. The rule is per surface. - **Prefer removing absence from key columns at the boundary.** A key that may be empty is not really a key, and enforcing that upstream makes every downstream surface's rule irrelevant.

  • A table with 300 absent-key rows is matched on that key to a table with 300 of its own. What can happen?
    Under a rule that treats absence as an ordinary key value, every absent-key row on one side pairs with every absent-key row on the other, so 300 by 300 becomes 90,000 result rows produced by cells that hold nothing. Under a rule where absence never matches absence, those rows pair with nothing and quietly disappear from the result. Count the absent keys on both sides before running it.
  • Does removing repeated rows use the comparison rule or the collective rule?
    The collective one. Deciding whether two rows are repeats needs an answer for every pair of cells, including two that both hold no value, so the surface defines them as identical rather than leaving the decision open. Two rows differing only in that one column is absent in both are therefore usually treated as the same row, even though comparing those two cells directly would say otherwise.
  • Why do ordering surfaces expose the placement of absent rows as an explicit setting?
    Because a comparison-based ordering has nothing to work with. A value that is neither less than, nor greater than, nor equal to anything supplies no ordering information, so its final position would be an accident of the algorithm's comparison sequence. Making the placement a setting turns an accident into a decision you can state and test.

A sorting office has to put every parcel on some shelf, so all the unaddressed ones end up together on one shelf. That shelf is the grouping rule. Now ask the clerk whether two unaddressed parcels are for the same household: he will say he has no way to know. That is the comparison. Both answers are correct at the same time, and neither one predicts the other.

saying these in an interview costs you the question

  • Generalises the scalar comparison rule to grouping and key matching
  • Assumes absent keys silently vanish at every surface
  • Expects the same rule in every tool and at every surface
  • Says rows with an absent key cannot form a group at all
  • Treats absent-key rows as harmless because the cells hold nothing