The same threshold comparison over a column with absent values yields a two-state result on one tool and a three-state result on another - why?
answer
- the operator is ordinary; the representation is not
- a borrowed sentinel compares unequal to everything
- three-valued semantics carry a third state
- true count plus false count against the length
- not equal to itself, under one design
basics
~20 sBecause the two tools represent absence differently. Where absence is a borrowed floating-point sentinel, every comparison against it is simply false, so the condition column has two states. Where absence is tracked separately, the comparison yields absence and a third state travels onward.
solid answer
~50 sA comparison operator is element-wise like any other: it walks both sides position by position and hands back a condition column of the same length. What it puts at a position whose value is absent depends on how the tool stores absence. One family of designs borrows a floating-point sentinel to mean "no value"; that sentinel is defined to compare unequal to everything, itself included, so every comparison involving it is false and the condition column holds only true and false. Another family records absence separately - a **validity bit** beside a value of the ordinary width - and runs the comparison under three-valued semantics, so the result at that position is neither true nor false but absent, and the condition column carries three states. Whatever consumes that condition column then has to be told what the third state means.
go deeper
Know that a comparison over a column returns a new column of per-position verdicts, the same length as what you compared, and that positions with no value are the ones worth asking about.
Explain both outcomes and tie each to how absence is stored - a borrowed floating-point sentinel against a separately tracked marker - rather than presenting one of them as what comparisons do.
Show that you would detect this rather than assume it: the true-plus-false count against the column length, and equality against itself, both answer the question in a line and both are cheap.
The lasting decision is whether your pipeline standardises on one absence representation before data reaches anyone's expression, so that comparison semantics stop being a per-tool property people rediscover.
## What a comparison hands back A comparison written over a whole column is an element-wise operator with a different result type. It walks the operands position by position and allocates a new column the same length, holding a per-position verdict rather than a number. Everything true of arithmetic element-wise operators is true of it: it is one dispatch from your program, the loop runs inside compiled code, and the operands are untouched. The interesting question is what it puts at a position where there is no value to compare. ## Two representations of absence, two answers | | Absence as a borrowed floating-point sentinel | Absence tracked separately | |---|---|---| | How "no value" is stored | a reserved floating-point value in the same buffer | a validity bit, or an equivalent marker, beside a value of the ordinary width | | Result of comparing it to a threshold | false | absent | | Result of comparing it to itself | false - the sentinel is unequal to everything | absent | | States in the condition column | two | three | | What the consumer must handle | nothing extra | what the third state means | The first design's behaviour is not an arbitrary choice: the sentinel it borrows is *defined* by the floating-point specification to compare unequal to every value including itself, so "false at an absent position" falls out of the representation rather than being decided by the library. The second design has no such sentinel to lean on, so it has to decide, and the coherent decision is that a comparison whose input is unknown has an unknown answer. Be careful not to fuse the two axes. **How a tool stores absence** and **what semantics its comparisons run under** are two separate choices, and tools bundle them differently. The pairing above is the common one, not a law. ## Why this is a senior question rather than trivia It is a silent divergence. Both tools accept the expression, both return a column of the right length, neither warns. The difference only shows up in whatever consumes the condition column, and by then the comparison is several steps behind you. Two specific consequences are worth carrying: - **The equality-with-itself trap.** Under a sentinel design, a value at an absent position is not equal to itself. Any code that identifies rows by comparing a column against a copy of itself, or that treats equality as reflexive, breaks exactly at those positions - and reports "different" rather than raising. - **The count that does not add up.** Under a sentinel design, the counts of true and false positions sum to the length of the column. Under three-valued semantics they do not, and the shortfall is exactly the absent positions. That arithmetic is the cheapest way to tell which design you are on without reading any documentation. ## What to check when you move between tools 1. Compare a small column containing an absent value against a threshold, and look at the states in the result. Two states or three answers the question in one line. 2. Compare that column against itself for equality. False at the absent position means a sentinel design; absent means three-valued semantics. 3. Add up the true count and the false count and compare that total against the column's length. A shortfall equal to the absent count is three-valued semantics confirming itself. ## The honest form of the claim Do not say "comparing against absence gives false" and do not say "it gives absence." Say which design you mean: "against a borrowed floating-point sentinel a comparison is simply false, so even equality with itself fails; under three-valued semantics the result is neither true nor false and the condition column carries absence onward." That is the answer that survives the interviewer switching tools on you, and it is also the answer that survives your team switching tools on you. ## What a strong candidate adds unprompted That the operator itself is unremarkable - it is the representation of absence that produces the divergence, not anything special about comparison - and that the result is a new column either way, allocated at the full length of the operands, with the originals unchanged. Locating the surprise in the representation rather than in the operator is the sign that someone understands the mechanism rather than having memorised one tool's behaviour.
- Without reading any documentation, how would you tell which of the two designs a tool uses?Compare a short column holding one absent value against a threshold and look at the states in the result. Two states means a sentinel design; three means three-valued semantics. As a cross-check, add the true and false counts: if they fall short of the column's length by exactly the number of absent values, the third state is real.
- Why is "not equal to itself" a sensible behaviour rather than a bug?It is not a library decision at all. The design borrows a reserved floating-point value to stand for absence, and that value is specified to compare unequal to every value including itself. The library inherits the behaviour along with the representation. It is coherent, but it does surprise anyone expecting equality to be reflexive.
- Does the comparison change the operand column in any way?No. Like any plain element-wise operator it allocates a new column of the same length and writes only there, leaving the operand exactly as it was, absent positions included. The result's representation is a per-position verdict rather than a number, which is usually much narrower per value.
saying these in an interview costs you the question
- Says a comparison against an absent value is false, with no condition attached
- Says a comparison against an absent value yields absence, with no condition attached
- Assumes absence is converted to zero before a numeric comparison runs
- Expects the true and false counts to always sum to the column length
- Treats equality as reflexive in every design, including sentinel-based ones