skip to content

Why does float('nan') compare unequal to itself, and what breaks as a result?

level: middleimportance: should knowfreq 45%

answer

  1. Undefined results get their own value
  2. The standard calls it unordered
  3. No trichotomy, so ordering breaks
  4. Containment tries identity first
  5. math.isnan is the only honest test

basics

~20 s

IEEE-754 defines nan as unordered, so ==, < and > are all False against it and != is True. That silently breaks sorting, min and max, and makes containment depend on whether the same nan object is reused.

solid answer

~40 s

A nan is the IEEE-754 result of an undefined operation, and the standard makes it *unordered*: `n == n`, `n < 1.0` and `n > 1.0` are all False, and only `!=` is True. Three consequences bite in real code. Sorting silently returns a wrong order rather than raising, because `sorted` relies on `<` being a consistent ordering and nan makes it inconsistent. `max` and `min` become order-dependent — `max(1.0, float('nan'))` is 1.0 but swapping the arguments gives nan. And containment is confusing: CPython's list containment check tries identity before equality, so `n in [n]` is True for the *same* object while a freshly built `float('nan')` is not found. Detect nan with `math.isnan` or `math.isfinite`; never with `== float('nan')`.

code

python · 10 lines
python
import math

n = float("nan")
print(n == n, n != n)                        # False True
print(n < 1.0, n > 1.0)                      # False False
print(n in [1.0, n])                         # True: `is` is tried first
print(float("nan") in [1.0, float("nan")])   # False: different objects
print(sorted([3.0, n, 1.0, 2.0]))            # [3.0, nan, 1.0, 2.0]
print(max(1.0, n), max(n, 1.0))              # 1.0 nan
print(math.isnan(n), math.isfinite(n))       # True False

go deeper

for a junior

Recall the single fact and its consequence: a nan is never equal to anything, itself included, so you test for it with math.isnan rather than ==. Know that float('inf') exists and is ordered normally.

for a middle

Explain unorderedness rather than just inequality — all of ==, < and > are False — and name the practical fallout: silently unsorted output, order-dependent max and min, and containment that hinges on whether the same nan object is reused.

for a senior

Demonstrate that you defend at the boundary: validate numbers where they enter the process, filter or reject nans before ordering or aggregating, and recognise the symptom of a silently unsorted result or an aggregate that flips with row order. Explain how an overflow upstream becomes a nan downstream.

for a principal

Decide the policy for missing and undefined numeric values in a system: whether they are represented as nan, as None, or rejected at the edge; who is responsible for the check; and how contracts between services keep an unordered value from reaching code that assumes a total order.

### Where a nan comes from A nan is the IEEE-754 value for "this operation has no meaningful result": `float('inf') - float('inf')`, `float('inf') * 0.0`, `0.0 * float('inf')`, and parsing `float('nan')` directly. Note that Python does *not* follow IEEE-754 for division by zero: `1 / 0` and `1.0 / 0.0` both raise `ZeroDivisionError` rather than producing an infinity, which is a deliberate Python departure. Nans also arrive from outside — a missing value in a data file, a numeric column with holes, a computation done in a lower-level library — which is why the behaviour matters even in code that never creates one on purpose. ### Unordered, not just unequal The standard classifies every comparison with a nan as **unordered**, and the required outcome is that `==`, `<`, `<=`, `>` and `>=` all return False while `!=` returns True. That includes comparing a nan with itself, which is the property people remember. It is worth stating the stronger form in an interview: the problem is not merely that equality fails, it is that nan breaks the *trichotomy* every ordering algorithm assumes — for any two values exactly one of less-than, equal, greater-than should hold, and with a nan none of them does. ### What that breaks **Sorting silently lies.** `sorted([3.0, float('nan'), 1.0, 2.0])` returns `[3.0, nan, 1.0, 2.0]` — not sorted, and no exception. Timsort makes a bounded number of comparisons and every comparison against the nan answers False, so elements simply stay where the nan left them. The output shape depends on where the nan sat in the input, which makes the bug reproduce inconsistently. **`max` and `min` become order-dependent.** They keep a running candidate and replace it only when a comparison says to. `max(1.0, float('nan'))` returns 1.0, while `max(float('nan'), 1.0)` returns nan. A summary statistic computed over data with one missing value can be a nan or not depending on row order. **Containment and deduplication depend on object identity.** CPython's containment checks and its hash-based lookups both try `is` before `==`, as a performance shortcut that also happens to make containers behave reflexively. So with `n = float('nan')`, `n in [1.0, n]` is True and `{n, n}` has one element, but `float('nan') in [1.0, float('nan')]` is False and `{float('nan'), float('nan')}` has two elements, because those are distinct objects that never compare equal. A set built from data containing several nans can therefore hold several of them; deduplication does not deduplicate. **Propagation.** Any arithmetic touching a nan yields a nan, so one bad value silently poisons a total, an average or an aggregate all the way to the output. That is a feature — the alternative is a plausible-looking wrong number — but only if something eventually checks. ### Infinities, by contrast, are well behaved `float('inf')` is ordered: it compares greater than every finite float, `-float('inf')` less than every one, and sorting works. `math.inf` is the same value spelled as a constant. Overflow reaches it by two different routes worth distinguishing: multiplying two very large floats such as `1e200 * 1e200` yields `inf`, and `float('1e400')` parses to `inf`, but `10.0 ** 400` raises `OverflowError`, as do the `math` functions when a result exceeds range. Subtracting equal infinities, or multiplying one by zero, produces a nan — which is how an overflow in the middle of a pipeline turns into the unordered case downstream. ### Detecting and defending Use `math.isnan(x)` to test for a nan and `math.isfinite(x)` to reject nan and both infinities in one call; `math.isinf(x)` isolates the infinities. Since a nan compares unequal to itself, the idiom `x != x` also detects one, and you will meet it in older code, but `math.isnan` says what it means. The defensive habits are to validate at the boundary where external numbers enter, to filter nans out before sorting or computing extremes rather than after, and never to write `x == float('nan')`, which is always False and is therefore a check that can never fire.

  • How do you actually test whether a float is a nan?
    `math.isnan(x)`, or `math.isfinite(x)` when you want to reject infinities too. The old idiom `x != x` works precisely because nan is unequal to itself, but it reads as a typo. What never works is `x == float('nan')`: that comparison is False for every input, including a nan, so the branch is dead code.
  • Why can a set built from float data end up holding several nan values?
    Set membership compares by hash then by identity then by equality. All nans hash alike, but two distinct nan objects are neither identical nor equal, so each one is stored separately. Reusing a single nan object collapses them to one, which is why the behaviour looks inconsistent between a literal-built set and one built from parsed data.
  • Does an infinity cause the same ordering problems as a nan?
    No. `float('inf')` and its negative are fully ordered against every finite float, so sorting, `max` and `min` behave. The danger with infinities is arithmetic: subtracting two equal infinities or multiplying one by zero produces a nan, so an overflow early in a pipeline becomes an unordered value later.

saying these in an interview costs you the question

  • Tests for a nan with x == float('nan')
  • Expects sorting a list with nan to raise
  • Assumes nan sorts to the start or the end
  • Thinks 1.0 / 0.0 returns inf in Python
  • Believes a set always deduplicates nan values
  • Confuses nan with None or with a missing key

context