How do NaN and Infinity readings break a max-finding scan over a sensor stream?
answer
- what does a comparison against NaN return
- all of them, equality with itself included
- so the branch simply never fires
- where the bad reading lands changes the symptom
- reject at ingest, not in every consumer
basics
~20 sEvery comparison involving NaN is false, so a NaN seed pins the running maximum at NaN forever, while a NaN arriving later is silently skipped. Infinity compares as an ordinary value and wins every comparison. Neither is reported as an error.
solid answer
~50 sNaN is *unordered* with everything: `<`, `<=`, `>`, `>=` and `==` are all false when either operand is NaN, and `!=` is the only true comparison — including against itself. So the corrupt reading gives two opposite symptoms depending on where it lands. If it seeds `best`, then `readings[i] > best` is false for every later element and the loop returns NaN having ignored the whole stream. If it arrives mid-scan it is skipped, and the result is a plausible maximum over data you did not know was incomplete — the dangerous case, because nothing looks wrong. Infinity is different: a legitimate value from overflow or division by zero, it wins every comparison and pins the maximum. The fix belongs at ingest — validate once, count rejects as a metric, decide deliberately between dropping and propagating — not as an improvised guard in each consumer.
code
pseudocode · 7 lines// running maximum over a stream of 64-bit float readings
best = readings[0]
for i in 1..n-1
if readings[i] > best
best = readings[i]
...
return bestgo deeper
Know that NaN is not equal to anything, itself included, and that Infinity is an ordinary comparable value rather than an error code. Neither one halts a computation on its own.
Trace the loop both ways: a NaN seed leaves every comparison false so the maximum never updates, while a later NaN is skipped and the result looks entirely plausible. Name where NaN and Infinity originate.
Show that you would validate readings at the ingest boundary, expose the reject count as a metric, and choose deliberately between dropping and propagating rather than letting each consumer improvise a guard.
Own the data contract: what a corrupt reading means to every downstream aggregate, whether the system reports an honest smaller sample or a loud non-value, and how that choice is made visible to whoever reads the dashboard.
## The two special values and where they come from **Infinity** is a real, ordered value in the format, with both signs. It arises from overflow past the largest representable magnitude and from dividing a non-zero value by zero. It compares normally: `+Infinity` is greater than every finite value, `-Infinity` less than every finite value. It is not an error indicator; the arithmetic produced it deliberately. **NaN** — "not a number" — is the result of an operation with no meaningful answer: zero divided by zero, Infinity minus Infinity, Infinity times zero, the square root of a negative value, and the parse of an unparseable numeric field. It also has two flavours: a *quiet* NaN, which propagates through arithmetic silently, and a *signalling* NaN intended to raise a condition — but in practice almost everything produced and moved around is quiet. Both propagate. Almost any arithmetic with NaN as an operand yields NaN, so a single corrupt reading poisons every aggregate downstream of it with no error raised anywhere along the way. ## Comparison: NaN is unordered This is the rule that produces the surprising behaviour. Any two values are in exactly one of four relations: less than, equal, greater than, or **unordered**. NaN is unordered with everything, itself included. Concretely: - `NaN < x`, `NaN <= x`, `NaN > x`, `NaN >= x`, `NaN == x` are all **false** for every `x`, including when `x` is NaN. - `NaN != x` is **true** for every `x`, including when `x` is NaN. That last line is why `x != x` is the classic NaN test: it is true exactly when `x` is NaN and false for every other value, both infinities included. It is also why a NaN cannot be found in a collection by an equality search, and why any ordering over data that may contain NaN is not a total order until you define where NaN sits. ## Tracing the scan ``` best = readings[0] for i in 1..n-1 if readings[i] > best best = readings[i] ``` **Case 1 — the corrupt reading is first.** `best` is NaN. Every `readings[i] > best` is false because the comparison is unordered, so the assignment never runs. The loop completes normally and returns NaN, having examined and ignored the whole stream. This is the *loud* failure: the output is obviously not a measurement. **Case 2 — the corrupt reading is in the middle.** `NaN > best` is false, so it is skipped. The loop returns the maximum of the remaining readings. Nothing signals anything; the dashboard shows a plausible number, and the fact that the sample was incomplete is invisible. This is the *quiet* failure, and it is the one that matters, because it produces a wrong-but-believable answer. Note that the symptom flips if you write the scan as a minimum with `<`, and flips again if you seed with negative Infinity instead of the first element. Seeding `best = -Infinity` removes case 1 but keeps case 2. The behaviour of your aggregate under corrupt input is therefore an accident of how the loop was written unless someone decided it on purpose. **Infinity in the same loop** behaves "correctly" in the sense that the comparison is well defined: `+Infinity > best` is true, so it becomes the maximum and stays there. The output is not a lie about the comparison; it is a faithful report of a value that came from an overflow you did not intend. ## Deciding rather than improvising The right question is not "which guard do I add here" but "what does a corrupt reading mean". Two defensible answers: - **Reject at ingest.** A reading that is NaN or infinite is not a measurement, so it is counted, dropped, and surfaced as a metric. The maximum is then honest about being computed over a smaller sample, and the reject count is the alarm. - **Propagate deliberately.** The aggregate becomes NaN loudly, and no one mistakes it for a measurement. Appropriate where an incomplete sample is worse than no answer. The failure mode to avoid is the third, accidental option: silently skipping the corrupt reading inside the loop, which reports a confident maximum over data you did not know was incomplete. Whichever you choose, make it a property of the ingest boundary so every consumer inherits it, and expose the reject count so silent data loss becomes visible. ## Even the standard found this hard How a maximum operation should treat NaN has no obvious right answer, and the standard itself changed position: the 2008 edition specified maximum operations that returned the non-NaN operand, and the 2019 edition deprecated those in favour of a family where one variant propagates NaN and another skips it — precisely because "what is the maximum of NaN and 3" depends on whether the NaN means "missing" or "broken". If the specification needed two answers, your pipeline needs to choose one explicitly too.
- Why does the test x != x identify NaN?NaN is defined to be unordered with every value, itself included, so equality is false and inequality is true when either operand is NaN. That makes `x != x` true exactly for NaN and false for everything else, both infinities included — the one comparison that identifies it without a dedicated predicate. The flip side is that a NaN stored in a collection can never be found by an equality search, and an ordering over data containing NaN is not a total order until you define where it sits.
- Where do Infinity and NaN come from in a numeric pipeline?Infinity comes from overflow beyond the largest representable magnitude and from dividing a non-zero value by zero. NaN comes from genuinely undefined operations: zero divided by zero, Infinity minus Infinity, Infinity times zero, the square root of a negative value, and unparseable numeric input. Once either appears it propagates, because almost any arithmetic with a NaN operand yields NaN — so one corrupt reading can poison an entire downstream aggregate with nothing raised anywhere.
- Should the scan reject a NaN reading or propagate it?Decide once, at ingest, and make it explicit. Rejecting means the corrupt reading is counted and dropped, so the maximum is honest about a smaller sample and the reject count is the alarm. Propagating means the aggregate loudly becomes NaN and nobody mistakes it for a measurement. The mode to avoid is the accidental third one — silently skipping it inside the loop — which reports a confident maximum over data you never knew was incomplete.
saying these in an interview costs you the question
- Says NaN equals NaN because the bit patterns match
- Assumes a NaN reading raises an error somewhere
- Treats Infinity as an error code rather than a value
- Adds a NaN guard in every consumer instead of at ingest
- Believes sorting reliably pushes NaN to one end