Why do 64-bit record identifiers get silently corrupted by a 64-bit-float numeric type?
answer
- the significand is not 64 bits wide
- eleven bits go to the exponent
- 53 bits of exact integer range
- above that, only every other integer exists
- 2^53 is roughly nine quadrillion
basics
~20 sA 64-bit float carries a 53-bit significand, so it represents integers exactly only up to 2^53, about 9.0e15. Larger identifiers round to the nearest representable value with no error raised, so distinct IDs shift by one or collide with each other.
solid answer
~50 sThe two types are both 64 bits wide but hold different sets of values: a 64-bit integer spends all its bits on magnitude, while a 64-bit float spends 11 on the exponent, leaving 53 bits of exact integer range. Every integer up to 2^53 (9,007,199,254,740,992) is representable; above that only even integers are, above 2^54 only multiples of four, and so on. A 19-digit identifier passed through a stage whose only numeric type is a 64-bit float is rounded to the nearest representable neighbour and handed back looking entirely plausible — no exception, no truncation warning. The symptoms downstream are sporadic lookup failures on high-numbered records and, eventually, two records answering to the same identifier. The remedy is to carry identifiers as opaque strings across any such boundary and as wide integers everywhere else: identifiers are names, not quantities, and nothing arithmetic is ever done to them.
go deeper
Remember that a 64-bit float and a 64-bit integer do not hold the same values: the float spends eleven bits on the exponent, leaving 53 bits of exact integer range at roughly nine quadrillion.
Explain that above 2^53 only every second integer is representable and above 2^54 only every fourth, so the conversion rounds rather than failing and the loss is completely silent.
Show the diagnosis: round-trip a known large identifier through one hop and compare integers, check parity of stored identifiers above the boundary, then find the hop whose numeric type is a 64-bit float.
Set the rule that identifiers are names rather than quantities and never pass through a float-typed field, and require a boundary test above 2^53 in every contract so the guarantee outlives the incident that found it.
## Two 64-bit types, two different value sets A signed 64-bit integer represents every whole number from about -9.22e18 to 9.22e18, spaced one apart, with nothing in between. A 64-bit binary float represents about the same range of magnitudes but spends 1 bit on sign and 11 on the exponent, leaving a 53-bit significand. Its values are `(53-bit integer) x 2^k`, which means they are spaced *one apart only while the exponent is small enough that the significand can carry the whole integer*. That boundary is `2^53 = 9,007,199,254,740,992`, roughly 9.0e15 — a 16-digit number. Below it, every integer is exactly representable and integer arithmetic that stays in range is exact. At and above it, the spacing doubles with each binary exponent: | range | integers that exist as 64-bit floats | |---|---| | below 2^53 | every integer | | 2^53 to 2^54 | every second integer | | 2^54 to 2^55 | every fourth integer | | 2^55 to 2^56 | every eighth integer | So the odd numbers simply have nowhere to land above 2^53. Converting one produces its nearest even neighbour. ## Why the corruption is silent Conversion from a wide integer to a 64-bit float is defined to round, not to fail. There is no overflow, no truncation flag anyone checks, no exception. The value that comes back is a perfectly valid float that prints as a plausible 19-digit number differing from the original by one or two. Every layer downstream treats it as the identifier it claims to be. This is why the failure looks like everything except what it is. A lookup by the round-tripped identifier misses, so the report is "record not found", which points at the store rather than the transport. Two neighbouring identifiers can round to the same value, so the report is "the wrong record came back", which points at a cache or a join. And the problem grows monotonically: identifiers issued later are larger, so a system that has been healthy for years starts failing as its sequence crosses the boundary, and then fails more often. ## Diagnosing it Prove the conversion rather than the lookup. Take a known identifier above 2^53, send it through one hop of the pipeline in isolation, and compare the decoded value against the original *as an integer*, not as a printed string that a formatter may have re-rounded. If the decoded value differs — usually by one — that hop is the culprit. A second signature: sample stored identifiers above the boundary and look at parity. If none of them are odd, they have already been rounded to even neighbours, and the corruption is in your data, not just in flight. Then walk the pipeline and find the hop whose numeric type is a 64-bit float. It is rarely the database and rarely the identifier generator; it is a boundary in between. ## The remedy Carry identifiers as opaque strings across any boundary whose numeric representation is a 64-bit float, and as wide integers everywhere inside. This costs a few bytes and buys exactness, and it loses nothing, because an identifier is a *name*: no one adds two of them, averages them or takes their square root. Ordering, when needed, is either lexicographic over fixed-width strings or done on the integer form on one side. Then make the guarantee enforceable rather than remembered: add a contract test that round-trips a specific value above 2^53 — a good choice is `2^53 + 1`, the smallest integer the format cannot represent — and asserts exact equality. That single test outlives everyone who was in the incident review. ## Ecosystems made different calls here This bug is a boundary bug, and boundaries are where ecosystems disagree. Text interchange formats commonly leave numeric precision to the implementation, and implementations chose differently: some runtimes decode every number to a 64-bit float regardless of whether it had a fractional part, while others decode integral literals to arbitrary-precision or 64-bit integers. Consequently the same payload can survive one hop intact and be mangled at the next, and reproducing the bug on a colleague's stack can fail for reasons that have nothing to do with your data. That divergence is exactly why the fix is to stop shipping the value as a number at all.
- Our identifiers are 16 digits and nothing has broken. Is this a real risk?Sixteen digits straddle the boundary: 2^53 is about 9.0e15, itself a 16-digit number. Some identifiers already issued are above it, so the corruption is present but rare enough to read as sporadic 'record not found'. It also worsens monotonically, because every newly issued identifier is larger than the last. Treat it as an active incident with a slow fuse rather than a future risk, and check the parity of stored identifiers above the boundary to see how far it has already spread.
- How do you prove the conversion is the culprit rather than a lookup bug?Isolate one hop. Take a known identifier above 2^53, encode it, decode it, and compare the result to the original as an integer — never as a printed string, which a formatter may re-round into agreement. A difference of one or two confirms the numeric conversion. As a second signal, sample stored identifiers above the boundary and check parity: if none are odd, they were rounded to even neighbours and the corruption is already in the data.
- What is the standard remedy, and why is it cheap?Carry identifiers as opaque strings across any boundary whose numeric type is a 64-bit float, and as wide integers everywhere inside. It is cheap because identifiers are names rather than quantities: nothing sums, averages or scales them, so the string form costs only bytes on the wire. Add a contract test that round-trips 2^53 + 1 — the smallest integer the format cannot hold — and asserts exact equality, so the guarantee is enforced rather than remembered.
A 64-bit float is a ruler whose marks spread further apart the further out you read. Past nine quadrillion the marks are two units apart, so an odd identifier has nowhere to land.
saying these in an interview costs you the question
- Assumes a 64-bit float holds every 64-bit integer
- Thinks the conversion would raise an error if it lost data
- Blames the store or the identifier generator
- Fixes it by shortening identifiers to fewer digits
- Believes the problem only starts at 2^63