skip to content

Why does converting a 19-digit int to a Python float make distinct IDs collide?

level: seniorimportance: nice to knowfreq 25%

answer

  1. int is unbounded, float is not
  2. Only 53 significand bits survive
  3. Nine quadrillion is the boundary
  4. Neighbouring integers share one double
  5. An identifier is not a quantity

basics

~10 s

A float keeps only 53 significand bits, so integers above 2**53 round to the nearest representable value. Two 19-digit identifiers can round to the same double, after which they are indistinguishable and compare equal.

solid answer

~40 s

Python's `int` is arbitrary precision, but `float` is binary64 with a 53-bit significand, so every integer above 2**53 (about 9.0e15) is rounded on conversion. `float(2**53 + 1) == float(2**53)` is True: two adjacent integers become one double. A 19-digit identifier has roughly 63 bits, so conversion throws away the low ten bits and neighbouring IDs collide silently — no exception, no warning, just two distinct records that now key the same. The usual cause is an identifier passing through a numeric type it was never meant to touch. The fix is to keep identifiers as `int` or `str` end to end and never let them be parsed as floats. Note that Python's own `int`-to-`float` *comparisons* are exact, so `10**23 == 1e23` is correctly False.

code

python · 7 lines
python
item_id = 9007199254740993          # 2**53 + 1
as_float = float(item_id)

print(as_float)                     # 9007199254740992.0
print(int(as_float) == item_id)     # False: the round trip lost a unit
print(as_float == float(item_id - 1))  # True: two distinct IDs collide
print(item_id.bit_length())         # 54: one bit too many for a double

go deeper

for a junior

Remember the asymmetry: Python integers grow without limit, but floats do not, so a very large integer loses digits when it becomes a float. Avoid putting identifiers through float conversion at all.

for a middle

Explain the mechanism: 53 significand bits, exactness up to 2**53, and rounding to the nearest representable value above it. Know that int(float(n)) == n is the round-trip test and that overflow beyond the float range raises rather than rounds.

for a senior

Diagnose the silent, data-dependent version of this: records merging because two identifiers collided after crossing a boundary that encoded numbers as doubles. Show that you fix it by keeping identifiers as int or string end to end rather than by adjusting tolerances.

for a principal

Own the contract at the edges of the system: which fields are quantities and which are identities, what each serialization and storage layer guarantees about integer width, and how that guarantee is asserted so a precision loss fails at the boundary instead of surfacing as merged records weeks later.

### The boundary A Python `int` has no width limit; it grows as needed. A Python `float` is binary64 with a 53-bit significand, so the integers it represents exactly are those with at most 53 significant bits: everything up to 2**53 = 9007199254740992, roughly 9.0e15, or sixteen decimal digits. Above that, consecutive integers are no longer all representable — from 2**53 to 2**54 only even integers are, then only multiples of four, and so on — and `float(n)` returns the nearest representable value. So `float(2**53 + 1)` is 9007199254740992.0, exactly equal to `float(2**53)`. Two distinct integers, one double. Nothing raises; conversion of an in-range integer is defined to round. ### Why identifiers are the classic victim A 19-digit identifier needs about 63 bits, so conversion to a double discards the low ten or so bits. Consider a warehouse pick-list builder that assembles a run for an 11-person picking team: it reads order lines, groups them by item identifier, and emits one pick per group. If those 19-digit identifiers ever pass through a float — because a numeric column was read as a floating-point type, because a serialization boundary encoded every number as a double, or because an encoding mismatch between two services turned an exact integer field into a numeric one — then identifiers that differ only in their last few digits map to the same double. The grouping merges lines for two different items, and the picker is sent to one bin for stock that lives in two. The failure is silent and data-dependent: it appears only when two IDs happen to be close enough to collide, which is why it survives testing and shows up in production. The repair is structural rather than numeric. Identifiers are not quantities: they are never added, averaged or compared for magnitude, so they should be `int` or `str` from end to end, and any boundary that cannot carry a 64-bit integer should carry the identifier as a string instead. Widening the float or rounding harder cannot help — the information is gone at the moment of conversion. ### Losing precision without leaving Python The same rounding happens in ordinary arithmetic. `float(10**23)` is 1e23, whose exact value is 99999999999999991611392, so `int(float(10**23))` differs from `10**23` by more than eight million. Mixing an `int` and a `float` in an arithmetic expression converts the int to a float first, so a large exact integer entering a float expression is quietly rounded even though nobody called `float()`. True division with `/` always produces a float, which is the most common accidental route out of exactness; `//` keeps two ints as an int. **Comparisons are the exception, and it is worth knowing.** CPython compares an `int` with a `float` *exactly*, not by converting the int to a float first. That is why `10**23 == 1e23` is False and why ordering between huge ints and floats is always right. So an equality check will tell you the truth even where an arithmetic conversion would not. ### The other end: overflow Beyond about 1.8e308 there is no representable double at all, and `float(10**400)` raises `OverflowError: int too large to convert to float` — a loud failure, unlike the silent rounding below it. Note that not every overflow behaves the same way: `1e200 * 1e200` yields `inf` and `float('1e400')` parses to `inf`, while `10.0 ** 400` raises `OverflowError`. Integer-to-float conversion is in the raising camp. ### Detecting it The check that matters is a round-trip: `int(float(n)) == n` tells you whether the conversion was lossless, and `n.bit_length() <= 53` tells you in advance. `float.is_integer` confirms a float holds an integral value but says nothing about *which* integer it originally came from, because that information is already lost. When exact large values are genuinely needed in a decimal context, `decimal.Decimal` accepts an int without loss; when a ratio must stay exact, `fractions.Fraction` does. But for identifiers the right answer is not a better numeric type — it is not treating an identifier as a number at all.

  • How do you check whether a particular integer survives conversion to a float?
    Round-trip it: `int(float(n)) == n` is True exactly when the conversion was lossless. To test in advance, `n.bit_length() <= 53` answers the same question without doing the conversion. Both are cheap enough to assert at a boundary where external numbers arrive.
  • Since `float(10**23)` loses precision, why is `10**23 == 1e23` False rather than True?
    CPython compares an `int` with a `float` exactly rather than converting the int first, so the comparison sees that 1e23 really holds 99999999999999991611392 and reports inequality. Ordering works the same way. Arithmetic is different: mixing the two in an expression does convert the int, and rounds it.
  • What happens above the float range, and how is that different from what happens below it?
    `float(10**400)` raises `OverflowError` because no double is anywhere near the value — a loud failure. Below 1.8e308 the conversion succeeds and silently rounds. The dangerous region is the quiet one between 2**53 and the overflow limit, where every conversion looks like it worked.

saying these in an interview costs you the question

  • Assumes any int converts to float losslessly
  • Expects an exception instead of silent rounding
  • Thinks 64-bit floats hold 64-bit integers exactly
  • Treats identifiers as numeric quantities
  • Believes rounding after the fact can recover the value
  • Confuses the 53-bit limit with the overflow limit

context