Why do 1, 1.0 and True collapse into a single dict key in Python?
answer
- A dict maps equality classes, not types
- Equal objects must hash equal
- bool is a subclass of int
- hash(1), hash(1.0), hash(True) are all 1
- Assignment replaces the value, keeps the key
basics
~20 sThey are equal and therefore hash the same: hash(1), hash(1.0) and hash(True) are all 1, and bool subclasses int. A dict keeps one entry for the whole equality class, retaining the first key object inserted and overwriting only its value.
solid answer
~40 sA dict is a mapping over equality classes, not over types. The invariant `x == y` implies `hash(x) == hash(y)` is what makes lookups work, and Python's numeric tower makes `1 == 1.0 == True == decimal.Decimal("1")`, so all of them must hash to `1` and share one slot. A lookup computes the hash, then checks `is` and `==` on the candidate. The detail people miss: assigning to an existing key replaces the **value** and keeps the **key object already stored**, so `{1.0: 'a'}` followed by `d[1] = 'b'` leaves `{1.0: 'b'}` — a float key carrying a value written through an int. Sets behave the same way: `{1, 1.0, True}` is `{1}`. The defence is normalizing key types at the parse boundary.
code
pycon · 10 lines>>> hash(1), hash(1.0), hash(True)
(1, 1, 1)
>>> samples = {1.0: "from the float parser"}
>>> samples[1] = "from the int parser"
>>> samples
{1.0: 'from the int parser'}
>>> samples[True]
'from the int parser'
>>> {1, 1.0, True}
{1}go deeper
Recall the surprise itself: 1, 1.0 and True are one dict key, and True equals 1 because bool subclasses int. Being able to predict the length of {1, 1.0, True} is enough at this level.
Explain the mechanism: equal objects must hash equal, so the numeric tower forces one slot, and assignment to an existing key keeps the stored key object while replacing the value. Show it in a REPL without hesitating.
Demonstrate the production angle: mixed int and float parses silently merging counters, key type flipping the serialized output between runs, and normalizing key types at the parse boundary instead of downstream.
Own the boundary contract — one canonical key type per domain concept, chosen where data enters the system — and be able to argue why encoding type in the key beats scattering defensive conversions across call sites.
### One equality class, one key Python's dict does not group keys by type. It groups them by **equality**, and the hash exists only to find candidates quickly. The invariant that makes that work is: if `x == y`, then `hash(x) == hash(y)`. Python's numeric tower deliberately makes numbers of different types compare equal when their mathematical values match, so the hash invariant forces them to share a slot: ```pycon >>> 1 == 1.0 == True True >>> hash(1), hash(1.0), hash(True) (1, 1, 1) ``` `bool` is a subclass of `int`, with `True` equal to `1` and `False` equal to `0`. `decimal.Decimal("1")` and `fractions.Fraction(1, 1)` join the same club. So a dict or set can hold **one** entry for that whole family, and `d[True]`, `d[1]` and `d[1.0]` are three spellings of one lookup. The float side is not a coincidence of small values. CPython hashes a rational number by mapping it into modular arithmetic with `sys.hash_info.modulus` (2**61 - 1 on a 64-bit build), which is why `hash(0.5) == hash(Fraction(1, 2))` as well. Any two numeric objects that are `==` hash the same, by construction. ### The lookup path, and which key object survives A lookup computes the hash, picks a slot, and for each candidate does an identity check (`is`) first and `==` second. The identity shortcut is a fast path — and the reason a `float('nan')` key can be retrieved with the *same* object even though `float('nan') != float('nan')`. The detail people get wrong: on assignment to an **existing** key, a dict replaces the *value* and keeps the *key object it already stored*. So insertion order decides the key's type: ```pycon >>> samples = {1.0: "a"} >>> samples[1] = "b" >>> samples {1.0: 'b'} ``` The dict now maps a `float` key to the value written through an `int`. `dict.update()`, `|=` and `setdefault()` behave identically. The same holds for sets: `{1, 1.0, True}` is `{1}`, retaining the first element inserted. ### How this bites in real code Take a metrics scraper that parses sample values out of text and keeps a per-value tally. The text arrives in a locale-dependent number format, so one collector's parser yields the `int` `1` and another's yields the `float` `1.0` for the very same reading. Two things happen at once and neither raises: 1. The two tallies **merge** into a single entry, so a count you expected to be split is doubled. 2. Which one merged into which depends on arrival order, so the key serializes as `1` on one run and `1.0` on the next — a dashboard label or a JSON payload flips shape with no code change. The same collision shows up with flag-shaped keys: a mapping written as `{0: "off", False: "disabled"}` has one entry, and `{1: "on", True: "enabled"}` likewise. The fix is not to fight the dict; it is to **normalize at the parse boundary**. Decide one canonical key type — a `str`, a `Decimal`, or a plain `int` — and convert as data enters. When the *type* is genuinely part of the identity, make it part of the key: `(type(value).__name__, value)` or an explicit enum tag. And use enum members or strings rather than `True`/`False` for dict keys whenever `0` or `1` might also appear. ### NaN, and one curiosity `float('nan')` is hashable and can be stored, but since it is not equal to itself, only the exact object you inserted will find the entry — a freshly built NaN never matches. Since **3.10**, distinct NaN objects hash by object identity rather than all hashing to `0`; before that, a dict holding many NaN keys degraded toward a linear scan because they all landed in one slot. And a favourite piece of trivia: `hash(-1)` is `-2`. `-1` is the error sentinel in CPython's C-level hash API, so the integer `-1` reports `-2` instead — which means `-1` and `-2` share a hash. They are still two separate keys, because they are not equal: hash collisions are normal and harmless; equality is what decides identity of a key. ### What the interviewer is testing Not the trivia. They want to hear that a dict is a mapping over equality classes, that the hash invariant is the mechanism, that the first-inserted key object is the one kept, and that the practical defence is normalizing key types at the boundary rather than hoping values stay in one lane.
- After d = {1.0: 'a'} and then d[1] = 'b', which key object does the dict hold, and why does it matter?It still holds the `float` `1.0`; only the value was replaced. It matters at the edges of the system: serializing that dict emits `1.0` where a run with the opposite insertion order emits `1`, so a JSON payload, a log line or a dashboard label changes shape with no code change. `update()`, `|=` and `setdefault()` all follow the same keep-the-existing-key rule.
- Does float('nan') work as a dict key?You can store it, but only the exact object you inserted will find it again, because NaN is not equal to itself and the dict's identity check is what rescues the lookup. A freshly constructed NaN never matches. Since 3.10 distinct NaN objects hash by identity; before that they all hashed to 0, so a dict holding many NaN keys degraded toward a linear scan.
- How would you keep an int reading and a float reading apart in one mapping?Make the distinction part of the key rather than hoping the dict preserves it — for example key on `(type(value).__name__, value)`, or on an explicit tag from an enum. Better still, normalize at the parse boundary: pick one canonical key type (a str, a Decimal, or an int) and convert as data enters, so downstream code never sees two spellings of the same reading.
A dict is a filing cabinet indexed by what a value means, not by how it was written. "1", "1.0" and "one true" are three handwritings of the same number, so they all open the same folder — and the label on that folder is whichever handwriting filed it first.
saying these in an interview costs you the question
- Claims True is not a number so it cannot collide
- Says the newer key object replaces the stored one
- Thinks comparing an int with a float raises TypeError
- Assumes two different types can never share a slot
- Believes float keys are unsafe because of precision
- Says dicts compare keys by hash alone, never with ==