Why do the dict keys (1, 'seats') and (True, 'seats') collapse into a single entry?
answer
- Nothing about the container is at fault
- What do the elements say about each other
- Python's numbers are more equal than they look
- bool is a subclass of int
- Equal objects must hash equally
basics
~20 sTuples compare and hash element-wise, and in Python True == 1 == 1.0 with equal hashes. The two tuples are therefore equal keys: the second write overwrites the first value, while the dict keeps the key object it stored first.
solid answer
~50 sA tuple has **value semantics**: its `__eq__` compares element by element and its `__hash__` is built from the elements' hashes. So the keys are equal exactly when their components are, and Python's numeric tower makes `True == 1 == 1.0 == Decimal(1)` with matching hashes — booleans are literally a subclass of `int`. Two keys that look distinct in a log are therefore one dict entry. In a subscription-billing run that aggregates usage by `(quantity_flag, meter)`, where one upstream feed sends integers and another sends booleans or floats, buckets silently merge and totals come out short with no exception anywhere. Worse for forensics: on overwrite a dict keeps the **first** key object and the **last** value, so the log shows `(1, 'seats')` for a write that arrived as `(True, 'seats')`. The fix is to normalise key components to a canonical type at ingest.
code
python · 6 linesusage = {}
usage[(1, "seats")] = 10
usage[(True, "seats")] = 99
print(usage, len(usage))
print(hash(1) == hash(True) == hash(1.0), (1, "seats") == (1.0, "seats"))
print(("1", "seats") == (1, "seats"))go deeper
Recall that bool is a subclass of int, so True equals 1 in comparisons and as a dict key. You are not expected to have debugged this, only to not be surprised by it.
Explain the two rules that combine: a tuple's equality and hash come from its elements, and equal numeric values must hash equally. Show that this makes the two keys one entry rather than a collision.
Demonstrate the diagnosis on a real aggregation: spotting the count mismatch, projecting key component types instead of values, and knowing that the retained first key object misleads whoever reads the log.
Own where key shape is decided - one normalisation boundary, one key-building helper, and a policy on whether a surprising component type is coerced or rejected - so aggregate identity is a stated contract rather than an accident of the feeds.
### The mechanism Tuples have no identity of their own as far as a dict is concerned. `tuple.__eq__` compares position by position using the elements' own equality, and `tuple.__hash__` mixes the elements' hashes. Two tuples are the same dict key precisely when their components are pairwise equal — regardless of how, where or when they were constructed. Layer Python's numeric model on top of that. `bool` is a subclass of `int`, with `True` equal to `1`. Python's unified numeric hashing, in place since Python 3.2 (see `sys.hash_info`), guarantees that any two numeric values that compare equal also hash equal — across `int`, `float`, `bool`, `decimal.Decimal` and `fractions.Fraction`. So `hash(1) == hash(True) == hash(1.0)` and `(1, 'seats') == (True, 'seats')` is `True`. One dict entry, not two. This is not a bug in either feature: the hash contract *requires* equal objects to hash equally, and the numeric tower defines these values as equal. The surprise is only in the combination. ### What it looks like in a billing run Consider a nightly subscription-billing job that aggregates usage into `totals[(quantity, meter)]`. One upstream feed emits integer quantities; another, newer feed emits a boolean flag for its single-seat plans; a third round-trips through a float-typed column. Every one of those produces keys that are *equal* to the integer key, so three logically distinct buckets merge into one. Nothing raises. `len(totals)` is simply smaller than the number of distinct raw keys, some invoices are short, and the reconciliation report is the first place anyone notices. The forensic trap makes this worse. When you assign to an existing key, CPython **keeps the key object already stored** and replaces only the value. So after the merge, iterating the dict shows `(1, 'seats')` — the shape of the first feed — even though the surviving value came from the boolean feed. Anyone reading the log concludes the boolean feed never arrived. ### The mirror-image trap The same value semantics fail in the opposite direction just as quietly. `'1'` and `1` are *not* equal, nor are `b'EU'` and `'EU'`, nor `(1,)` and `1`. Those keys look identical in a log line — both render as `1` or `EU` — and stay stubbornly separate, splitting one bucket into two. Type drift across ingest paths therefore causes both symptoms at once: numeric variants merging where they should not, string-versus-bytes and string-versus-int variants splitting where they should not. A `float('nan')` component is the pathological extreme: it is unequal to itself, so a NaN-containing key can only be retrieved with the identical object that inserted it. ### Diagnosing it The cheap signal is a count mismatch: the number of records processed versus `len(totals)` versus the number of distinct keys you expected. To pin it down, project the *types* of the key components rather than their values — `collections.Counter(tuple(type(part).__name__ for part in key) for key in totals)` — and a single row of `('bool', 'str')` next to thousands of `('int', 'str')` tells the whole story in one line. Comparing two suspect keys directly with `==` and `hash()` confirms it. ### Fixing it, and where Normalise key components at the boundary, not at the point of use. Coerce each component to exactly one canonical type as records enter the pipeline — `int(quantity)`, `str(meter)` — and validate rather than coerce where a surprise type should be an error instead of a silent conversion. Build one small helper that constructs the key, so there is a single place that decides the key's shape, and no caller can invent a variant. Where the components genuinely span types, tag the key: include a short discriminator element, so `('int', 1, 'seats')` and `('bool', True, 'seats')` can never be equal. And if the key is really a domain concept rather than an anonymous pair, giving it a proper record type with an explicit `__eq__` and `__hash__` makes the equality rule reviewable instead of emergent. ### The interview one-liner Tuple keys are equal when their elements are, and Python says `True`, `1` and `1.0` are equal with equal hashes. That single fact merges buckets, hides the merge behind a retained first key object, and is invisible to every test that uses one consistent type.
- Which key object does the dict keep after the second write, and why does that matter?The one stored first. Assignment to an equal key replaces the value but leaves the existing key object in place, so iterating shows `(1, 'seats')` even when the surviving value arrived as `(True, 'seats')`. During an incident that reads as the second feed never having arrived, which sends the investigation in the wrong direction.
- How would you prove in production that key collapse is happening?Compare the record count with the number of distinct keys you expected and with len() of the aggregate. Then project the component types rather than the values - a Counter over tuples of type names per key - so a handful of ('bool', 'str') keys stand out against thousands of ('int', 'str'). Confirm the pair with == and hash().
- What is the mirror-image failure of the same rule?Keys that look identical in a log but are not equal, so one bucket splits into two: '1' versus 1, b'EU' versus 'EU', a one-element tuple versus a bare value. Type drift across ingest paths produces both symptoms at once, which is why normalising at the boundary fixes both.
- How would you make these keys structurally incapable of colliding?Normalise each component to one canonical type in a single key-building helper, so no caller can invent a variant. Where components genuinely span types, add a short discriminator element to the key so an int variant and a bool variant can never compare equal - or promote the key to a domain record type whose equality is written down rather than emergent.
Two order forms filed under the same customer number end up in one folder, however differently the number was typed. The folder keeps the first label it was given and the last document put inside it.
saying these in an interview costs you the question
- Says a hash collision overwrote the value
- Claims True is not an int in Python
- Assumes different reprs guarantee different keys
- Believes the dict stores the newest key object
- Thinks dict compares keys by identity
- Says 1 and 1.0 are always distinct dict keys