skip to content

Which Python built-in types have randomized hashes, and which stay stable?

level: middleimportance: nice to knowfreq 22%

answer

  1. Two families, only one is salted
  2. Bytes get a keyed function; numbers get arithmetic
  3. Containers inherit from what they hold
  4. Ask sys.hash_info rather than guessing
  5. hash(1) is 1, hash(-1) is -2

basics

~20 s

Only byte-oriented values are salted: str, bytes, read-only memoryview, and types whose hash derives from their bytes, such as datetime objects. Numbers are not — hash(1) is 1 in every process. Containers inherit salting from their elements.

solid answer

~50 s

The salt is applied where a hash is computed over raw bytes: `str`, `bytes`, hashable read-only `memoryview` objects, and types that hash their byte representation, notably `datetime.date` and `datetime.datetime`. Numeric hashing is a fixed mathematical rule instead — an `int` hashes to its value reduced modulo `2**61 - 1` on a 64-bit build, so `hash(1)` is `1` everywhere, and `float` and `complex` follow the same rule so that numerically equal values of different types hash equal. Immutable containers combine their elements' hashes, so `('a', 'b')` is salted transitively while `(1, 2)` is not, and the same holds for `frozenset`. Objects that inherit the default `__hash__` derive it from identity, so their values vary between runs for a different reason — address allocation, not the seed. `sys.hash_info` reports the width, modulus, algorithm and seed size in force.

code

pycon · 7 lines
pycon
>>> import sys
>>> hash(1), hash(1.0), hash(True), hash(-1)
(1, 1, 1, -2)
>>> hash(2 ** 61 - 1)
0
>>> print(sys.hash_info)
sys.hash_info(width=64, modulus=2305843009213693951, inf=314159, nan=0, imag=1000003, algorithm='siphash13', hash_bits=64, seed_bits=128, cutoff=0)

go deeper

for a junior

Remember the short version: text-like values have hashes that move between runs, numbers do not. You do not need the algorithm names, only the habit of never storing or comparing a hash() value across processes.

for a middle

Explain the two hashing families — a keyed function over bytes versus a modulus rule for numbers — and how tuples and frozensets inherit salting from their elements. Know that sys.hash_info is where you look the parameters up.

for a senior

Use the distinction diagnostically: it tells you whether a value that differs across processes is moving because of the salt or because of an identity hash, and therefore whether pinning the seed will reproduce the failure at all.

for a principal

Frame the rule for the codebase: hash() is an in-memory detail for every type, stable numeric values included, and nothing durable may be derived from it. Decide which stable digest the organization standardizes on for sharding and cache keys.

## The dividing line CPython has two quite different hashing stories, and the seed touches only one of them. **Byte-oriented hashing.** For `str` and `bytes`, the hash is computed by running a keyed pseudorandom function over the object's bytes, with the per-process salt as the key. `sys.hash_info` names the function; on Python 3.14 it is `siphash13`, which became the default in 3.11, replacing SipHash-2-4 introduced by PEP 456 in 3.4. The same path covers a hashable read-only `memoryview`, and it covers types that define their hash as the hash of a byte representation — `datetime.date`, `datetime.time` and `datetime.datetime` are the ones people notice, because a set of dates reorders between runs just as a set of strings does. **Numeric hashing.** Integers, floats, `complex`, `bool`, `decimal.Decimal` and `fractions.Fraction` all share one mathematical rule: the value is mapped into the field of integers modulo a prime, which is `2**61 - 1` on a 64-bit build. That rule is what makes `1 == 1.0 == True` hash identically, which they must, because they compare equal and equal objects are required to hash equally. `hash(7)` is `7`; `hash(2**61 - 1)` is `0`, since the modulus reduces it; `hash(-1)` is `-2`, because `-1` is reserved internally as an error signal. None of this involves the seed, so numeric hashes are the same in every process, on every machine of the same word size. ## Containers inherit `tuple` and `frozenset` do not hash their contents' bytes; they combine the hashes of their elements with a mixing function. That means salting propagates. `hash(('a', 'b'))` changes between runs because `hash('a')` does; `hash((1, 2))` does not, because neither element is salted. A frozenset of strings is randomized, a frozenset of integers is not. This is worth internalizing because it explains a puzzling asymmetry in real code. A cache keyed by `(customer_id, currency)` where both are integers appears to give stable hashes forever, and someone concludes hashes are stable in general. Change `currency` to the string `'EUR'` and the same expression starts moving between processes. ## Identity-based hashes An ordinary class that defines neither `__hash__` nor `__eq__` inherits `object.__hash__`, which derives a value from the object's identity. Those values also differ between runs, but for an unrelated reason: the object lives at a different address. No environment variable makes them stable, and pinning the seed does not touch them. Defining `__eq__` without `__hash__` sets `__hash__` to `None` and makes instances unhashable, which is a separate rule about the equality contract and belongs to the object-model material rather than here. ## What sys.hash_info tells you `sys.hash_info` is a named tuple describing the interpreter's hashing parameters: the width in bits, the modulus used for numeric hashing, the sentinel values used for infinity and for the imaginary unit, the algorithm name, the number of hash bits and seed bits, and the cutoff length below which a different algorithm would be used. Reading it is the honest way to answer 'is this build randomized and with what' rather than guessing from the version number, and it is how you would confirm an alternative implementation's behaviour. One nearby detail: since Python 3.10, `hash(float('nan'))` is derived from the object's identity rather than being a fixed constant, so two distinct NaN objects hash differently. That change was made so that a dict holding many NaN keys does not collapse into one collision chain. It is unrelated to the seed but it lands in the same conversation, because it is another case where 'the hash of this value' is not a pure function of the value. ## Why the distinction matters in practice Three consequences follow. First, tests. A set of integers often iterates in the same order every run, purely because small integers hash to themselves and land in predictable slots. Code that was 'proved' order-stable with an integer fixture breaks the moment the fixture becomes strings. Second, diagnosis. If you see a value that differs across processes, knowing which family it belongs to tells you whether the seed is responsible or whether you are looking at an identity hash — and therefore whether pinning the seed will reproduce the failure at all. Third, the boundary rule stays the same regardless: `hash()` of any type is an in-memory implementation detail. Even the values that happen to be stable today — integer hashes — are guaranteed only within an interpreter build, not across versions or word sizes. Anything that must persist or cross a process boundary belongs in `hashlib` or `zlib.crc32`.

  • Why must hash(1) and hash(1.0) be equal in the first place?
    Because `1 == 1.0` is true, and Python requires that objects comparing equal hash equally, or a dict keyed by `1` could not be looked up with `1.0`. The numeric hash rule maps every numeric type into integers modulo `2**61 - 1` on a 64-bit build precisely so that all numerically equal values across `int`, `float`, `complex`, `Decimal` and `Fraction` agree.
  • Does pinning the seed make an ordinary object's hash reproducible across runs?
    No. A class that inherits `object.__hash__` gets an identity-derived value, so it moves with the allocation address rather than the salt; the seed is irrelevant to it. If you need reproducible hashing for your own type, define `__hash__` over stable field values — and remember it will still only be stable within one interpreter process if any of those fields are strings.
  • How would you check at runtime whether the interpreter you are on randomizes hashes?
    Read `sys.hash_info`, which reports the algorithm name and the number of seed bits, and check `sys.flags` for the recorded hash-randomization setting. Comparing `hash('x')` across two freshly launched subprocesses is the empirical version of the same check and works on any implementation.

saying these in an interview costs you the question

  • Says every hash in Python is randomized
  • Expects hash(1) to differ between runs
  • Thinks a tuple of strings hashes stably because tuples are immutable
  • Confuses an identity-based hash with a salted one
  • Assumes integer hashes are portable across builds and versions
  • Believes datetime objects are exempt from the salt

context