skip to content

Why does Python randomize str hashes per process, and when do you set PYTHONHASHSEED?

level: seniorimportance: nice to knowfreq 24%

answer

  1. A defence, not a performance feature
  2. Only string-like and datetime hashes move
  3. Seeded once per interpreter startup
  4. Sets vary run to run, dicts do not
  5. Never persist hash() across a process

basics

~20 s

Python seeds the hash for str, bytes and datetime objects randomly at interpreter startup to defeat collision-flooding attacks, so hash of a string differs every run. Set PYTHONHASHSEED only to reproduce order-dependent failures, never as a production fix.

solid answer

~40 s

Since 3.3 CPython picks a random hash seed at startup for `str`, `bytes` and `datetime` objects, so `hash("scrape_ok")` differs in every process; `sys.hash_info.algorithm` reports the algorithm and `sys.flags.hash_randomization` whether it is on. It defends against collision flooding, where an attacker who controls many key strings degrades dict inserts toward quadratic. Integers are not randomized — `hash(42)` is `42`. Two consequences matter: set iteration order varies between runs (dicts keep insertion order since 3.7, so they do not), and `hash()` must never be persisted into a cache name, shard index or database column. `PYTHONHASHSEED` is read at startup — `0` disables randomization, an integer pins it — and is for reproducing a flaky ordering bug, not for production. Use `hashlib` for anything that must survive a process.

code

console · 8 lines
console
$ python3 -c "print(hash('scrape_ok'))"
-4488815621594165800
$ python3 -c "print(hash('scrape_ok'))"
-4952339357242205410
$ PYTHONHASHSEED=0 python3 -c "print(hash('scrape_ok'))"
900072125980686718
$ PYTHONHASHSEED=0 python3 -c "print(hash('scrape_ok'))"
900072125980686718

go deeper

for a junior

Just know that hash of a string is not stable between runs and must never be written to a file or database. Reach for hashlib whenever a value has to outlive the process.

for a middle

Explain the why — collision-flooding defence — and the scope: str, bytes and datetime are seeded, integers are not, and set ordering varies while dict insertion order does not.

for a senior

Diagnose the symptoms: an intermittent order-dependent test, a shard or cache mapping that drifts after a single worker restarts, and a child process disagreeing with its parent under spawn or forkserver. Know that PYTHONHASHSEED is a reproduction tool, not a fix.

for a principal

Own the policy: hash() never crosses a process boundary anywhere in the codebase, stable digests are the sanctioned alternative, and a pinned seed is confined to CI reproduction runs rather than shipped into a service handling untrusted keys.

### What is randomized, and why Since **3.3**, CPython seeds the hash function for `str`, `bytes` and `datetime` objects with a random value chosen at interpreter startup. `hash("scrape_ok")` therefore returns a different integer in every new process. `sys.hash_info.algorithm` reports the algorithm (`siphash13` on a stock 3.14 build) and `sys.flags.hash_randomization` reports whether the seeding is active. The motivation is a denial-of-service class, not cryptography. Dict and set operations are O(1) *on average*, and degrade toward O(n) per operation when many keys land in the same region. An attacker who controls a lot of key strings — form field names, JSON object keys, HTTP header names, metric label names scraped from a remote target — could precompute thousands of strings that collide under a fixed hash function and turn one request into quadratic work. Randomizing the seed per process makes that precomputation impossible: the adversary cannot know this process's hash mapping. **Numbers are not randomized.** `hash(42) == 42` in every process, and the numeric hash invariant described by `sys.hash_info` is fixed. Only the string-like and datetime hashes move. ### The consequences you actually live with **Sets and set-derived orderings vary between runs.** Dicts have iterated in insertion order since **3.7**, so dict iteration is *not* affected — this is the nuance that separates a real answer from a memorized one. But `list({"cpu", "mem", "disk"})` genuinely differs run to run, and so does anything built from it: a rendered golden file, a comparison of `set` output, an "any element" pick. The classic symptom is a test that fails on roughly one CI run in five with no code change. **`hash()` is an in-process value with no persistence guarantee.** It is not stable across restarts, across processes, across interpreter versions or across builds. Writing it into a cache filename, a sharding column, a bloom filter on disk, or a partition key is a latent outage. Concretely: a metrics scraper distributes series across worker processes with `hash(series_name) % worker_count`. It looks correct in a single-process test and even survives a restart of the whole fleet, because everything is recomputed at once. Then one worker restarts on its own, computes a different mapping, and starts scraping series that another worker still owns — duplicate samples for some series, gaps for others, and a cache hit-rate collapse that shows up as a breach of the 92nd-percentile scrape-latency budget. The bug is not in the sharding math; it is in using a per-process value as if it were stable. ### `PYTHONHASHSEED` `PYTHONHASHSEED` is read **at interpreter startup**, so setting `os.environ["PYTHONHASHSEED"]` inside a running process does nothing — it must be in the environment of the command that launches Python. Its values: - `random` — the default: pick a fresh seed per process. - an integer in `[0, 4294967295]` — use that seed, making string hashes and therefore set ordering reproducible across runs. - `0` — disable randomization entirely. Legitimate uses are all about **reproducibility**: pinning a seed in CI to reproduce an order-dependent test failure, bisecting a flaky ordering bug, or producing a byte-identical build artefact. Pinning it permanently for a service that hashes untrusted input re-opens the collision DoS the feature exists to prevent — and, more mundanely, it hides an ordering bug rather than fixing it. The real fix for an order-dependent test is to sort the output or compare sets, not to freeze the seed. ### Child processes Seeding happens per interpreter, so the start method decides what a worker sees. A `fork` child inherits the parent's memory and therefore the parent's seed; `spawn` and `forkserver` start a fresh interpreter that seeds itself, so `hash("x")` differs between parent and child. That matters more on **3.14**, where the default `multiprocessing` start method on Unix other than macOS became `forkserver` (macOS and Windows remain `spawn`, and `fork` must now be requested explicitly). Code that quietly relied on parent and child agreeing on `hash()` under `fork` will start misbehaving on the new default. ### What to use instead Anything that must survive a process boundary needs a documented, stable digest: `hashlib.blake2b(name.encode(), digest_size=8)` or `hashlib.sha256` for correctness-critical or adversarial contexts, or a small non-cryptographic algorithm you own and version. Reserve `hash()` for what it is: an in-memory implementation detail of `dict` and `set`, valid for exactly as long as the process that produced it.

  • Does hash randomization make dict iteration order vary between runs?
    No. Dicts have iterated in insertion order since 3.7, and that is independent of the seed. Sets are the ones that move, along with anything derived from set ordering — a rendered golden file, a list built from a set, an arbitrary 'pick any element'. That is the usual source of a test that fails on one CI run in five, and the fix is to sort or compare as sets rather than to pin the seed.
  • A worker process computes a different hash for the same string than its parent — how?
    The seed is chosen per interpreter. A `fork` child inherits the parent's memory and therefore its seed, while `spawn` and `forkserver` start a fresh interpreter that seeds itself. On 3.14 the multiprocessing default is `forkserver` on Unix other than macOS, and `spawn` on macOS and Windows, so code that quietly relied on parent and child agreeing under `fork` breaks on the new default. Pin the seed before launch, or stop using `hash()` for that.
  • What should you use instead of hash() for a persistent shard key?
    A digest with a documented, stable definition: `hashlib.blake2b(key.encode(), digest_size=8)` for speed, or `hashlib.sha256` where the input is adversarial. Both are stable across processes, restarts and interpreter versions. `hash()` carries no such guarantee — it is an in-memory implementation detail of dict and set, valid only for the life of the process that produced it.

Every process shuffles its own deck before dealing keys into slots. That is what stops an attacker stacking the deck in advance — and also why a card's position from yesterday's deal means nothing today.

saying these in an interview costs you the question

  • Says randomization makes dict iteration order random
  • Persists hash() output in a cache key or database column
  • Thinks integer hashes are randomized too
  • Pins PYTHONHASHSEED in production to stop flaky tests
  • Believes setting os.environ mid-process changes the seed
  • Describes hash() as a cryptographic hash function

context