Why does hash('abc') differ between two Python processes, and what does PYTHONHASHSEED=0 cost?
answer
- Something is seeded at interpreter startup
- Only text-like objects are affected
- Defends against crafted colliding keys
- Numbers hash to themselves regardless
- PYTHONHASHSEED pins or disables it
basics
~20 sCPython mixes a random per-process seed into the hash of str and bytes so an attacker cannot precompute keys that all collide in a dict. PYTHONHASHSEED=0 turns that randomization off, and removing it re-opens the collision-flood attack.
solid answer
~50 sSince Python 3.3, hash randomization is on by default: the interpreter draws a random seed at startup and mixes it into the hash of `str` and `bytes` (and anything derived from them, such as a tuple of strings). Without it, an attacker who knows the hash function can craft thousands of distinct keys that all land in the same bucket, so building a dict from client-controlled keys — JSON object keys, form fields, headers — degrades into quadratic work. Numbers are not randomized: `hash(1234)` is 1234, by design, so integer keys get no protection from this. The consequences for your own code are that `hash()` values must never be persisted, shared between processes or used for sharding, and that set iteration order changes between runs. `PYTHONHASHSEED=0` disables randomization and any other value pins the seed; both are debugging tools, not production settings.
code
console · 4 lines$ python3 -c 'print(hash("payroll"))'
$ python3 -c 'print(hash("payroll"))' # a different value
$ PYTHONHASHSEED=0 python3 -c 'print(hash("payroll"))'
$ PYTHONHASHSEED=0 python3 -c 'print(hash("payroll"))' # identical to the line abovego deeper
Recall that hash() of a string is not stable across runs and must never be stored or compared between processes. If you need a stable fingerprint, reach for hashlib instead.
Explain what the seed defends against, which types it touches and which it deliberately does not, that PYTHONHASHSEED is read only at interpreter startup, and why set iteration order rather than dict order is what visibly varies.
Show that you would treat a pinned seed in a shared configuration as a finding, fix the order-dependent test at its source, and point out that randomization is no substitute for capping how many client-controlled keys a request may carry.
Own where such settings are allowed to live. Argue for keeping debugging-only environment variables out of images that also serve traffic, and for making key-count limits a platform-level control rather than something each service remembers.
## What randomization is defending A dict or a set places a key by its hash. When many distinct keys share a hash, the container has to fall back on comparing them, and the cost of inserting n such keys grows quadratically rather than linearly. That is fine as an accident and dangerous as a choice: if the hash function is fixed and public, an attacker can compute a large set of distinct strings that all collide and send them as the keys of one request — object keys in a JSON body, fields in a form, header names, query parameters. The receiving service does no more work than "parse a request into a dict", and that request costs seconds of CPU instead of microseconds. The defence is to make the function unpredictable per process. CPython draws a random seed at interpreter startup and mixes it into the hash of text-like objects. An attacker can still cause collisions inside a *single* known process if they can observe hashes, but they cannot precompute a payload that works against a process they have never seen, which is what makes the attack cheap. This has been the default since Python 3.3. The algorithm is pluggable — `sys.hash_info` reports which one the build uses — and the default has been SipHash since 3.4, with SipHash-1-3 replacing SipHash-2-4 in 3.11 for speed. Note the goal: it is a *keyed* hash for collision resistance against an adversary who cannot see the key. It is not a cryptographic digest, and `hash()` must never be used where you want one; that is what `hashlib` is for. ## What is randomized, and what is not Randomized: `str`, `bytes`, and by extension every object whose hash is derived from them — a tuple of strings, a frozenset of strings, a `datetime` value. Not randomized: numbers. `hash(1234) == 1234`, and small floats hash to the same value as the equal int. This is required by the language's own rule that objects comparing equal must hash equal, so that `1`, `1.0` and `True` interchange as dict keys. The security consequence is worth saying out loud: **if your attacker-controlled keys are integers, hash randomization gives you nothing**, and the bound has to come from limiting how many keys you accept in the first place. That last point generalizes. Randomization raises the cost of one specific attack; it is not a substitute for a cap on the number of fields, keys or headers a request may carry. ## What it costs you Because the seed changes every run, `hash()` is only meaningful inside one process's lifetime. Three rules follow: - **Never persist a `hash()` value.** Not as a cache key in a shared store, not as a column, not in a file. It will not match tomorrow. Use `hashlib.sha256` and take a digest if you need a stable fingerprint. - **Never shard on `hash()` across processes.** `hash(key) % n` routes differently in each worker, because each worker drew its own seed at startup. - **Set iteration order varies between runs.** Dict iteration has been insertion-ordered since 3.7, so dicts look stable; sets do not, and that is where the flakiness lands. ``` $ python3 -c 'print(list({"a", "b", "c", "d", "e"}))' ['c', 'b', 'd', 'a', 'e'] $ python3 -c 'print(list({"a", "b", "c", "d", "e"}))' ['d', 'b', 'e', 'a', 'c'] ``` ## The trap worth naming A CSV import for a payroll system has one test that asserts on the order of a set of validation messages. It fails about one run in five. The 27-minute suite is already the slowest thing in the pipeline, so nobody wants to re-run it, and someone finds that `PYTHONHASHSEED=0` makes the ordering stable. It goes into the CI configuration, the fix works, the flake stops. Then that environment block is copied — as environment blocks always are — into the deployment manifest, and production now runs every worker with a known, fixed hash function. Nothing breaks. No test fails. The collision-flood defence is simply gone, and the only evidence is one line in a YAML file that reads like a test setting. The fix for the flaky test was never the seed. Sort the messages, or compare sets to sets rather than lists to lists. If you genuinely need determinism to reproduce one confusing failure, set the variable for that single command, not for the pipeline and never for an image that also ships to production. Two mechanical details. `PYTHONHASHSEED` is read when the interpreter starts, so assigning `os.environ["PYTHONHASHSEED"]` inside a running program does nothing to that program; it affects only children you launch afterwards. And the seed is per interpreter process — separately started workers do not share one, so anything that requires several processes to agree on a hash must use an explicit stable function rather than the built-in. ## Reviewing for it Three questions cover almost all of it. Does any stored value or cross-process decision come from `hash()`? Does any test depend on set iteration order? And is `PYTHONHASHSEED` set anywhere outside a developer's own shell? A "yes" to the third is worth a conversation every time.
- Why does dict iteration look stable across runs while set iteration does not?Dicts have preserved insertion order since 3.7 — iteration walks the entries in the order they were added, which has nothing to do with the hash seed. Sets have no such guarantee and iterate in bucket order, which the seed determines. That is why order-dependent tests almost always turn out to be asserting on a set, or on something derived from one.
- Does hash randomization protect a service whose request keys are integers?No. `hash(n)` for an integer is derived from its value and is not seeded, because the language requires equal numbers to hash equally across int, float and bool. If a client controls integer keys, they can still craft collisions. The defence there is a cap on how many keys a request may carry, which is the control you want regardless of key type.
- Someone sets PYTHONHASHSEED=0 in the CI config to stop a flaky test. What do you say?That it hides the real bug and travels badly. The test is asserting on set iteration order; sort the values or compare sets instead. Pinning the seed makes the assertion pass without making it correct, and environment blocks get copied into deployment manifests, at which point production runs with a known hash function and loses the collision-flood defence.
It is a bank that reshuffles which counter serves which surname every morning: a queue-jammer who studied yesterday's layout arrives with a plan that no longer applies.
saying these in an interview costs you the question
- Thinks randomization changes dict iteration order in 3.7+
- Sets PYTHONHASHSEED=0 in a production image
- Persists hash() values as cache keys or columns
- Believes integer keys are randomized too
- Calls hash() cryptographically secure
- Assumes all worker processes share one seed