skip to content

Why does joining a Python set of stop ids give a route-optimisation job a different cache key each run?

level: seniorimportance: should knowfreq 38%

answer

  1. Same code, different order each run
  2. Something is randomised per process
  3. String hashing is seeded at interpreter startup
  4. Sets place entries by hash, not insertion
  5. Sort the elements before joining the key

basics

~20 s

CPython randomises the hash of str and bytes per process, and a set stores entries by hash, so its iteration order differs between processes. Any key built from that order differs too. Join sorted(ids) instead.

solid answer

~40 s

Hash randomisation has been on by default since 3.3: the interpreter picks a seed at startup and folds it into every `str`, `bytes` and `datetime` hash, which makes hash-collision denial-of-service attacks impractical. Set and dict placement follows those hashes, so iterating a set of ids yields a different order in every process — and a key built as `"|".join(stop_ids)` therefore differs per worker. Nothing crashes; an 83% hit rate simply decays as the job spreads across workers or restarts, because old entries are filed under keys nobody recomputes. The fix belongs in the key: `"|".join(sorted(stop_ids))`, or `json.dumps(payload, sort_keys=True)` for a mapping, plus an explicit `encode("utf-8")` so a `str` key and a `bytes` key never diverge. `PYTHONHASHSEED` pins the seed for reproducing a failure locally, but it is a diagnostic, not a production fix.

code

console · 4 lines
console
python3 -c "print('|'.join({'lisbon','porto','faro','braga'}))"
python3 -c "print('|'.join({'lisbon','porto','faro','braga'}))"
PYTHONHASHSEED=0 python3 -c "print('|'.join({'lisbon','porto','faro','braga'}))"
PYTHONHASHSEED=0 python3 -c "print('|'.join({'lisbon','porto','faro','braga'}))"

go deeper

for a junior

Know that a set has no defined order and that anything you serialise from one should be sorted first. Recognise that a value differing between two runs of unchanged code usually means something in the process is randomised.

for a middle

Explain the mechanism: a per-process hash seed for str and bytes, set placement derived from hashes, and why sorted() or a sorted JSON dump makes the derived key stable everywhere.

for a senior

Diagnose from the symptom — a hit rate that decays across workers rather than an error — trace it to key construction, and know that PYTHONHASHSEED reproduces the failure but does not fix it, because pinning the seed restores the collision predictability it defends against.

for a principal

Own the standard for cross-process identity: which values are contractual keys, how they are canonicalised and encoded, and a CI policy of varying the hash seed so order dependence surfaces before a cache or a deploy exposes it.

### What actually varies CPython randomises the hash of `str`, `bytes` and `datetime` objects. The interpreter picks a hash seed during startup, and every hash of those types folds that seed in, so the same string hashes to a different value in two different processes. Sets and dictionaries place their entries by hash, so the *iteration order* of a set of strings is a function of that per-process seed (plus insertion history and resize points). Nothing about your code changed; the bucket layout did. This defence exists for a reason. Without it an attacker who can choose keys — form fields, JSON keys, header names — can craft thousands of strings that collide, turning an average O(1) dictionary into a quadratic one and taking a service down with a single request. Randomising the seed makes that set of collisions unguessable. It has been on by default since 3.3 and remains so on 3.14. ### Why the cache key drifts A route-optimisation job that builds its key like `"|".join(stop_ids)` from a `set` is asking the hash table for an order the language never promised. Each worker process starts with its own hash seed, so the *same* logical set of stops produces a different key string in each one. The result is not a crash but a slow bleed: an 83% cache-hit rate that only holds while a single process serves the traffic, and collapses as soon as the job is spread over several workers or restarted between deployments. The entries are all still there — they are simply filed under key strings nobody will ever compute again. ### The fix is in the key, not the environment Make the serialisation order-independent: ```python stops = {"lisbon", "porto", "faro", "braga"} key = "|".join(sorted(stops)) ``` `sorted()` gives a total order that does not depend on hashing at all. The same move applies to mappings — `json.dumps(payload, sort_keys=True)` — and to any digest computed over a collection. While you are there, pin the text encoding explicitly: a key hashed from `key.encode("utf-8")` in one place and from a differently-encoded byte string in another produces two distinct cache entries for the same logical value, which is the same bug in a different disguise. `str` and `bytes` never compare equal to each other either, so a cache written with one and read with the other misses every time. ### Where PYTHONHASHSEED belongs `PYTHONHASHSEED` is read at interpreter startup. Setting it to `0` disables randomisation altogether; setting it to an integer fixes the seed to that value. Two constraints follow. First, it must be set in the environment *before* the process starts — assigning to `os.environ` inside the running program has no effect on the seed already chosen, and only affects children you spawn afterwards. Second, it is a diagnostic and reproducibility tool, not a production fix: pinning the seed in production restores exactly the collision predictability the randomisation was introduced to prevent, and it leaves the underlying bug — code that depends on set order — in place. Use it to reproduce a failing CI run locally, and fix the ordering dependence in the code. `sys.flags` records whether randomisation was enabled for the current interpreter, which is worth checking before concluding that "it is deterministic here". ### Catching this class of bug earlier Order dependence hides well because a single-process test run is internally consistent. Two habits expose it. Run the suite under at least two explicit, different hash seeds in CI, so any test that silently relies on set order fails on one of them. And for anything that must be stable across processes — a cache key, a checksum, a file name, a snapshot fixture — build it in a subprocess or a second run and assert the two agree, rather than asserting only that the function returns *something*. ### The neighbouring myth Dictionaries preserve insertion order — an implementation detail in 3.6 that became a language guarantee in 3.7. Sets got no such guarantee and never will: their order is an artefact of hashing and resizing. A candidate who has generalised "dicts are ordered now" into "sets are ordered too" is carrying exactly the assumption that produces this defect.

  • Can you set PYTHONHASHSEED from inside the running program?
    No. The seed is chosen while the interpreter starts, so assigning to `os.environ` afterwards changes nothing for the current process — it only affects children you spawn later. To pin it you set the variable in the environment before launch, or re-execute the interpreter. `sys.flags` records whether randomisation was enabled for this run, which is worth checking before concluding that ordering looks stable.
  • Which objects does hash randomisation actually affect?
    `str`, `bytes` and `datetime` objects. Integers hash to values derived from the number itself, so a set of small ints tends to look stable across runs. That stability is an implementation detail, not a promise: insertion history and resize points still determine layout, so relying on the iteration order of any set is a defect waiting for a different input distribution.
  • How would you catch this class of bug before it reaches production?
    Run the suite under at least two different explicit hash seeds in CI, so any test that quietly depends on set order fails on one of them. And for anything that must be stable across processes — a cache key, a checksum, a generated filename — compute it in a second process and assert the two agree, rather than asserting only that the function returned something.

saying these in an interview costs you the question

  • Claims sets preserve insertion order like dicts do
  • Blames the cache library instead of key construction
  • Sets PYTHONHASHSEED=0 in production as the fix
  • Thinks assigning os.environ mid-run changes the seed
  • Assumes string hashing is stable across processes
  • Treats str and bytes keys as interchangeable

context