skip to content

A payment reconciliation job shards accounts with hash(account_id) % 8 — why does it lose rows?

level: seniorimportance: should knowfreq 34%

answer

  1. The invariant assumed was never promised
  2. Works in one process, breaks in two
  3. The salt is per interpreter, per run
  4. Inheritance depends on the start method
  5. hashlib or crc32 for anything durable

basics

~20 s

Because hash() of a string is salted per interpreter process. Each worker and each restarted run maps the same account to a different shard, so rows written under one shard are searched for under another. Use hashlib.blake2b instead.

solid answer

~50 s

`hash()` is only defined within one interpreter process: the salt applied to `str` hashing is drawn fresh at start-up, so the same `account_id` maps to a different shard in a worker process, in a restarted run, or on another machine. The job therefore writes a row into shard 3 today and searches shard 6 for it tomorrow, and a resumed run re-partitions everything it has already emitted. The tell is that it works when everything runs in one process and breaks under a process pool or after a restart. The fix is a stable function of the bytes: `hashlib.blake2b(account_id.encode('utf-8'), digest_size=8)` reduced with `int.from_bytes`, or `zlib.crc32` where a checksum-grade spread is enough and the per-row budget is tight. Keep `hash()` for in-memory `dict` and `set` membership only; never persist it, name a file with it, or send it across a process boundary.

code

python · 14 lines
python
import hashlib
import zlib


def shard_digest(account_id: str, shards: int) -> int:
    digest = hashlib.blake2b(account_id.encode('utf-8'), digest_size=8).digest()
    return int.from_bytes(digest, 'big') % shards


def shard_checksum(account_id: str, shards: int) -> int:
    return zlib.crc32(account_id.encode('utf-8')) % shards


print(shard_digest('ACCT-4417', 8), shard_checksum('ACCT-4417', 8))

go deeper

for a junior

Take away the rule rather than the diagnosis: hash() belongs to dicts and sets inside one running program. If a value has to mean the same thing in another process or tomorrow, compute it with hashlib instead.

for a middle

Explain why a second process changes the answer, and be able to show it with two subprocesses. Know that hashlib.blake2b or zlib.crc32 over explicitly encoded bytes is the stable replacement, and why the encoding must be explicit.

for a senior

Diagnose before fixing: prove the divergence across processes, rule out lookalike causes such as a writer left unclosed, then replace the shard function and add a subprocess test that pins the mapping. Be able to relate the start method to whether children inherit the salt.

for a principal

Own the boundary rule and make it reviewable: nothing derived from hash() may cross a process, a restart or a machine. Decide the standard digest, where the shard count and algorithm are recorded, and how a future re-partition is rolled out without corrupting reconciled history.

## The failure A reconciliation job partitions the day's transactions across eight workers by `shard = hash(account_id) % 8`, and each worker writes matched and unmatched rows into its own output file. Totals come out short, and rerunning produces a different set of missing accounts each time. The cause is that `hash()` of a `str` is salted with a secret drawn per interpreter process. Every process in the pipeline — each pool worker, the driver, tomorrow's rerun, the machine in the other region — computes a different hash for `'ACCT-4417'` and therefore a different shard number. A row emitted into shard 3 by yesterday's run is looked for in shard 6 by today's. Worse, a partial rerun re-partitions the records it has already written, so accounts are double-counted in one shard and absent from another. Nothing in the code is wrong to read; the invariant it relies on — that a given account always maps to the same bucket — is simply not one that `hash()` provides. ## Why it looked fine Two things mask it in development. Single-process runs are self-consistent. The salt is fixed for the life of the process, so a job that shards, writes and reconciles inside one interpreter never sees a mismatch. The bug appears only when a second process is involved. And the way children are created matters. A child created with the `fork` start method inherits the parent's memory, salt included, so parent and children agree on `hash()`. A child created with `spawn` or `forkserver` is a fresh interpreter with its own salt, so it does not. That distinction is why this class of bug has a habit of appearing after an upgrade or a platform move rather than at the moment the code is written: since Python 3.14 the default start method on Unix platforms other than macOS is `forkserver`, macOS and Windows have used `spawn` for years, and `fork` must now be asked for explicitly. Code that quietly depended on inherited salts on Linux changes behaviour on upgrade. ## Confirming it before you fix it Do not guess. Print `hash('ACCT-4417')` from two freshly launched subprocesses and show they differ; then print it from a pool worker and from the driver in the same run and show that too. That takes a minute and it distinguishes this cause from the other candidates — a partitioner that drops rows on a boundary, a worker that dies mid-write, or an output file that was never flushed. That last one deserves a note: when a shard worker is killed part-way, an output file left unclosed keeps its final buffered block unwritten, which produces short totals that look exactly like a partition bug. Rule it out by checking that every writer runs under a `with` block or an explicit close, so a crash cannot be confused with a mis-shard. ## The fix Derive the shard from a function that is stable everywhere. `hashlib.blake2b(account_id.encode('utf-8'), digest_size=8)` gives eight bytes you can turn into an integer with `int.from_bytes` and reduce modulo the shard count; it is fast, it is in the standard library, and it produces the same number on every interpreter, version and machine. `zlib.crc32` over the same encoded bytes is cheaper still and is fine where you need spread rather than collision resistance — worth knowing when the per-batch latency budget is tight and you are measuring at a high percentile rather than at the mean, because a per-row digest that is invisible in the average can still show up at the 92nd percentile of a large batch. Three details make the fix durable. Encode explicitly — `str` has no bytes until you choose an encoding, and hashing an implicitly encoded value is how the same account gets two shards on two locales. Pin the shard count somewhere the readers and the writer share, because changing it re-partitions everything just as surely as a changing salt does. And write the shard number into the output alongside the row, so a mismatch is detectable rather than silent. ## The general rule `hash()` exists to place objects in the in-memory `dict` and `set` implementations. Its contract is limited to one process: not stable across runs for text, not guaranteed across interpreter versions or word sizes even for the values that happen to be stable today, and not something an alternative implementation must reproduce. So the boundary is easy to state and easy to review for: any hash that outlives the process — a shard key, a cache key in a shared store, a filename, a partition column, a bloom filter written to disk, a deduplication fingerprint — comes from `hashlib` or a checksum, never from `hash()`. The same rule catches a subtler variant: `PYTHONHASHSEED` pinned in the deployment to 'fix' this. It papers over the symptom, ties correctness to how the process was launched, and re-opens the collision denial-of-service for any part of the service that parses untrusted keys. Fix the function, not the environment.

  • Why does the same code appear to work when the pool uses the fork start method?
    A forked child is a copy of the parent's memory, so it inherits the already-chosen salt and computes identical `hash()` values. That makes the defect invisible until something changes the start method — a move to macOS or Windows, a switch to `spawn` for thread safety, or the 3.14 default of `forkserver` on other Unix platforms. It is coincidence, not a guarantee, and it never survives a restart of the parent.
  • Would exporting PYTHONHASHSEED for the whole pipeline be an acceptable fix?
    No. It makes correctness depend on how every process was launched, so a cron entry, a container override or an isolated start-up mode that ignores the environment silently re-breaks partitioning. It also disables the collision defence for any process in the pipeline that parses untrusted keys. A three-line stable digest removes the dependency entirely.
  • Where else does an accidental dependency on hash() typically hide?
    Filenames and cache keys built from `hash()`, deduplication fingerprints persisted between runs, bloom filters written to disk, partition columns pushed into a warehouse, and anything logged in one process and matched in another. A quick grep for `hash(` near serialization or path construction usually finds them; anything crossing a process boundary should be using `hashlib` or a checksum.
  • How would you keep this from recurring across the codebase?
    Put the shard function behind one named helper that everything imports, make it take bytes or encode explicitly, record the shard count and the algorithm in the output metadata, and add a test that computes the shard in a freshly launched subprocess and asserts it matches the in-process value. That test fails immediately if someone reintroduces `hash()`.

It is like numbering warehouse aisles with a code that each shift's supervisor invents fresh: everything the shift files is findable that day, and unfindable to the next shift.

saying these in an interview costs you the question

  • Blames the modulus or negative hash values rather than the salt
  • Fixes it by exporting PYTHONHASHSEED across the pipeline
  • Believes hash() is stable because the type is immutable
  • Thinks the salt changes within a running process
  • Persists hash() output in filenames or a shared cache
  • Cannot say why a forked child agrees with its parent

context