skip to content

How do you hash a text address with hashlib so two services agree on the digest?

level: middleimportance: should knowfreq 38%

answer

  1. Bytes, and the text did not pick them
  2. Two hosts, two byte forms, no error
  3. Same glyph, different code points
  4. Compose first, then encode
  5. unicodedata.normalize before str.encode

basics

~10 s

Encode explicitly and identically everywhere: pin one encoding (UTF-8), normalise the text first, and hash the resulting bytes. hashlib refuses a str with TypeError precisely because the encoding choice changes the digest.

solid answer

~40 s

`hashlib` hashes bytes, so `hashlib.sha256("Münchner Straße")` raises `TypeError: Strings must be encoded before hashing`. The fix is not just `.encode()` — it is agreeing on a **canonical byte form** across every producer and consumer. Encoding the same text as UTF-8 and as Latin-1 yields two completely different digests for any non-ASCII character, and Unicode gives you a second trap: a composed `ü` (U+00FC) and a decomposed `u` plus a combining diaeresis are the same text and different bytes. So the recipe is: normalise with `unicodedata.normalize("NFC", s)`, apply whatever case and whitespace folding the domain requires, then `.encode("utf-8")` — and write that rule down, because a service that skips a step produces a digest nobody else can reproduce.

code

python · 13 lines
python
import hashlib, unicodedata

addr = "M\u00fcnchner Stra\u00dfe 5"
print(hashlib.sha256(addr.encode("utf-8")).hexdigest()[:16])
print(hashlib.sha256(addr.encode("latin-1")).hexdigest()[:16])

try:
    hashlib.sha256(addr)
except TypeError as exc:
    print(exc)

canonical = unicodedata.normalize("NFC", addr).casefold().encode("utf-8")
print(hashlib.sha256(canonical).hexdigest()[:16])

go deeper

for a junior

Remember that hashlib takes bytes: call .encode("utf-8") on the text yourself, and know that the TypeError means Python will not guess an encoding for you.

for a middle

Explain why the encoding is part of the digest, show that UTF-8 and Latin-1 diverge on any non-ASCII character, and add unicodedata.normalize for text that looks identical but uses different code points.

for a senior

Diagnose the silent version of this: cache misses or duplicate rows with no exception. Compare encoded bytes rather than strings, log encoded length beside the digest, and fix the producer so exactly one canonical form exists.

for a principal

Own the canonical-form contract across services: one documented normalisation and encoding, one shared serialiser implementation, and a plan for re-keying stored digests if the contract ever changes.

### Why the TypeError is a feature `hashlib.sha256("abc")` does not hash the text; it raises `TypeError: Strings must be encoded before hashing`. Python could have picked UTF-8 silently, and a lot of people wish it had. It refuses because the digest of a `str` is not well defined: a digest is a function of bytes, and turning text into bytes is a choice with several defensible answers. Making you write `.encode("utf-8")` puts that choice in your source, where a reviewer can see it and a second service can copy it. Consider a geocoding batch whose services share a content-addressed cache keyed by the digest of a normalised address. One service was written on a host that reads its input files with the platform's default encoding and re-encodes as Latin-1; the rest use UTF-8. Every ASCII-only address agrees. Every address containing `ü`, `é` or `ß` produces two different cache keys, so those addresses miss the cache forever and are re-geocoded on every run. Nothing errors. The symptom is a cost line and a latency tail, not a stack trace — which is why encoding mismatches survive so long in production. ### Two layers of canonicalisation **Layer one is the encoding.** Pick UTF-8 and pin it explicitly at every boundary. Never rely on a default: `str.encode()` defaults to UTF-8, but `open()` in text mode historically followed the locale, and anything that round-trips through another system may not have been UTF-8 at all. Passing the encoding by name in every call is one word of typing and removes a whole class of "works on my machine". **Layer two is Unicode normalisation.** Unicode can represent the same character in more than one way. `"ü"` may be the single code point U+00FC, or `"u"` followed by U+0308 COMBINING DIAERESIS. They render identically, compare unequal as `str`, and encode to different bytes — so they hash differently. `unicodedata.normalize("NFC", s)` folds text into a single composed form; `"NFKC"` goes further and also collapses compatibility variants such as a full-width digit into its ASCII equivalent. NFC is the usual choice for identifiers; NFKC when your inputs come from varied keyboards and you want aggressive equivalence. On top of that sits domain canonicalisation, which is not Unicode's problem but yours: trimming whitespace, collapsing runs of spaces, `casefold()` when case is not significant, ordering fields. `str.casefold()` rather than `str.lower()` matters here — it handles cases such as `ß` folding to `ss` that `lower()` leaves alone. ### The general rule: hash a serialisation, not an object The same discipline scales past single strings. To key a cache on a record, do not hash `repr(obj)` or `str(obj)`: repr is not a stable contract, dict ordering feeds through it, and float formatting can vary. Serialise deliberately — for JSON, `json.dumps(obj, sort_keys=True, separators=(",", ":"), ensure_ascii=False)` then `.encode("utf-8")` — so the byte form is a documented function of the values. Write that serialiser once and share it; two implementations of "the canonical form" will diverge. A related trap: Python's builtin `hash()` is not this. It is randomised per process for `str` and `bytes` unless `PYTHONHASHSEED` is fixed, so it is fine for an in-process dictionary and useless as a stored or shared key. Anything that must be reproduced by another process or another host goes through `hashlib`. ### Diagnosing a mismatch When two sides disagree, compare the **bytes**, not the strings. Print `s.encode("utf-8")` on both sides and look at the escapes; a Latin-1 `ü` shows up as a single `\xfc`, UTF-8 as `\xc3\xbc`, and decomposed text as an extra `\xcc\x88` after an ASCII `u`. Logging the length of the encoded bytes alongside the digest is a cheap permanent tripwire — it catches the normalisation difference without logging the data itself. Once you have found it, fix the producer to emit the canonical form rather than teaching the consumer to accept both, or you will be maintaining two canonical forms forever. ### The filesystem-path special case Paths are the one place where "decode, normalise, encode" can lose information. A filename on a POSIX filesystem is a bag of bytes that need not be valid UTF-8; Python surfaces the undecodable parts as surrogate code points, and re-encoding those with plain `.encode("utf-8")` raises `UnicodeEncodeError`. Use `os.fsencode()` when the thing you are hashing is a path: it applies the filesystem encoding with the surrogate-escape error handler, so the bytes you hash are the bytes the kernel holds. And note that normalisation is actively wrong here — some platforms store decomposed filenames, so folding a path to NFC before hashing would produce a digest for a name that does not exist on disk.

  • Why not key a cache on Python's builtin hash() of the string instead?
    Because hash() of a str is randomised per process unless PYTHONHASHSEED is set, so two processes disagree and the same process disagrees after a restart. It is also a short, non-cryptographic value with no collision guarantees. Any digest that crosses a process, a host or a storage boundary belongs to hashlib.
  • When would you choose NFKC over NFC before hashing?
    NFC only composes characters that have a composed form, preserving distinctions such as full-width versus ASCII digits. NFKC additionally folds compatibility variants together, so visually equivalent inputs from different keyboards or legacy systems collapse to one form. Choose NFKC when you want aggressive equivalence for user-entered identifiers, and NFC when the distinctions carry meaning.
  • How would you hash a whole record rather than a single string?
    Define a canonical serialisation and hash that. For JSON, json.dumps with sort_keys=True and fixed separators, then .encode("utf-8"); never hash repr() or str() of an object, whose format is not a stable contract. Implement the serialiser once and share it, because two independent versions of the canonical form will eventually diverge.

saying these in an interview costs you the question

  • Expects hashlib to encode a str automatically
  • Relies on the platform default encoding at any boundary
  • Ignores Unicode normalisation for equal-looking text
  • Uses the builtin hash() for a stored or shared key
  • Hashes repr() of an object as a cache key
  • Fixes a mismatch in the consumer instead of the producer

context