skip to content

Why do two Python str values that render identically sometimes compare unequal, and how does unicodedata.normalize fix it?

level: juniorimportance: should knowfreq 40%

answer

  1. Two spellings, one picture
  2. Equality compares code points, not glyphs
  3. Precomposed versus base plus combining mark
  4. Four forms; NFC is the interchange default
  5. unicodedata.normalize at the boundary, not at compare time

basics

~20 s

Python compares str values code point by code point, and the same visible text has several spellings: 'e' with an acute accent can be one code point or 'e' plus a combining accent. unicodedata.normalize('NFC', s) rewrites both into one canonical spelling first.

solid answer

~50 s

Equality on `str` compares code point sequences, not rendered glyphs. Unicode deliberately allows more than one encoding of the same character: U+00E9 is a precomposed e-acute, while U+0065 followed by U+0301 (COMBINING ACUTE ACCENT) draws the same thing from two code points. So `==` is `False` and even `len()` differs, while a terminal shows them as identical. `unicodedata.normalize(form, s)` converts to one of four forms — `'NFC'` composes to the shortest canonical spelling, `'NFD'` fully decomposes, and `'NFKC'`/`'NFKD'` add compatibility folding on top. The practical rule is to normalize to `'NFC'` once, at the boundary where text enters the system, and store and compare only the normalized value; `unicodedata.is_normalized('NFC', s)` is a cheap pre-check that avoids building a copy when the input is already canonical. Producers disagree by default, so an equality check that never normalizes eventually misses a match.

code

python · 8 lines
python
import unicodedata

composed = "café"                          # e-acute as one code point
decomposed = unicodedata.normalize("NFD", composed)   # e + U+0301

print(composed == decomposed, len(composed), len(decomposed))
print(unicodedata.is_normalized("NFC", decomposed))
print(unicodedata.normalize("NFC", decomposed) == composed)

go deeper

for a junior

Recall that a Python str is a sequence of code points and that == compares those, so the same-looking word can exist in two spellings with different len(). Name unicodedata.normalize and the 'NFC' form.

for a middle

Explain canonical equivalence, the composed/decomposed and canonical/compatibility axes behind the four forms, and why the call belongs at the boundary where text enters rather than inside a comparison helper.

for a senior

Show where in a real system the normalization is applied and enforced: what gets stored, what the uniqueness key is, and how non-canonical input is detected with unicodedata.is_normalized rather than silently rewritten deep in the code.

for a principal

Own the policy question of whether the system normalizes on input or rejects non-canonical input, since one loses the sender's exact bytes and the other pushes work onto clients. Decide once and make it uniform across every entry point.

## The model: str is a sequence of code points A Python `str` is a sequence of Unicode **code points**, and `==` on two `str` objects is a length-plus-element comparison over that sequence. Nothing in that comparison knows about fonts, glyphs, or how many marks end up stacked over one letter on screen. That is the whole source of the surprise: *rendering* is a many-to-one function of code points, so two different sequences can produce the same picture. Unicode makes this unavoidable rather than accidental. To stay round-trippable with the older national character sets it replaced, Unicode encodes many accented letters twice: once as a **precomposed** character (U+00E9 LATIN SMALL LETTER E WITH ACUTE) and once as a **base plus combining mark** (U+0065 LATIN SMALL LETTER E, then U+0301 COMBINING ACUTE ACCENT). The standard calls these two sequences **canonically equivalent**: they mean the same character and conforming software must treat them as the same text. Python's `==` does not, because `==` is defined on code points, so you get `len(a) == 4` and `len(b) == 5` for the same visible word. ## The four normalization forms `unicodedata.normalize(form, s)` returns a new `str` rewritten into a chosen normal form. There are four, and they are two independent choices crossed: * **Composed vs decomposed.** `'NFC'` pushes combining marks back into precomposed characters where one exists; `'NFD'` splits every precomposed character into base plus marks. Both preserve canonical equivalence — the text still means the same thing. * **Canonical vs compatibility.** `'NFC'`/`'NFD'` only apply canonical mappings. `'NFKC'`/`'NFKD'` additionally apply **compatibility** mappings, which are *lossy* in a formatting sense: a fullwidth letter becomes an ASCII letter, a ligature becomes its component letters, a superscript becomes a plain digit. That extra folding is powerful and dangerous, and it is a separate subject from simply making equal-looking text compare equal. For "these two strings should be the same string", `'NFC'` is the default answer. It is the form the web platform and most interchange formats assume, it is the shortest of the four for typical text, and it changes nothing for pure ASCII, so applying it costs you nothing on the common path. ## Where to put the call Normalization is a **boundary** operation, not a comparison-time operation. If you normalize inside your equality helper, then every dict key, every database row, every index entry and every log line still holds whichever spelling arrived, and the moment some other code path compares them directly — a `set` membership test, a `dict` lookup, a uniqueness constraint in storage — the mismatch is back. Normalize once where untrusted or foreign text enters (request parsing, file reading, decoding), keep the normalized form as the only form you store, and let ordinary `==` work from then on. `unicodedata.is_normalized(form, s)` answers "is this already in that form?" without allocating the converted copy, which matters when the answer is nearly always yes. It is also useful as a *policy* check rather than a fix: some systems prefer to reject a non-canonical input at the edge instead of silently rewriting it, so that the sender learns about the problem. ## Related pieces of the module `unicodedata` also exposes the character database itself: `unicodedata.name(ch)` gives the official character name, `unicodedata.category(ch)` gives the general category (`'Lu'`, `'Nd'`, `'Mn'` for a non-spacing combining mark, and so on), and `unicodedata.combining(ch)` returns the canonical combining class, which is nonzero exactly for the marks that reordering rules apply to. Those are the tools for *inspecting* why two strings differ, and printing `[unicodedata.name(c) for c in s]` for both sides is usually the fastest way to end an argument about a mystery mismatch. `unicodedata.unidata_version` reports which Unicode data release the interpreter was built against, which is why a normalization result can in principle differ between old and new interpreters for very recently added characters — stable, long-standing characters do not move. ## What this is not Normalization is not sanitization, not case handling and not a lookalike filter. Two characters can be visually confusable and still be entirely distinct after any normal form — a Cyrillic small letter that looks exactly like a Latin `a` is a different letter, canonically and compatibly, and NFC will happily leave both alone. Normalization only collapses the spellings Unicode itself declares equivalent. Treating it as "make this text safe" is the mistake that produces the next bug.

  • Which normal form would you apply to text arriving from an untrusted client, and why not one of the others?
    `'NFC'` for general text: it preserves meaning, is the form interchange formats assume, and is a no-op for ASCII. `'NFD'` is fine internally but stores more code points and surprises anyone who reads the data later. `'NFKC'` and `'NFKD'` apply compatibility folding, which rewrites characters into different characters — useful for identifier matching, wrong as a blanket transform for text you must reproduce faithfully.
  • Does normalizing a string make it safe to render or store?
    No. Normalization only collapses spellings Unicode declares equivalent. Visually confusable characters from different scripts survive every normal form untouched, and normalization does nothing about control characters, bidirectional overrides, or output encoding for whatever consumes the value. It is a comparison fix, not a sanitization step.
  • Why check unicodedata.is_normalized before calling unicodedata.normalize?
    `is_normalized` runs a quick-check without building the converted copy, so on the common path — text that is already canonical — you skip an allocation per value. It also expresses a different policy when you want it: instead of silently rewriting a non-canonical input, you can reject it at the edge so the sender finds out.

Two writers spell the same name with a single accented letter or with a letter plus a separately typed accent. The page looks the same; the keystroke logs do not, and Python is reading the keystroke log.

saying these in an interview costs you the question

  • Claims == on str compares rendered glyphs
  • Thinks two equal-looking strings must have equal len()
  • Uses NFKC as the default form for arbitrary text
  • Normalizes inside the comparison instead of at the boundary
  • Says normalization removes confusable or lookalike characters
  • Believes normalization is a sanitization step

context