Why can two Python strings that look identical compare unequal with ==?
answer
- Equality on str is not equality of shapes
- The same accent has two legal spellings
- One code point, or a base plus a mark
- Composed and decomposed canonical forms
- unicodedata.normalize at the input boundary
basics
~20 sBecause == compares code points, not shapes. An accented letter can be one precomposed code point or a base letter plus a combining mark; both render the same but differ. unicodedata.normalize('NFC', s) folds them to one form before comparing.
solid answer
~40 sPython's `==` on `str` compares sequences of code points, so `'\u00e9'` (precomposed é) and `'e\u0301'` (e plus a combining acute) are unequal even though they render identically — and their `len()` values are 1 and 2. Unicode calls those two sequences **canonically equivalent**, and `unicodedata.normalize` is what makes them comparable: **NFC** composes to the shortest precomposed form, **NFD** fully decomposes into base plus marks. Both are lossless and round-trip. The fix is to normalize at the boundary — the moment text arrives from a form, a filename, a file, or another service — pick one form (NFC is the usual choice for storage and the web), and compare and store only normalized text. `unicodedata.is_normalized(form, s)` is a cheap pre-check that avoids rebuilding strings that are already in the target form.
code
python · 9 linesimport unicodedata
composed = 'é' # e-acute as one code point
decomposed = 'é' # e plus combining acute
print(composed == decomposed) # False
print(len(composed), len(decomposed)) # 1 2
print(unicodedata.normalize('NFC', decomposed) == composed) # True
print(unicodedata.is_normalized('NFC', decomposed)) # Falsego deeper
Recall that a Python str is a sequence of code points and that == compares those code points, so identical-looking text can differ. Naming unicodedata.normalize as the fix is what is expected at this level.
Explain the mechanics: precomposed versus base-plus-combining-mark, what NFC and NFD each produce, that both are lossless, and why len() disagrees. Expect to be asked where in the program the normalize call belongs.
Show the production judgment of normalizing once at every input boundary and storing a single form, plus the diagnosis path when a lookup fails intermittently: compare repr and len of the two values before blaming the storage layer.
Own the cross-cutting policy: one documented form for stored text, a defined match key for identity comparisons, and a plan for the day the interpreter's Unicode database advances and old and new keys must coexist.
## What `==` actually compares A Python `str` is a sequence of Unicode **code points**. `a == b` is true when the two sequences have the same length and the same code point at every position. It has no idea what the text looks like when rendered, which is why two strings that a human would swear are the same word can compare unequal. The commonest cause is **canonical equivalence**. Unicode encodes many accented letters twice: once as a single precomposed code point (U+00E9, é) and once as a base letter followed by a combining mark (`e` + U+0301 COMBINING ACUTE ACCENT). Renderers draw both identically, and the standard declares them equivalent — but they are different code point sequences, so `==` says no and `len()` reports 1 versus 2. ## The four normalization forms `unicodedata.normalize(form, s)` rewrites a string into one of four canonical shapes: - **NFD** — canonical *decomposition*: every precomposed character is split into base plus combining marks, and the marks are reordered into a canonical order. - **NFC** — canonical decomposition followed by canonical *composition*: the shortest equivalent form, which is what most keyboards, most web content and most databases already contain. - **NFKD** and **NFKC** — the same two shapes but using *compatibility* decomposition, which additionally rewrites characters that merely mean the same thing (a ligature into its letters, a superscript digit into a digit). Those are lossy and belong to a different decision. For the equality problem, NFC and NFD are the relevant pair. Both are **lossless**: normalizing to NFD and back to NFC returns the original, and two canonically equivalent inputs always produce identical output. Pick one and be consistent; NFC is the conventional choice because it is shortest and because most text already arrives in it. ## Where non-normalized text comes from It rarely comes from the keyboard. It comes from the edges of the system: filenames read from a filesystem whose convention is decomposed, text pasted from a document producer that composes differently, form input typed with dead keys, records exported by a service that normalized to the other form, and text assembled by your own code from pieces of different provenance. Two of those paths can feed the same field in the same table, and then a lookup written as an exact match finds a row half the time. ## The fix: normalize at the boundary Treat normalization the way you treat decoding: do it once, at the point where text enters the program, and keep everything inside in one known form. ```python import unicodedata def clean(s: str) -> str: return unicodedata.normalize('NFC', s) ``` If a value will be used to *decide identity* rather than to be displayed, fold its case as well and normalize again afterwards, because case folding can leave its output unnormalized. Keep that key separate from the display string — normalization for comparison and normalization for storage are the same operation here only because NFC is lossless. ## Cost, and how to avoid paying it Normalization walks the string and consults the Unicode database, so it is not free on hot paths. Two cheap guards help. `unicodedata.is_normalized(form, s)`, added in Python 3.8, answers the question without building a new string, and returns quickly for text that is already in the requested form. `str.isascii()`, added in Python 3.7, is cheaper still and is a complete answer: an ASCII-only string is already in every normalization form, so an early `if s.isascii(): return s` skips the work for the overwhelming majority of real input. ## What normalization does not fix It does not fold case — `'É'` and `'é'` remain different after NFC. It does not remove zero-width or formatting characters, which are their own code points and survive. It does not make two different letters that merely look alike compare equal; visually confusable characters from different scripts are genuinely different characters and no canonical form merges them. And it does not fix a `bytes` value: normalization is defined on text, so decode first. The interview-ready summary: `==` on `str` is code point equality, the same rendered text has more than one legal encoding as code points, `unicodedata.normalize` is the operation that collapses those spellings, and the right place to call it is the input boundary rather than every comparison site.
- Which normalization form would you store, NFC or NFD, and why?NFC by default. It is the shorter form, it is what most input already arrives in, and it is the form the web and most storage conventions expect, so it minimizes conversions. Either works because both are lossless and round-trip; what actually matters is that one form is chosen and applied at every boundary, so writes and lookups cannot disagree.
- Does unicodedata.normalize also fix case differences?No. Normalization only reconciles canonically equivalent spellings of the same characters; É and é stay distinct because they are different characters, not different encodings of one. Case-insensitive matching needs str.casefold() in addition, applied together with normalization since folding can leave its own output unnormalized.
- How would you avoid paying normalization cost on a hot path?Check first. str.isascii() is a fast scan and an ASCII string is already in every normalization form, so return it unchanged. Otherwise unicodedata.is_normalized(form, s) answers without allocating a new string. Beyond that, normalize once when the value enters the system and cache or store the normalized form rather than re-normalizing at each comparison.
Two spellings of the same surname on two forms: a human reads them as one person, a strict string match reads them as two, until you rewrite both into a single agreed spelling.
saying these in an interview costs you the question
- Believing == on str compares rendered appearance
- Assuming len() is the same for both spellings
- Thinking NFC and NFD lose information
- Normalizing at every comparison instead of at input
- Expecting normalization to also fold case
- Calling normalize on a bytes value without decoding