skip to content

When is unicodedata.normalize with NFKC the right choice, and what does it destroy?

level: seniorimportance: nice to knowfreq 20%

answer

  1. Two kinds of Unicode equivalence, not one
  2. The K forms trade fidelity for matchability
  3. Ligatures and superscripts flatten one way
  4. Right for an index, wrong for stored text
  5. Python folds source identifiers this way

basics

~20 s

NFKC applies compatibility folding on top of NFC: ligatures split into letters, superscripts become plain digits, full-width forms become ASCII, no-break space becomes a space. Use it for search and matching keys only — it is lossy and must never overwrite stored text.

solid answer

~50 s

The K forms trade information for matchability. **NFC/NFD** merge only *canonically equivalent* spellings and are lossless. **NFKC/NFKD** additionally fold *compatibility* characters — a ligature into its component letters, a superscript digit into a plain digit, full-width Latin into ASCII, a no-break space into a space — and that is a one-way conversion: the original distinction is gone. So NFKC is right for building a **match key**: a search index, a deduplication pass, or a normalized identifier where full-width and ASCII should collide. It is wrong for anything you will store or display, because it silently rewrites mathematical notation, spacing choices and typographic detail. Python itself uses this rule: the parser NFKC-normalizes identifiers, so a source name written with a ligature binds the same object as its plain-letter spelling. For caseless search keys, combine NFKC with `str.casefold()` and normalize again, since folding can leave its result unnormalized.

code

python · 7 lines
python
import unicodedata

print(unicodedata.normalize('NFKC', 'file'))        # file
print(unicodedata.normalize('NFKC', 'x²'))         # x2
print(unicodedata.normalize('NFKC', 'A1'))    # A1
print(unicodedata.normalize('NFKC', ' ') == ' ')   # True
print(unicodedata.normalize('NFC', 'x²') == 'x²')  # True, NFC keeps it

go deeper

for a junior

Recall that unicodedata.normalize offers four forms and that the two containing K do more than the others, folding characters that merely look or mean alike into plain equivalents.

for a middle

Explain the mechanics: canonical equivalence is lossless and reversible, compatibility equivalence is one-way, and name concrete foldings such as a ligature to letters or a no-break space to a plain space.

for a senior

Show the production judgment of storing the canonical form and deriving the compatibility key beside it, and of pairing NFKC with str.casefold() between two normalization passes for caseless lookup.

for a principal

Own the data-lifecycle question: which form is the value of record, who can rebuild derived keys when the interpreter's Unicode database advances, and where the boundary sits between normalization and an explicit character-policy check.

## Canonical versus compatibility equivalence Unicode defines two kinds of "the same". **Canonical equivalence** says two sequences encode the same character: a precomposed é and `e` plus a combining acute are the same letter, encoded differently. **Compatibility equivalence** is weaker and says two characters mean the same thing while differing in form: the ligature fi and the two letters `fi`, the superscript ² and the digit `2`, full-width A and ASCII `A`, a no-break space and an ordinary space. `unicodedata.normalize` exposes both. NFD and NFC apply canonical decomposition (and, for NFC, recomposition) and are **lossless** — you can round-trip between them forever. NFKD and NFKC apply *compatibility* decomposition first, and are **lossy**: after NFKC there is nothing left in the string that says the author wrote a superscript, a ligature or a no-break space. ```python import unicodedata print(unicodedata.normalize('NFKC', 'file')) # file print(unicodedata.normalize('NFKC', 'x²')) # x2 print(unicodedata.normalize('NFKC', 'A1')) # A1 print(unicodedata.normalize('NFKC', '\u00a0') == ' ') # True ``` ## What NFKC is for The K forms exist to build **keys**, not text. Three uses justify them: **Search and matching.** A user typing `file` should find a document typeset with the fi ligature, and a user typing an ASCII product code should find a record entered with full-width digits. Folding both sides through NFKC makes those collide by design. **Deduplication.** Two records that differ only by a no-break space or by full-width punctuation are the same record for the purpose of merging, and the folded key says so. **Identifier normalization.** Python itself does this: since Python 3.0 the parser normalizes identifiers to NFKC, so a name written with a ligature and the same name written with plain letters bind the identical object. That is a deliberate application of exactly this rule — the identifier is a key, not display text. ```python fi = 1 print(fi) # 1 -- the parser folded the identifier to NFKC ``` ## What it destroys Because compatibility decomposition is one-way, applying it to text you keep is data loss: - Mathematical and scientific notation flattens: exponents, subscripts and many symbol variants become ordinary characters, and a formula rendered from the folded text is wrong. - Deliberate spacing disappears: a no-break space inserted to keep a number with its unit becomes an ordinary space that can be broken across a line. - Typography flattens: ligatures and specialised forms become plain letters, which changes text a publisher chose deliberately. - Some scripts are affected far more than Latin, so a change that looks harmless on English test data can be destructive on real multilingual input. The rule that follows: **store the canonical form, derive the compatibility form.** Keep NFC text as the value of record, compute the NFKC key alongside it for indexing and lookup, and never write the key back over the value. ## Combining with case folding A case-insensitive search key needs both operations, and the order matters because case folding can produce output that is no longer normalized. Fold between two normalization passes: ```python import unicodedata def search_key(s: str) -> str: return unicodedata.normalize('NFKC', unicodedata.normalize('NFKC', s).casefold()) ``` ## The limits worth naming NFKC is not a sanitizer. It does not remove zero-width characters, it does not merge characters from different scripts that merely look alike, and it does not strip control characters. Treating it as a cleanup pass for untrusted input gives false confidence: whatever policy you need about confusable or invisible characters has to be written explicitly, typically by inspecting `unicodedata.category` and rejecting or stripping the categories you disallow. It is also not free, and the folded result depends on the interpreter's Unicode database — 16.0.0 in CPython 3.14, reported by `unicodedata.unidata_version`. Keys persisted under an older database are not guaranteed identical to keys computed today for rare characters, which is a real consideration for a long-lived index and an argument for being able to rebuild it. The one-sentence interview answer: NFC for what you keep, NFKC for what you look things up by, and never the reverse.

  • Why is NFKC unsafe as the stored form of user-entered text?
    Because compatibility decomposition is irreversible. Superscripts, ligatures, full-width forms and no-break spaces all collapse into plain equivalents, and nothing in the result records what was there before. Store NFC as the value of record and derive the NFKC key beside it; a system that overwrites the stored value has silently destroyed notation and spacing its users chose.
  • Does Python apply any normalization to source code identifiers itself?
    Yes. Since Python 3.0 the parser normalizes identifiers to NFKC, so a name written with a ligature or a full-width letter refers to the same object as its plain-letter spelling. It is the language applying the same rule recommended for match keys: an identifier is a lookup key, not display text, so compatibility folding is the correct choice there.
  • Is running NFKC over untrusted input a reasonable sanitization step?
    No, not by itself. NFKC folds compatibility characters but leaves zero-width and control characters in place and never merges lookalike characters from different scripts. Any policy about invisible or confusable characters has to be written explicitly — for example by inspecting unicodedata.category and rejecting the categories you disallow — rather than assumed from normalization.

NFC is correcting a spelling; NFKC is filing the document under a simplified code — useful for finding it, ruinous if you throw away the original.

saying these in an interview costs you the question

  • Calling NFKC a stricter but still lossless NFC
  • Storing the NFKC form as the value of record
  • Believing NFKC strips zero-width characters
  • Assuming NFKC merges lookalike characters across scripts
  • Not knowing the K forms are one-way
  • Applying NFKC to text that will be displayed

context