skip to content

How do NFC and NFD differ from NFKC and NFKD in unicodedata.normalize?

level: middleimportance: should knowfreq 40%

answer

  1. Two axes: compose, and the K
  2. C and D round-trip; K does not
  3. K folds ligatures and formatting variants
  4. Store one form, match on another
  5. NFC to store, NFKC for keys

basics

~20 s

NFC and NFD are canonical: they only compose or decompose characters that are already equivalent, and they round-trip. NFKC and NFKD add compatibility folding, which rewrites formatting variants such as ligatures and Roman numerals and is lossy.

solid answer

~50 s

The four forms of `unicodedata.normalize` vary along two axes. **C versus D** is composition: D decomposes a precomposed character into base plus combining marks, C decomposes then recomposes. **With or without K** is the far bigger difference. The K forms apply *compatibility* mappings: the single-character ligature `\ufb01` becomes `fi`, the Roman numeral `\u2168` becomes `IX`, superscripts flatten to digits, full-width forms narrow. Those characters are not canonically equivalent to their replacements — they are formatting variants — so the mapping loses information and cannot be reversed. The rule of thumb: **store NFC**, because it preserves the user's text exactly and is the form most systems and the web expect. Use NFKC only to build a derived *matching key* — a search index entry, a lookup or comparison key — and keep the original alongside it. Choosing NFKC for storage means handing a user back text they did not type.

code

python · 8 lines
python
import unicodedata

s = "\ufb01le \u2168"           # 'fi' ligature, ROMAN NUMERAL NINE

print(unicodedata.normalize("NFC", s))    # unchanged
print(unicodedata.normalize("NFKC", s))   # file IX
print(unicodedata.is_normalized("NFC", s))  # True - NFC has no opinion here
print(len(s), len(unicodedata.normalize("NFKC", s)))  # 5 7

go deeper

for a junior

Know that unicodedata.normalize takes a form name and that NFC is the usual choice for storing text. Recognising that the K forms do something more aggressive is enough at this level.

for a middle

Explain both axes: composition versus decomposition, and canonical versus compatibility. Be able to give a concrete K example such as a ligature folding to two letters, and to say why that is not reversible.

for a senior

Show the operational rule: store NFC, derive an NFKC (usually casefolded) key for matching, keep both. Be ready to say where normalization runs on a hot path and why it is done once on write rather than per comparison.

for a principal

Own the platform-wide decision about which form is canonical in storage, how it is enforced across services and languages, and what the migration looks like for data already written in mixed forms.

## Two axes, four forms `unicodedata.normalize(form, text)` accepts `"NFC"`, `"NFD"`, `"NFKC"` and `"NFKD"`. The names decode cleanly: - **D = decomposition**: split precomposed characters into a base character plus combining marks, then sort the marks into canonical order. - **C = composition**: decompose first, then recombine base plus marks into precomposed characters where they exist. - **K = compatibility** (K for *kompatibility*, since C was taken): before composing or decomposing, apply compatibility mappings as well as canonical ones. So NFD and NFC are the canonical pair, NFKD and NFKC the compatibility pair. ## Canonical equivalence: lossless Two strings are **canonically equivalent** when they represent exactly the same written character with no difference in meaning or appearance — a precomposed e-acute versus `e` plus COMBINING ACUTE ACCENT, or the same base carrying two marks stacked in either order. NFC and NFD both map every member of such a group onto one representative, and no information is lost: text normalized to NFD can be normalized back to NFC and you get the original. Nothing a reader would notice changes. ## Compatibility equivalence: lossy on purpose **Compatibility equivalence** is a deliberately weaker relation: characters that mean the same thing but differ in formatting. Unicode encodes many of these for round-tripping with older character sets: - typographic ligatures — `\ufb01` (a single character) folds to the two characters `fi` - Roman numerals — `\u2168` folds to `IX` - superscripts and subscripts fold to their plain digits - full-width and half-width forms fold to the plain ASCII forms - some spaces fold to a plain space, and some fractions expand The K forms perform these mappings; the non-K forms leave them alone. Note what that means for `unicodedata.is_normalized("NFC", text)`: a string containing a ligature is already valid NFC — NFC has no opinion about it at all. Compatibility folding is **not reversible**. Once `\u2168` has become `IX` you cannot tell it from a genuine `IX` the user typed, and re-normalizing will never bring it back. That is the single most important operational fact about the K forms. ## Which form do you actually use? **Store NFC.** It is what the web platform recommends for interchange, it is what most text on most systems already is, it is compact, and it preserves the user's input character for character. Data written in NFC survives a round trip through your service unchanged. **Use NFKC to build a matching key, never as the stored value.** When you want a search box to find a title whose author typed a ligature, or a lookup to treat a full-width identifier as equal to its ASCII form, compute an NFKC (often plus `str.casefold()`) key, index that, and keep the original text next to it. The moment you overwrite the stored value with its NFKC form you have edited the user's data — a display name, a legal name, a document title — with no way to undo it. **NFD has narrow, real uses.** It is the form in which combining marks are visible as separate characters, which is what you want when you inspect or filter marks — for example dropping every character whose `unicodedata.combining(ch)` is nonzero, or whose `unicodedata.category(ch)` is `Mn`, to build an accent-insensitive index key. Some platform APIs also hand you filenames in NFD, which is a classic source of "the file is right there but the lookup fails". **NFKD is rarest**: compatibility folding plus visible marks, useful mainly as the first step of an aggressive fold-everything search key. ## Cost and caching Normalization walks the string and consults the character database, so it is not free, but the implementation is C and there is a quick-check fast path: `unicodedata.is_normalized` can often answer without allocating, and normalizing text that is already in the requested form is cheap. On a hot path — say a triage bot classifying every inbound ticket title against a 92nd-percentile latency budget — normalize once when the record is created rather than on each comparison, and store the derived key so later requests do no Unicode work at all. ## A precise way to say it in an interview "C and D differ only in whether marks are composed; that difference is invisible to a reader and fully reversible. K is the real fork: it folds formatting variants like ligatures and Roman numerals, which is lossy. So NFC for storage, NFKC for a derived matching key, and never the other way round."

  • Why is it wrong to store the NFKC form of a user's display name?
    Because compatibility folding is lossy and irreversible. A name containing a ligature, a full-width character or a Roman numeral comes back rendered differently from what the user typed, and there is no mapping home. Store NFC, which preserves the text exactly, and keep the NFKC value only as a derived key for search and comparison.
  • Is a string containing a ligature already in NFC?
    Yes. `unicodedata.is_normalized("NFC", text)` returns True for it, because NFC applies only canonical mappings and a ligature is not canonically equivalent to the letters it folds to. That is a common trap: passing an NFC check does not mean the text is free of compatibility variants.
  • How would you build an accent-insensitive lookup key?
    Normalize to NFD or NFKD so marks become separate characters, drop every character whose `unicodedata.combining(ch)` is nonzero (equivalently, whose `unicodedata.category(ch)` is `Mn`), apply `str.casefold()`, then normalize the result back to NFC. Keep the original text; this key is only for matching, and it is language-insensitive, so it is a heuristic rather than a correct collation.

Canonical forms are like rewriting a handwritten address in neat block capitals: same information, reversible. Compatibility folding is like retyping it as plain text with all the letterhead and typography thrown away: useful for matching, impossible to restore.

saying these in an interview costs you the question

  • Says the four forms differ only in composition
  • Treats NFKC as the safe default for storage
  • Thinks compatibility folding can be reversed
  • Believes NFC removes ligatures or full-width characters
  • Calls NFD wrong rather than differently useful

context