What does unicodedata.normalize('NFKC', s) fold that NFC leaves alone, and why can that turn a rejected username into a reserved one?
answer
- The K stands for compatibility
- Fullwidth, ligatures, circled digits collapse
- Lossy: you cannot fold back
- The bug is check-then-fold ordering
- Python folds source identifiers with NFKC itself
basics
~20 sNFKC adds compatibility mappings on top of NFC: fullwidth letters, ligatures, superscripts, circled digits and abbreviation characters all collapse to plain ASCII equivalents. A string that failed a reserved-name check before folding can therefore equal a reserved name after it.
solid answer
~50 s`'NFC'` applies only canonical mappings — spellings Unicode says are the same character. `'NFKC'` also applies **compatibility** mappings, which deliberately discard formatting distinctions: fullwidth forms become ASCII, the fi ligature becomes `fi`, a Roman-numeral character becomes the ASCII letters, `℡` becomes `TEL`. That folding is exactly what you want for matching identifiers, and exactly what creates a bypass if it happens on the wrong side of a check. If code compares a raw string against a reserved-name set, stores it, and only later folds it with `'NFKC'` — or hands it to a component that folds — a fullwidth spelling passes the check and becomes the reserved name downstream. The fix is ordering plus a single owner: fold with `'NFKC'` once at the entry point, run every check on the folded value, and store the folded value so nothing downstream can fold again and change it.
code
python · 11 linesimport unicodedata
RESERVED = {"admin", "root"}
raw = "admin" # fullwidth letters
print(raw in RESERVED) # False - the check passes
folded = unicodedata.normalize("NFKC", raw)
print(folded, folded in RESERVED) # admin True
print(unicodedata.normalize("NFKC", "file")) # file
print(unicodedata.normalize("NFKC", "℡")) # TELgo deeper
Recall that unicodedata.normalize takes a form argument and that the K forms fold extra characters, such as fullwidth letters, into plain ASCII ones. Know that this changes what the string compares equal to.
Explain the canonical-versus-compatibility split, name concrete mappings NFKC performs, and show why folding after a check rather than before it is what creates a bypass rather than the mapping table itself.
Demonstrate ownership of ordering across a real flow: one fold at the entry point, all checks and the length limit on the folded value, the folded value stored, and no second fold downstream in a search index or a second service.
Own the policy across services: whether identifiers are silently folded or non-folded input is rejected outright, which script or character policy sits on top of folding, and how you keep two teams from disagreeing about the canonical form.
## Canonical versus compatibility Unicode defines two kinds of equivalence, and the K in `'NFKC'`/`'NFKD'` is the switch between them. **Canonical equivalence** covers sequences that are the same character written differently: a precomposed e-acute and an `e` followed by a combining acute accent. Applying `'NFC'` never changes what the text *is*; it only picks a spelling. **Compatibility equivalence** covers characters that Unicode encodes separately for formatting or legacy round-trip reasons but which *mean* the same thing: fullwidth `a` and ASCII `a`, the `fi` ligature and the two letters `f` `i`, superscript `²` and digit `2`, circled `①` and `1`, `℡` and the three letters `TEL`, `Ⅸ` and `IX`. `'NFKC'` rewrites all of those into the plain form. That is a *lossy* transformation: you cannot get the fullwidth spelling back, and text you must reproduce faithfully — a person's name, a document body — should not be NFKC-folded on the way into storage. ## Why the security failure is an ordering failure The interesting property is that NFKC can turn a string that is not a reserved word into one that is. `"admin"` (five fullwidth letters) is not equal to `"admin"` under any comparison Python does by default. Fold it with `'NFKC'` and it *is* `"admin"`. So a check and a fold form a pipeline, and only one order is safe: 1. **Fold, then check, then store the folded value.** The check sees the same string every later consumer sees. Safe. 2. **Check, then store the raw value, then fold somewhere downstream.** The check saw one string and the system uses another. This is the bypass, and it does not require the attacker to know where the downstream fold happens — a display layer, a search index, a second service, or Python's own compiler can supply it. The second pattern is easy to arrive at accidentally, because folding often lives in a helper written by someone else. A search index that folds for matching, a login flow that folds before lookup while registration did not, or two services with different opinions about normalization, all reproduce it without any single line of code looking wrong. ## Python folds identifiers this way itself This is not a hypothetical mapping table. CPython applies NFKC to **source identifiers** (PEP 3131), so a name written with fullwidth letters and the ASCII name are the *same name* to the compiler: assigning to a fullwidth `value` rebinds `value`. `str.isidentifier` reports `True` for both spellings. Anyone generating Python source, evaluating names, or matching attribute names against an allowlist has to fold with `'NFKC'` first for the same reason the compiler does. The other place the standard library folds for you is hostnames. Encoding a name with the `'idna'` codec — `"example.com".encode("idna")` — runs a preparation step that includes NFKC, so a fullwidth spelling and the ASCII spelling encode to the *same* wire form and therefore the same host. A hostname allowlist that compares the pre-encoding string against `"example.com"` and then lets a client encode it later is the same ordering bug in a different suit. ## Choosing the form deliberately * `'NFC'` for text you store and reproduce: names, descriptions, message bodies. It preserves meaning. * `'NFKC'` for **identifiers you match on**: usernames, hostnames, lookup keys, anything compared against a reserved or allowed set. Fold and store the folded value. * `'NFKD'`/`'NFD'` mainly when you then want to inspect or strip combining marks yourself. `unicodedata.is_normalized('NFKC', s)` gives the reject-instead-of-rewrite option: an entry point can refuse an identifier that is not already in the folded form, which pushes the decision back to the sender and leaves an unambiguous audit trail instead of a silent rewrite. For registration flows that is often the better policy, because a silent rewrite means the account the user believes they created is not the one that exists. ## What NFKC still does not do Folding is not a confusable filter. Compatibility mappings are a fixed, published table; they do not touch characters that merely *look* alike across scripts. A Cyrillic letter shaped like a Latin `a` survives `'NFKC'` unchanged, so an identifier policy that cares about impersonation needs a script or allowed-character policy on top — typically restricting identifiers to a single script or an explicit character set after folding. And because compatibility folding changes length as well as content, any length limit must be applied *after* the fold, or a value that measured acceptable can grow or shrink into something the storage layer then truncates.
- Where exactly would you put the NFKC call in a registration flow, and what do you store?Fold at the entry point, before any check runs: parse the request, apply `unicodedata.normalize('NFKC', name)`, then run the reserved-name, character-policy and length checks on the folded value, and store that folded value as the account's canonical identifier. Nothing downstream should ever fold again, because a second fold on stored data can change the value the checks approved.
- Why is NFKC the wrong default for text a system must reproduce faithfully?Compatibility folding is lossy by design: it discards distinctions the author chose, rewriting fullwidth letters, ligatures, superscripts and abbreviation characters into plain equivalents that cannot be recovered. For a person's name or a document body that is data corruption. Use `'NFC'` for stored prose and reserve `'NFKC'` for identifiers you compare against a set.
- Does NFKC folding protect against lookalike characters from another script?No. Compatibility mappings are a fixed published table covering formatting variants, not visual confusables. A Cyrillic letter shaped like a Latin one passes through `'NFKC'` untouched. Preventing impersonation needs a separate policy after folding, typically restricting an identifier to a single script or an explicit allowed character set.
- Why must a length limit be applied after folding rather than before?Compatibility mappings change length: one character can fold to several, so a value that measured within the limit before folding can exceed it after. Checking first and folding later means the stored value violates the constraint the check approved, and the storage layer may then truncate it into a different value than either side expected.
A door check reads the badge exactly as printed, but the turnstile behind it silently expands abbreviations before matching. Print the name in an unusual typeface and the guard sees a stranger while the turnstile sees the boss.
saying these in an interview costs you the question
- Thinks NFKC and NFC produce the same result
- Calls NFKC folding lossless and reversible
- Checks the raw string and normalizes afterwards
- Uses NFKC as the storage form for names and prose
- Claims NFKC removes lookalike characters from other scripts
- Applies the length limit before folding