Two strings that both display as "café" compare false with === in JavaScript, and one reports length 4 while the other reports 5. What is going on, and how do you compare them correctly?
answer
- identical on screen, different underneath
- the length difference is the clue
- one base letter, one combining mark
- canonical form before comparing
- compatibility folding is lossy
basics
~20 sUnicode allows two encodings of the same accented text: a precomposed é (U+00E9) or an e followed by a combining acute accent (U+0301). JavaScript's === compares code units, so the two differ. Call normalize('NFC') on both before comparing or storing.
solid answer
~50 sUnicode gives some characters more than one legal encoding. `é` can be the single precomposed code point U+00E9, or the two-code-point sequence `e` + combining acute U+0301. Both render identically, but `===` on strings is an exact code-unit comparison, so they are unequal — and one string is a code unit longer, which is your clue. The fix is `String.prototype.normalize` (ES2015): normalise both sides to the same form, conventionally NFC, and compare then. In practice you normalise once at the boundary — on input, before storing, and before hashing or using the value as a key — rather than at every comparison site, so the canonical form is what lives in your system. NFD is the decomposed counterpart; NFKC and NFKD additionally fold compatibility differences and are lossy, so keep them for search keys rather than for text you must round-trip.
code
javascript · 11 linesconst precomposed = 'café'; // é as a single code point
const decomposed = 'café'; // e + combining acute accent
console.log(precomposed === decomposed); // false
console.log(precomposed.length, decomposed.length); // 4 5
console.log(
precomposed.normalize('NFC') === decomposed.normalize('NFC')
); // true
console.log('fi'.normalize('NFKC')); // "fi" — lossy compatibility folding
console.log('①'.normalize('NFKC')); // "1"go deeper
Know that the same visible character can have two encodings, that === compares raw code units, and that calling normalize() on both strings before comparing is the fix.
Explain canonical equivalence concretely: precomposed U+00E9 versus e plus combining acute U+0301, why the lengths differ, and what each of NFC, NFD, NFKC and NFKD does.
Show the diagnosis path from the symptom — identical-looking values that fail lookup or de-duplication, with different lengths — to normalising once at the ingest boundary so storage, hashing and map keys all hold the canonical form.
Own the policy: which normalization form is canonical for the system, that compatibility forms are lossy and belong to derived search keys rather than stored text, and how normalization, case folding and collation are three separate decisions that must not be conflated.
## The same text, two legal encodings Unicode is not injective from appearance to encoding. Many accented characters exist both as a single *precomposed* code point and as a *base character plus one or more combining marks*: ```javascript const precomposed = 'café'; // é as U+00E9 const decomposed = 'café'; // e + U+0301 COMBINING ACUTE ACCENT console.log(precomposed); // "café" console.log(decomposed); // "café" — visually identical console.log(precomposed === decomposed); // false console.log(precomposed.length, decomposed.length); // 4 5 ``` JavaScript's `===` on two strings is a code-unit-by-code-unit comparison with no Unicode awareness at all. Same appearance, different bytes, unequal. The length discrepancy is the diagnostic that usually cracks the case: identical-looking text with different lengths means different encodings, not a rendering bug. ## Where each form comes from You rarely choose the form; your input does. Different keyboards, input methods, operating systems, filesystems and clipboards produce different forms, and text that has travelled through several of them can be mixed. That is why the symptom is so often "it works when I type it and fails when I paste it" — two code paths, two encodings, one comparison. ## normalize and the four forms `String.prototype.normalize(form)` (ES2015) implements the Unicode normalization algorithms. The argument defaults to `'NFC'` and accepts four values: - **NFC** — canonical decomposition, then canonical composition. The composed form, and the usual default for storage and interchange. - **NFD** — canonical decomposition only. Useful when you want to strip marks, since removing the combining range from a decomposed string leaves the base letters. - **NFKC / NFKD** — the same, plus *compatibility* folding. Canonical forms preserve meaning: NFC and NFD of the same text always render the same and can be converted back and forth. Compatibility forms deliberately discard distinctions: ```javascript console.log('fi'.normalize('NFKC')); // "fi" — the fi ligature is folded apart console.log('①'.normalize('NFKC')); // "1" — circled digit one becomes a plain 1 console.log('A'.normalize('NFKC')); // "A" — fullwidth A becomes ASCII A ``` That folding is lossy and irreversible. It is excellent for a search index or a duplicate-detection key, and wrong for text you must give back to the user unchanged. ```javascript console.log(precomposed.normalize('NFC') === decomposed.normalize('NFC')); // true ``` ## Normalise at the boundary, not at every comparison The robust design is to normalise once, on the way in: validate and `normalize('NFC')` user input, store the canonical form, and let everything downstream compare with plain `===`. Normalising at each comparison site is easy to forget in one place, and that one place is where the bug lives. It matters most for: - equality checks and de-duplication; - keys in a `Map` or `Set`, and any hashing, since two forms hash differently; - anything compared across systems that may have normalised differently. A second, non-obvious consequence: normalization can change length. Never normalise *after* enforcing a length limit — measure the canonical form. ## Normalization is not case folding, and not collation These are three different problems and candidates routinely blur them: - **Normalization** removes encoding differences for the *same* characters. - **Case folding** (`toLowerCase`, `toUpperCase`) removes case differences, and is locale-sensitive for some scripts. - **Collation** decides ordering and "equal enough" for a human, which is what `localeCompare` is for — `a.localeCompare(b, 'en', { sensitivity: 'base' })` returns 0 for text differing only in accents and case. A case-insensitive, accent-insensitive match needs normalization *plus* a folding decision; `normalize` alone answers only the first question. ## Stripping accents, since it always comes up NFD plus a regex over the combining range is the standard idiom, and it is worth knowing why it works: decomposition splits `é` into `e` and a mark, and the marks live in a contiguous block. ```javascript const deaccent = (s) => s.normalize('NFD').replace(/[̀-ͯ]/g, ''); console.log(deaccent('café')); // "cafe" ``` This is a pragmatic slug/search helper, not a general transliteration — it does nothing for scripts whose distinctions are not expressed as combining marks. ## The interview-grade answer Name the mechanism (canonical equivalence), show the length clue, reach for `normalize('NFC')`, and then add the judgment: normalise at the boundary, keep compatibility forms for keys rather than for stored text, and do not confuse normalization with case folding or collation.
- When would you choose NFKC over NFC?For derived keys rather than stored text: search indexes, slugs, duplicate detection. NFKC additionally folds compatibility characters, turning the fi ligature into `fi`, a circled digit into a plain digit and fullwidth Latin into ASCII, which is exactly what you want for matching. It is lossy and irreversible, so never store it as the user's own text.
- Where in an application would you actually call normalize?At the ingest boundary — once, on input, before validation, storage, hashing or use as a key — so the canonical form is what lives in the system and downstream code can use plain `===`. Normalising at each comparison site is fragile, because the one site you forget is where the bug appears.
- Does normalize handle case-insensitive comparison too?No. Normalization removes encoding differences between the same characters; it does nothing about case. Case folding via `toLowerCase` is a separate step, and for a human-facing accent- and case-insensitive match you would instead use `localeCompare` with a sensitivity option, which answers the collation question rather than the encoding one.
- Why can normalization change a string's length, and when does that bite?NFC composes a base letter plus a mark into one code point, and NFD does the reverse, so the code-unit count moves. It bites when a length limit is enforced before normalising: text that passed the check on the way in can exceed it once stored, or vice versa. Always measure the canonical form.
saying these in an interview costs you the question
- Blames the font or the terminal for the mismatch
- Thinks === does any Unicode-aware comparison
- Uses NFKC for text that must be returned unchanged
- Confuses normalization with lowercasing
- Normalises at each comparison instead of at input