In PHP, why can two identical-looking UTF-8 strings such as 'café' fail an === comparison, and how does intl's Normalizer fix it?
answer
- precomposed versus combining sequence
- NFC composes, NFD decomposes
- Normalizer::normalize defaults to FORM_C
- NFKC folds full-width forms
- MB_CASE_FOLD for caseless keys
basics
~10 s'é' can be one code point, U+00E9, or 'e' plus a combining accent U+0301; both render the same but differ in bytes, so === fails. Normalizer::normalize($s, Normalizer::FORM_C) converts both to the same composed form.
solid answer
~40 sUnicode allows several code-point sequences for the same text. `é` can be the **precomposed** U+00E9 or `e` followed by the **combining** acute accent U+0301. Text from macOS file names, some keyboards and copy-paste often arrives decomposed, while most web input is composed, so two visually identical product names compare unequal with `===` and miss each other in lookups. The intl extension's `Normalizer::normalize($s, $form)` converts text to a canonical form; the default, `Normalizer::FORM_C` (NFC), composes where possible, while `FORM_D` decomposes. The **compatibility** forms `FORM_KC` and `FORM_KD` also fold variants such as full-width Latin letters and half-width katakana, which suits search keys but loses distinctions, so do not store them for display. Normalize at the input boundary, store NFC, and use `mb_convert_case($s, MB_CASE_FOLD)` when keys must also ignore case.
code
php · 17 lines<?php
declare(strict_types=1);
$composed = "caf\u{00E9}"; // é as one code point
$decomposed = "cafe\u{0301}"; // e + combining acute accent
var_dump($composed === $decomposed); // bool(false)
$nfc = Normalizer::normalize($decomposed, Normalizer::FORM_C);
var_dump($nfc === $composed); // bool(true)
var_dump(Normalizer::isNormalized($decomposed)); // bool(false)
$key = Normalizer::normalize('ABC123', Normalizer::FORM_KC);
var_dump($key); // string(6) "ABC123"
var_dump(mb_convert_case('Straße', MB_CASE_FOLD, 'UTF-8')
=== mb_convert_case('STRASSE', MB_CASE_FOLD, 'UTF-8')); // bool(true)go deeper
Recall that the same visible text can be stored as different code points, so equal-looking strings may fail ===.
Explain precomposed versus combining sequences, the NFC, NFD, NFKC and NFKD forms, and the Normalizer::normalize signature and failure value.
Show you would normalize to NFC at input, build NFKC plus case-folded search keys, and check isNormalized cheaply on hot paths.
Set the organisation-wide rule for text identity: which form is stored, which form keys use, and which boundary guarantees it.
## One text, several encodings Unicode often offers more than one sequence of code points for the same visible text: - **Precomposed**: `é` as the single code point U+00E9 LATIN SMALL LETTER E WITH ACUTE. - **Decomposed**: `e` (U+0065) followed by U+0301 COMBINING ACUTE ACCENT. Both render identically. In UTF-8 the first is two bytes, the second three, and PHP's `===` compares bytes, so `'café' === "cafe\u{0301}"` is `false`. `strlen()` gives 5 and 6, and `mb_strlen()` gives 4 and 5. This is not an exotic case. Decomposed text commonly arrives from file names created on some operating systems, from certain input methods and from copy-paste between applications. In a marketplace it shows up as duplicate sellers, product searches that miss exact matches and unique indexes that accept two "identical" names. ## The four normalization forms Unicode defines **normalization forms** that map equivalent sequences to one canonical representation. PHP exposes them through the intl extension's `Normalizer` class, with long and short constant names: | Constant | Short name | Effect | |---|---|---| | `Normalizer::FORM_C` | `Normalizer::NFC` | canonical decomposition, then composition; the usual storage form | | `Normalizer::FORM_D` | `Normalizer::NFD` | canonical decomposition | | `Normalizer::FORM_KC` | `Normalizer::NFKC` | compatibility decomposition, then composition | | `Normalizer::FORM_KD` | `Normalizer::NFKD` | compatibility decomposition | | `Normalizer::FORM_KC_CF` | `Normalizer::NFKC_CF` | NFKC plus case folding | **Canonical** equivalence covers sequences that mean exactly the same thing. **Compatibility** equivalence is broader: it treats full-width `ABC123` as `ABC123`, half-width katakana `カ` as `カ`, and the ligature `fi` as `fi`. Those mappings lose information, so compatibility forms are for comparison keys, not for text you will display again. ## The API - `Normalizer::normalize(string $string, int $form = Normalizer::FORM_C): string|false` returns the normalized string, or `false` on failure, for example on invalid UTF-8. - `Normalizer::isNormalized(string $string, int $form = Normalizer::FORM_C): bool` checks without converting, a cheap guard on hot paths. - `normalizer_normalize()` and `normalizer_is_normalized()` are the procedural equivalents. Always check the return value: a `false` passed on as a string becomes an empty name. ## Where to normalize 1. **At the input boundary.** Normalize seller-supplied names, search queries and usernames to NFC as they enter the system, so stored data is consistent. 2. **For comparison keys.** Build a separate search key with NFKC, then case-fold it, and index that key. Keep the NFC original for display. 3. **Before hashing or signing.** Two equivalent strings hash differently unless normalized first. ## Case-insensitive keys Normalization does not remove case differences. For caseless comparison, fold case after normalizing: - `mb_strtolower()` lowercases, but lowercase forms can still differ: `'Straße'` stays `'straße'` while `'STRASSE'` becomes `'strasse'`. - `mb_convert_case($s, MB_CASE_FOLD, 'UTF-8')` applies full Unicode **case folding**, mapping `ß` to `ss`, which is designed for comparisons. - `Normalizer::FORM_KC_CF` combines compatibility normalization and case folding in one call. A robust search key for a Japanese and European catalogue is therefore NFKC plus case folding, applied identically to stored names and to queries. ## Pitfalls - **Normalizing twice with different forms.** Mixing NFC in one service and NFD in another recreates the original mismatch; pick one storage form. - **Normalizing only one side.** A search key built with NFKC must be applied to both stored names and incoming queries, or nothing matches. - **Ignoring the failure value.** `Normalizer::normalize()` returns `false` for input it cannot process; writing that into a record produces an empty name. - **Assuming the database collation does it.** Whether a database compares composed and decomposed forms as equal depends on its collation; normalizing in PHP makes the behaviour explicit. - **Missing extension.** `Normalizer` belongs to intl, which may be absent on a minimal server image. ## What an interviewer is checking The question tests whether you know that equal-looking strings can differ in bytes, can name NFC and NFD, and know that NFKC is lossy. Mentioning that normalization is not a security filter on its own, and that it belongs at a defined boundary rather than scattered through the code, shows production experience.
- Why should you not store product names in NFKC?Compatibility normalization is lossy: it turns full-width letters into ASCII, half-width katakana into full-width, and ligatures into separate letters. That is right for a search key but changes what sellers typed. Store NFC for display and keep an NFKC, case-folded copy only as an index key.
- Why is mb_strtolower() not enough for a case-insensitive comparison key?Lowercasing is not designed for equivalence: `'Straße'` lowercases to `'straße'` while `'STRASSE'` lowercases to `'strasse'`. `mb_convert_case($s, MB_CASE_FOLD)` applies Unicode case folding, mapping `ß` to `ss`, so both produce the same key.
saying these in an interview costs you the question
- Two strings that render identically are always equal under ===.
- Normalizer::normalize uses NFD unless told otherwise.
- NFKC is lossless, so it is the best form to store for display.
- mb_strtolower is equivalent to Unicode case folding.
- Normalization is only needed for Asian scripts.