skip to content

In PHP, why can two identical-looking UTF-8 strings such as 'café' fail an === comparison, and how does intl's Normalizer fix it?

level: middleimportance: nice to knowfreq 22%

answer

  1. precomposed versus combining sequence
  2. NFC composes, NFD decomposes
  3. Normalizer::normalize defaults to FORM_C
  4. NFKC folds full-width forms
  5. MB_CASE_FOLD for caseless keys

basics

~10 s

'é' can be one code point, U+00E9, or 'e' plus a combining accent U+0301; both render the same but differ in bytes, so === fails. Normalizer::normalize($s, Normalizer::FORM_C) converts both to the same composed form.

solid answer

~40 s

Unicode allows several code-point sequences for the same text. `é` can be the **precomposed** U+00E9 or `e` followed by the **combining** acute accent U+0301. Text from macOS file names, some keyboards and copy-paste often arrives decomposed, while most web input is composed, so two visually identical product names compare unequal with `===` and miss each other in lookups. The intl extension's `Normalizer::normalize($s, $form)` converts text to a canonical form; the default, `Normalizer::FORM_C` (NFC), composes where possible, while `FORM_D` decomposes. The **compatibility** forms `FORM_KC` and `FORM_KD` also fold variants such as full-width Latin letters and half-width katakana, which suits search keys but loses distinctions, so do not store them for display. Normalize at the input boundary, store NFC, and use `mb_convert_case($s, MB_CASE_FOLD)` when keys must also ignore case.

code

php · 17 lines
php
<?php
declare(strict_types=1);

$composed = "caf\u{00E9}";        // é as one code point
$decomposed = "cafe\u{0301}";     // e + combining acute accent

var_dump($composed === $decomposed);                  // bool(false)

$nfc = Normalizer::normalize($decomposed, Normalizer::FORM_C);
var_dump($nfc === $composed);                         // bool(true)
var_dump(Normalizer::isNormalized($decomposed));      // bool(false)

$key = Normalizer::normalize('ABC123', Normalizer::FORM_KC);
var_dump($key);                                       // string(6) "ABC123"

var_dump(mb_convert_case('Straße', MB_CASE_FOLD, 'UTF-8')
    === mb_convert_case('STRASSE', MB_CASE_FOLD, 'UTF-8')); // bool(true)

go deeper

for a junior

Recall that the same visible text can be stored as different code points, so equal-looking strings may fail ===.

for a middle

Explain precomposed versus combining sequences, the NFC, NFD, NFKC and NFKD forms, and the Normalizer::normalize signature and failure value.

for a senior

Show you would normalize to NFC at input, build NFKC plus case-folded search keys, and check isNormalized cheaply on hot paths.

for a principal

Set the organisation-wide rule for text identity: which form is stored, which form keys use, and which boundary guarantees it.

## One text, several encodings Unicode often offers more than one sequence of code points for the same visible text: - **Precomposed**: `é` as the single code point U+00E9 LATIN SMALL LETTER E WITH ACUTE. - **Decomposed**: `e` (U+0065) followed by U+0301 COMBINING ACUTE ACCENT. Both render identically. In UTF-8 the first is two bytes, the second three, and PHP's `===` compares bytes, so `'café' === "cafe\u{0301}"` is `false`. `strlen()` gives 5 and 6, and `mb_strlen()` gives 4 and 5. This is not an exotic case. Decomposed text commonly arrives from file names created on some operating systems, from certain input methods and from copy-paste between applications. In a marketplace it shows up as duplicate sellers, product searches that miss exact matches and unique indexes that accept two "identical" names. ## The four normalization forms Unicode defines **normalization forms** that map equivalent sequences to one canonical representation. PHP exposes them through the intl extension's `Normalizer` class, with long and short constant names: | Constant | Short name | Effect | |---|---|---| | `Normalizer::FORM_C` | `Normalizer::NFC` | canonical decomposition, then composition; the usual storage form | | `Normalizer::FORM_D` | `Normalizer::NFD` | canonical decomposition | | `Normalizer::FORM_KC` | `Normalizer::NFKC` | compatibility decomposition, then composition | | `Normalizer::FORM_KD` | `Normalizer::NFKD` | compatibility decomposition | | `Normalizer::FORM_KC_CF` | `Normalizer::NFKC_CF` | NFKC plus case folding | **Canonical** equivalence covers sequences that mean exactly the same thing. **Compatibility** equivalence is broader: it treats full-width `ABC123` as `ABC123`, half-width katakana `カ` as `カ`, and the ligature `fi` as `fi`. Those mappings lose information, so compatibility forms are for comparison keys, not for text you will display again. ## The API - `Normalizer::normalize(string $string, int $form = Normalizer::FORM_C): string|false` returns the normalized string, or `false` on failure, for example on invalid UTF-8. - `Normalizer::isNormalized(string $string, int $form = Normalizer::FORM_C): bool` checks without converting, a cheap guard on hot paths. - `normalizer_normalize()` and `normalizer_is_normalized()` are the procedural equivalents. Always check the return value: a `false` passed on as a string becomes an empty name. ## Where to normalize 1. **At the input boundary.** Normalize seller-supplied names, search queries and usernames to NFC as they enter the system, so stored data is consistent. 2. **For comparison keys.** Build a separate search key with NFKC, then case-fold it, and index that key. Keep the NFC original for display. 3. **Before hashing or signing.** Two equivalent strings hash differently unless normalized first. ## Case-insensitive keys Normalization does not remove case differences. For caseless comparison, fold case after normalizing: - `mb_strtolower()` lowercases, but lowercase forms can still differ: `'Straße'` stays `'straße'` while `'STRASSE'` becomes `'strasse'`. - `mb_convert_case($s, MB_CASE_FOLD, 'UTF-8')` applies full Unicode **case folding**, mapping `ß` to `ss`, which is designed for comparisons. - `Normalizer::FORM_KC_CF` combines compatibility normalization and case folding in one call. A robust search key for a Japanese and European catalogue is therefore NFKC plus case folding, applied identically to stored names and to queries. ## Pitfalls - **Normalizing twice with different forms.** Mixing NFC in one service and NFD in another recreates the original mismatch; pick one storage form. - **Normalizing only one side.** A search key built with NFKC must be applied to both stored names and incoming queries, or nothing matches. - **Ignoring the failure value.** `Normalizer::normalize()` returns `false` for input it cannot process; writing that into a record produces an empty name. - **Assuming the database collation does it.** Whether a database compares composed and decomposed forms as equal depends on its collation; normalizing in PHP makes the behaviour explicit. - **Missing extension.** `Normalizer` belongs to intl, which may be absent on a minimal server image. ## What an interviewer is checking The question tests whether you know that equal-looking strings can differ in bytes, can name NFC and NFD, and know that NFKC is lossy. Mentioning that normalization is not a security filter on its own, and that it belongs at a defined boundary rather than scattered through the code, shows production experience.

  • Why should you not store product names in NFKC?
    Compatibility normalization is lossy: it turns full-width letters into ASCII, half-width katakana into full-width, and ligatures into separate letters. That is right for a search key but changes what sellers typed. Store NFC for display and keep an NFKC, case-folded copy only as an index key.
  • Why is mb_strtolower() not enough for a case-insensitive comparison key?
    Lowercasing is not designed for equivalence: `'Straße'` lowercases to `'straße'` while `'STRASSE'` lowercases to `'strasse'`. `mb_convert_case($s, MB_CASE_FOLD)` applies Unicode case folding, mapping `ß` to `ss`, so both produce the same key.

saying these in an interview costs you the question

  • Two strings that render identically are always equal under ===.
  • Normalizer::normalize uses NFD unless told otherwise.
  • NFKC is lossless, so it is the best form to store for display.
  • mb_strtolower is equivalent to Unicode case folding.
  • Normalization is only needed for Asian scripts.