Why can two payee names that render identically fail an equality check after crossing a wire contract as text?
answer
- same glyph, different code points
- precomposed letter versus base plus mark
- comparison sees sequences, not shapes
- normalize once, at the producer
- compatibility forms lose distinctions
basics
~20 sUnicode allows more than one code-point sequence for the same rendered text: a precomposed accented character, or a base letter followed by a combining mark. Both encode faithfully, so equality by bytes or code points fails while the display matches.
solid answer
~40 sThe two values are **canonically equivalent** but not identical. An accented letter can be one precomposed code point, or a base letter plus a combining mark; a renderer draws both the same way, while byte and code-point comparison sees different sequences. Producers on different platforms and input methods legitimately emit different forms, so a field written by two systems arrives in two shapes. The consequences are equality failures, lookups that miss, deduplication that does not deduplicate, and length limits that behave differently per form, since the decomposed form uses more code points and more bytes. The contract must therefore state the character encoding, **which normalization form the field is in**, and that the producer normalizes at the edge before the value enters the contract.
go deeper
Recall that the same written character can have more than one valid code-point sequence, so two names can look identical and still compare unequal.
Explain the mechanism: precomposed character versus base plus combining mark, comparison over sequences rather than glyphs, and what the canonical forms do to each.
Show where it breaks a real system — duplicate payees, uniqueness rules that let both through, length limits that pass in one form and fail in the other — and put normalization at the producer's edge.
Set the policy across the contract: one declared form, producer-side normalization, limits measured after it, and an explicit exemption list for identifiers that must be reproduced unchanged.
## One appearance, several encodings Unicode assigns code points, and for historical compatibility with older character sets it assigns **more than one sequence for the same written character**. An accented letter may appear as a single precomposed code point, or as the unaccented base letter followed by a combining mark that draws the accent over it. Both are correct, both encode faithfully in UTF-8, and a renderer shows the same glyph for each. Comparison does not work at the level of glyphs. It works on bytes, or on code points, so two canonically equivalent names compare unequal. Nothing in the pipeline is broken: the producer is right, the encoder is right, the decoder is right, and the equality check is right about the sequences it was given. ## Where it bites in a contract - **Equality and joins.** Two records for the same payee do not match, so an account is duplicated or a lookup silently returns nothing. - **Deduplication and uniqueness.** A uniqueness rule permits two entries that a human reads as identical. - **Length limits.** The decomposed form uses more code points and more bytes than the precomposed one, so a field of the same visible length passes the limit in one form and fails in the other. - **Search.** A query typed in one form fails to find text stored in the other. - **Cross-boundary diffs.** A value that was never edited appears changed after a round trip through an intermediary that normalized it. ## The normalization forms, and which one to name | Form | What it does | Reversible in effect | Use it for | |---|---|---|---| | Composition (NFC) | Combines base plus mark into precomposed characters where they exist | Yes, canonically equivalent | The default storage and wire form | | Decomposition (NFD) | Splits precomposed characters into base plus marks | Yes, canonically equivalent | Processing that inspects marks | | Compatibility composition (NFKC) | Also folds compatibility variants, such as a ligature into its letters or a full-width form into a plain one | **No** — distinctions are lost | Search keys and matching only | | Compatibility decomposition (NFKD) | The same folding, decomposed | **No** | Matching pipelines | The canonical forms preserve the text's identity: converting between them does not change what the text *is*. The compatibility forms deliberately destroy distinctions in order to make more things match, which is useful for a search index and wrong for a field that must reproduce what the producer wrote. A payee name run through a compatibility form may come back with a ligature expanded, and a downstream system that compares it to the source of record will report a mismatch that no one introduced. ## What the contract should say 1. **The character encoding** of every text scalar, so the bytes have one defined meaning. 2. **The normalization form** the field is in — in practice the composed canonical form — and that it is the producer's job to normalize before the value crosses, not each consumer's job afterwards. Normalizing at every consumer is wasteful and, worse, only works if every consumer remembers. 3. **Where limits are measured**: after normalization, and in which units — code points or bytes — since the two disagree for any text outside the single-byte range. 4. **Which fields are exempt.** Opaque identifiers, keys and anything that must be reproduced byte for byte should not be normalized at all; folding characters in an identifier changes the identifier. 5. **A boundary vector** in the conformance suite: the same name in both canonical forms, which must be accepted and stored identically. ## Two things this is not Normalization is **not case folding** and not trimming. Those are separate transformations with their own hazards, and applying them silently to a name field loses information the producer meant to send. Normalization is also **not a validation step**: it does not reject anything, and a value can be perfectly normalized and still contain characters — invisible marks, direction controls, characters from a different script that look like familiar ones — that a name field should refuse on its own terms.
- Which normalization form is wrong for a field that must reproduce the producer's text exactly, and why?A compatibility form. Unlike the canonical forms, it deliberately folds distinct characters together — a ligature into its component letters, a full-width character into a plain one — so the output is not equivalent to the input, merely more matchable. That is right for a search key and wrong for a name of record, because the stored value no longer equals what the producer sent.
- Who should normalize, and what breaks if the answer is every consumer?The producer, once, at the edge, before the value enters the contract. If it is left to consumers, the stored data holds mixed forms forever, every consumer pays the cost repeatedly, and the one consumer that forgets reintroduces the mismatch for everybody. It also makes any length limit ambiguous, because the value crosses the wire in a form different from the one being measured.
- A field has a maximum length. Does normalization interact with it?Yes. The decomposed form uses more code points, and therefore more bytes, than the composed form for the same visible text, so a name can pass the limit in one form and fail in the other. The contract has to say that limits are measured after normalization and whether the unit is code points or bytes.
saying these in an interview costs you the question
- Assumes identical-looking text always has identical bytes
- Treats normalization as the same thing as lowercasing
- Applies a compatibility form to a name of record
- Leaves normalization to each consumer instead of the producer
- Measures a length limit before normalizing the value
- Normalizes opaque identifiers that must round-trip unchanged