How do you decide whether to depend on golang.org/x/text and normalise stored names to NFC?
answer
- the standard library stops at encoding
- same glyph, two byte sequences
- it is also a go.mod decision
- cheap before data, a migration after
- identity or decoration decides it
basics
~20 sGo's standard library validates UTF-8 but does not normalise, so normalisation means adding golang.org/x/text to go.mod. Decide it before user data exists, because normalising later rewrites stored bytes and can collide rows that were previously distinct.
solid answer
~40 sThe forcing fact: `unicode/utf8` and `unicode` can say text is well formed and what category each rune is in, but nothing in the standard library composes "e" plus a combining acute into é. Normalisation forms come from `golang.org/x/text/unicode/norm`, so "should we normalise" is inseparably "should we take this dependency". I decide it on data-model grounds: if names are compared, deduplicated, uniquely indexed or used as lookup keys, two byte sequences that render identically must not be two values, and NFC at ingress is the fix; if names are only displayed, the case is weak. Timing matters more than the choice — normalising after data exists rewrites stored bytes and can collide rows on a unique index. And the module is real: pinned, upgraded and scanned like any other.
go deeper
Know that two strings can look identical on screen yet hold different bytes, and that Go's standard library will call both of them valid.
Explain composed versus decomposed encodings with a concrete example, and name where normalisation forms come from, since no standard library package provides them.
Show where the rule belongs: one ingress or repository boundary every write path crosses, plus a test fixture carrying both encodings so CI can actually detect a regression.
Own the call end to end: whether a name is an identity or a label, the irreversibility once data exists, who signs off on the added module and its pinning, and what you would insist on even if the answer is no.
## Why this is a decision and not a lookup Go's standard library deliberately stops at encoding. `utf8.ValidString` tells you the bytes are well-formed UTF-8; `unicode.IsLetter` tells you what a rune is; `strings.ToValidUTF8` repairs damage. None of them normalise. The Unicode normalisation forms — NFC, NFD, NFKC, NFKD — live in `golang.org/x/text/unicode/norm`, which is a separate module. So the engineering question "do we normalise names" is simultaneously a `go.mod` question, and it lands on whoever owns dependency policy. ## The problem normalisation solves Unicode lets the same visible text be encoded more than one way. The letter é can be: - the single code point U+00E9 (composed, 2 bytes), or - "e" U+0065 followed by U+0301 COMBINING ACUTE ACCENT (decomposed, 3 bytes). Both are valid UTF-8. Both render identically. Neither `==` nor a database unique index nor a hash map will treat them as equal. Which one you get depends on the user's platform: some input methods produce composed text, and some file systems and clients hand you decomposed text. The consequences show up wherever text is used as an identity rather than as decoration: - A user registers with one form, signs in with the other, and their account is not found. - Two accounts exist with names that look identical to every human who sees them. - A cache keyed by name misses constantly for a subset of users, and only for them. - A search index does not match text a user copied straight from your own page. **NFC** composes where possible and is the form the web platform recommends for interchange; it is the default choice. NFD decomposes; NFKC and NFKD additionally apply compatibility mappings, which fold things like the ligature fi and superscript digits together — useful for search matching, destructive if applied to data you must give back to the user unchanged. ## How I would frame the decision **Start from the data model, not the dependency.** The question is whether a name is an identity or a label. - If the value is compared, deduplicated, uniquely indexed, used as a lookup key, or shown as proof that two things are the same, encoding variation is a correctness bug and normalisation is the fix. - If the value is only rendered, the failure mode is cosmetic and the dependency may not be worth it. **Decide before data exists.** This is the part that is genuinely hard to reverse. Normalising a live column rewrites stored bytes; rows that were distinct can become identical, which a unique index will reject halfway through the migration, and any external system holding the old bytes now disagrees with you. Doing it on an empty table is a one-line policy; doing it on a populated one is a project with a dedup decision attached, and somebody has to choose which of two colliding accounts survives. **Separate the storage form from the comparison form.** A common settlement is to store what the user gave you, normalised only to NFC (which preserves the text), and to compute a separate, more aggressive folded key for uniqueness and search. That keeps the user's own bytes intact while still catching duplicates. It costs a column and it means two rules to explain, and that tradeoff is exactly what the discussion should be about. **Price the dependency honestly.** `golang.org/x/text` is a Go-team-maintained module and about as safe a dependency as exists outside the standard library, but it is still a module in `go.mod`: it gets pinned, it gets upgraded, it appears in the vulnerability report and the licence audit, and its Unicode data updates can change behaviour between versions. Because Go's minimal version selection picks the lowest version satisfying all requirements rather than the newest, the version you actually build is a property of the whole module graph, not of your `require` line alone — so "we pinned it" needs to mean the graph agrees. **Push it to a boundary you control.** Whatever the rule is, it must live in one place that every path crosses — ingress validation, or the repository layer — not in each handler. A normalisation applied in two of three write paths is worse than none, because the inconsistency is now invisible. ## What would make me say no - Names are display-only, never keys, and the product has no dedup requirement. - The organisation has a hard rule against non-standard-library dependencies in a particular binary, and the cost of arguing it exceeds the benefit. - An upstream system already normalises and owns the identity; doing it again downstream just creates two authorities that can drift. ## What I would insist on either way - Validate UTF-8 at ingress with `utf8.ValidString`, whatever we decide about normalisation. Well-formedness is not optional and needs no dependency. - Write the decision down, including what a name means and what makes two names the same, so the next team does not answer it differently. - Test with a fixture containing both encodings of the same name; without it, nothing in CI can tell the two apart. ## What an interviewer is checking Whether you can hold three things at once: a Unicode fact, a dependency policy, and an irreversible data migration. A strong answer starts with what the data is for, treats the dependency as a cost to be justified rather than an obstacle or a freebie, and is explicit that the window to make this choice cheaply closes the day real users sign up.
- Why is it so much cheaper to make this decision before user data exists?Because normalising later is a migration that rewrites stored bytes. Rows that were distinct can become identical, so a unique index rejects the update partway and somebody has to decide which account wins. Any external system, cache or export holding the old bytes now disagrees with you. On an empty table the same decision is one line of validation code.
- What would you normalise to, and would you ever store a second form?NFC for the stored value: it is the interchange form the web recommends and it preserves the text the user typed. Where uniqueness or search matters I would additionally compute a separate folded key, using a compatibility form, and index that. Storing the user's own text and the matching key separately avoids handing people back a name they did not write.
- How do you argue the dependency itself to a team that resists adding modules?By pricing both sides. The module is maintained by the Go team, is widely used and adds no transitive weight of consequence, but it still has to be pinned, upgraded and scanned. Against that, the alternative is either shipping the correctness bug or hand-rolling normalisation, which means carrying Unicode data tables ourselves. Framed that way the dependency is usually the cheaper of the three.
- Does normalisation solve look-alike impersonation between scripts?No, and conflating the two is a common mistake. Normalisation reconciles different encodings of the same character; it does nothing about a Cyrillic letter that merely looks like a Latin one, because those are genuinely different characters. Mixed-script confusables are a separate policy needing script restriction or a confusable-skeleton comparison, decided on its own terms.
It is like choosing a spelling convention for a directory: trivial to declare before the first entry, and a painful reconciliation once thousands of entries exist in both spellings.
saying these in an interview costs you the question
- Assumes the standard library can normalise text
- Treats normalisation as reversible once data exists
- Confuses normalisation with confusable-script defence
- Applies compatibility folding to text shown back to users
- Normalises in some write paths but not all