skip to content

questions

4

How do you strip accents from a Python str to build a search or slug key?

level: middleimportance: must knowfreq 52%

answer

  1. You cannot delete what is not separate
  2. One precomposed code point hides the mark
  3. Decompose first, then filter
  4. unicodedata.combining is non-zero on marks
  5. Recompose with NFC before storing

basics

~20 s

Normalize to NFD first so each accent becomes its own combining code point, drop every character whose unicodedata.combining value is non-zero, then normalize back to NFC. Without the NFD step the accent is fused inside one precomposed code point.

solid answer

~40 s

Call `unicodedata.normalize("NFD", s)` to split each precomposed letter into a base character plus its combining marks, join back only the characters where `unicodedata.combining(ch)` returns 0 (equivalently, where `unicodedata.category(ch)` is not `"Mn"`), then `unicodedata.normalize("NFC", ...)` the result. `str.replace` per accented letter only covers the letters you enumerated, and `s.encode("ascii", "ignore")` does not strip the accent — it deletes the whole letter, turning "Renée" into `b'Rene'` and wiping out every non-Latin character in the string. Note that NFD only undoes canonical decompositions: `ø`, `ł` and `ß` have none, so they survive and need an explicit, language-dependent mapping. Keep the folded value as a lookup key next to the original text, not as a replacement for it.

code

python · 14 lines
python
import unicodedata


def fold(s: str) -> str:
    decomposed = unicodedata.normalize("NFD", s)
    without_marks = "".join(
        ch for ch in decomposed if not unicodedata.combining(ch)
    )
    return unicodedata.normalize("NFC", without_marks)


name = "Ren\u00e9e Fabr\u00e8ge"
print(fold(name))                        # Renee Fabrege
print(name.encode("ascii", "ignore"))    # b'Rene Fabrge'

go deeper

for a junior

Recall that Python can spell an accented letter as one code point or as a base letter plus a separate accent, and that unicodedata.normalize moves between the two. Know that encode("ascii", "ignore") deletes characters rather than simplifying them.

for a middle

Be ready to write the three-line function at the keyboard and justify each step: NFD to expose the marks, a unicodedata.combining filter to drop them, NFC to put the result in a stable form. Explain why filtering a precomposed string is a no-op.

for a senior

Show that you know where the technique stops: letters with no canonical decomposition, the language-dependence of any ASCII mapping, and folding once on write instead of inside a query. Demonstrate keeping the folded key beside the original rather than replacing it.

for a principal

Own the policy question. Decide whether the product needs a search key, a slug or a uniqueness constraint - they tolerate different amounts of collapsing - and note that a many-to-one fold applied to identity or authorization values is how two distinct users end up sharing one record.

### The problem You want a *folding key*: a version of a Python `str` you can compare, index or put in a URL slug, in which "Renee" finds "Renée" and "Fabrege" finds "Fabrège". The two approaches people reach for first both fail, and they fail for the same underlying reason. The first is a hand-written table: `s.replace("é", "e").replace("è", "e")...`, or the same thing via `str.maketrans` and `str.translate`. It works only for the accented letters you thought of, and Latin script alone has hundreds of them across a dozen diacritics. The second is `s.encode("ascii", "ignore").decode()`. On `"Renée Fabrège"` that returns `b'Rene Fabrge'` — the accented letters are not *stripped*, they are **deleted**, because the codec sees one un-encodable code point and the `ignore` error handler drops it whole. The same call silently deletes every CJK character, every emoji and every Cyrillic letter in the string too, so a non-Latin title collapses to an empty key. ### Two representations of the same letter Unicode can spell "é" in two ways. **Precomposed** (NFC) is one code point, U+00E9 LATIN SMALL LETTER E WITH ACUTE. **Decomposed** (NFD) is two: `e` followed by U+0301 COMBINING ACUTE ACCENT. `unicodedata.normalize` converts between them. Text arriving from a form, a database or an HTTP request is usually precomposed, and that is exactly why filtering does nothing: there is no separate accent character to remove — the accent is *inside* the single code point. ### The recipe ```python import unicodedata def fold(s: str) -> str: d = unicodedata.normalize("NFD", s) return unicodedata.normalize("NFC", "".join( ch for ch in d if not unicodedata.combining(ch) )) ``` Three steps, each load-bearing: 1. **`normalize("NFD", s)`** applies canonical decomposition, splitting every precomposed letter into a base character plus its combining marks. `len("café")` is 4; `len(normalize("NFD", "café"))` is 5. 2. **Filter the marks.** `unicodedata.combining(ch)` returns the canonical combining class as an integer — non-zero exactly for characters that stack onto a preceding base character, zero for everything else. Testing `unicodedata.category(ch) == "Mn"` (nonspacing mark) is the usual equivalent; `combining` is the more precise test for "this reorders/stacks during normalization", and `Mn` is the more precise test for "this is a nonspacing mark". For accent stripping either is fine, and they agree on Latin diacritics. 3. **`normalize("NFC", ...)`** recomposes what is left. After removing the marks there is usually nothing to recompose, but it matters when the string mixes scripts: NFC is the form you want to store and compare, and returning a half-decomposed string means the next `==` against a precomposed value fails for the very reason you started here. ### What it does not do `NFD` only undoes *canonical* decompositions, and plenty of letters have none: `unicodedata.decomposition("ø")` is the empty string, and so is `đ`, `ł`, `ß`, `þ` and `æ`. NFD leaves them as they are, so the folded key still contains non-ASCII. That is a feature — it stops you silently mangling text — but it means "strip accents" is not the same as "make ASCII". If you truly need ASCII you must add an explicit mapping (`str.maketrans` over the letters you care about), and you are then making a **language-dependent** choice: German expects `ö → oe`, Swedish treats `ö` as a distinct letter that sorts after `z`, and Turkish distinguishes dotted and dotless `i`. A general transliteration library exists for this, but it is a transliteration policy, not a Unicode operation. `NFKD` decomposes more — compatibility mappings turn `fi` into `fi`, `²` into `2`, and full-width forms into ASCII — which is often desirable in a search key and disastrous in stored text, since it is not round-trippable. It still does nothing for `ø`. ### Using it correctly The folded value is a **key beside the original, never a replacement for it**. Store and display the user's real text; compute the key for lookup, uniqueness or a URL slug, and index that column. For a case-insensitive key, combine with `str.casefold()` (not `str.lower()`), which handles `ß → ss` and Greek final sigma. A defensible order is: `NFD` → drop marks → `casefold()` → `NFC`, because casefolding some characters can itself change the normalization form, so normalizing last leaves the key in a stable shape. Cost is a full pass over the string plus a database lookup per character in the worst case, so fold once when the record is written rather than on every comparison in a query loop. Note also that a folded key is **lossy and many-to-one**: "Renee", "Renée" and "Renèe" collapse together, which is what you want for search and what you must not do for authentication or authorization identifiers, where collapsing distinct inputs onto one value is how impersonation happens.

  • Your folded key still contains ø and ß. What do you do about them?
    Nothing in `unicodedata` will fix those: `unicodedata.decomposition("\u00f8")` is empty, so NFD leaves the letter whole and there is no mark to drop. Reaching ASCII means an explicit mapping, typically `str.maketrans` plus `str.translate`, and that is a language policy rather than a Unicode operation - German expects ö to become oe, while Swedish treats it as a distinct letter. Either accept a non-ASCII key, or pick a transliteration table deliberately and document the locale it assumes.
  • Would you store the folded value or compute it on every lookup?
    Store it alongside the original in its own indexed column, computed once on write. Folding is a full character-by-character pass with a database lookup per character, so doing it inside a query predicate defeats the index and scans the table. Keep the user's real text for display - the folded form is lossy and many-to-one, so it can never be the source of truth.
  • What changes if you use NFKD instead of NFD for the decomposition step?
    NFKD adds compatibility mappings on top of canonical ones: the fi ligature becomes fi, superscript two becomes 2, full-width Latin becomes ASCII. That is often what you want in a search key, since it collapses visually equivalent spellings. It is wrong for stored text because it is not round-trippable - you cannot recover the original form. It still does nothing for letters that have no decomposition at all.

A precomposed letter is like a stamp that prints the vowel and its accent as one piece of metal — you cannot rub the accent off. NFD swaps in two stamps, base and accent, and then you can simply not press the second one.

saying these in an interview costs you the question

  • Hand-writes a str.replace chain per accented letter
  • Claims s.encode('ascii', 'ignore') strips accents rather than deleting letters
  • Filters combining marks without normalizing to NFD first
  • Assumes every accented letter decomposes, so ø and ł fold to ASCII
  • Overwrites the stored display text with the lossy folded key
  • Uses str.lower instead of str.casefold when case-folding the key

context

open as a page

What does Python's str.casefold() do that str.lower() does not?

level: juniorimportance: should knowfreq 38%

basics

~10 s

str.lower() maps characters to their lowercase form for display. str.casefold() applies the stronger Unicode case-folding mappings meant for caseless matching, so "Stra\u00dfe".casefold() gives "strasse" while .lower() leaves the sharp s alone.

open as a page

How do NFC and NFD differ from NFKC and NFKD in unicodedata.normalize?

level: middleimportance: should knowfreq 40%

basics

~20 s

NFC and NFD are canonical: they only compose or decompose characters that are already equivalent, and they round-trip. NFKC and NFKD add compatibility folding, which rewrites formatting variants such as ligatures and Roman numerals and is lossy.

open as a page

A ticket-triage bot slices each Python str title to a fixed character budget and accents vanish - why, and how do you cut safely?

level: seniorimportance: should knowfreq 30%

basics

~20 s

len() and slicing count code points, not what a reader sees. If a title is decomposed, an accent is a separate combining code point after its base letter, so a cut can land between them and drop the accent. Normalize to NFC first and back off before a combining mark.

open as a page