What is the difference between str.casefold() and str.lower() in Python?
answer
- Two ways to lowercase a Python string
- One is for display, one for matching
- Folding may change the string's length
- German sharp s expands to two letters
- casefold, since Python 3.3
basics
~20 sstr.lower() applies simple lowercasing for display. str.casefold() applies Unicode full case folding, which is more aggressive: German ß becomes ss and Greek final sigma becomes sigma. Use casefold for case-insensitive comparison, lower for showing text.
solid answer
~40 sBoth return a new `str`, but they answer different questions. `str.lower()` maps each character to its lowercase form and is meant for **display** — it leaves German ß alone, so `"Straße".lower()` and `"STRASSE".lower()` still differ. `str.casefold()` implements Unicode full case folding, whose only job is to produce a **comparison key**: it expands ß to `ss`, maps Greek final sigma ς to σ, and folds other characters that simple lowercasing cannot reconcile. So case-insensitive matching should compare `a.casefold() == b.casefold()`. Casefolded text is deliberately not human-facing and is not reversible, so never store it back over the original. Neither method is locale-aware — Turkish dotted/dotless i is wrong under both — and neither normalizes combining marks, so a robust comparison key applies `unicodedata.normalize` around the fold as well.
code
pycon · 6 lines>>> 'Straße'.lower() == 'STRASSE'.lower()
False
>>> 'Straße'.casefold() == 'STRASSE'.casefold()
True
>>> len('Straße'), len('Straße'.casefold())
(6, 7)go deeper
Be ready to state the one-line rule: lower() for display, casefold() for comparing. Knowing one concrete example — German ß folding to ss — is enough to show you understand why two methods exist.
Explain the mechanics: simple per-character lowercase mapping versus Unicode full case folding that may change length and is irreversible. Expect to be asked why a folded key must not be written back over stored text.
Demonstrate the production judgment of building one comparison key — normalize, fold, normalize again — and applying it consistently at every place that decides identity, so a lookup and the write that created it cannot disagree.
Own the policy: which stored column is display text, which is the match key, who recomputes keys when the interpreter's Unicode database advances, and whether identity comparisons live in application code or in the storage layer at all.
## Two operations that look like one `str.lower()` and `str.casefold()` both return a new string with the case "removed", and on plain ASCII they are indistinguishable: `"HELLO".lower()` and `"HELLO".casefold()` are both `"hello"`. The difference only appears outside ASCII, and it is a difference of **purpose**, not of quality. `str.lower()` implements Unicode's simple lowercase mapping. It is a per-character table lookup: every code point maps to at most one code point, and the result is meant to be *shown to a human*. Simple mappings are conservative on purpose — they preserve the identity of a letter rather than rewriting it. `str.casefold()`, added in Python 3.3, implements Unicode **full case folding**. Case folding is defined by the Unicode standard for exactly one job: producing a key that compares equal for two strings that differ only by case. It is allowed to change the length of the string and it is allowed to be lossy, because nobody is supposed to read the output. ## The examples an interviewer expects German sharp s is the canonical case. `"Straße".lower()` is `"straße"`, while `"STRASSE".lower()` is `"strasse"` — the two compare unequal even though a German reader would call them the same word. `casefold()` expands ß to `ss`, so both sides become `"strasse"` and the comparison succeeds. Greek final sigma is the second example: ς appears only at the end of a word and σ elsewhere, `lower()` keeps them distinct, and `casefold()` maps ς to σ so the same word written in either position matches. ## What casefold does not do **It is not locale-aware.** Turkish treats dotted and dotless i as separate letters, and neither `lower()` nor `casefold()` knows that; the Turkish capital İ folds to a two-code-point sequence (`i` plus a combining dot) under both. If you genuinely need Turkish casing rules, they are not in the standard library's string methods. **It does not normalize.** Case folding operates on the code points it is given. A name typed with a precomposed accent and the same name typed with a base letter plus a combining accent still differ after `casefold()`, because folding never merges canonically equivalent sequences. A comparison key that must survive real user input therefore folds *and* normalizes, and the Unicode definition of caseless matching applies the normalization on both sides of the fold, because folding can itself leave the result unnormalized: ```python import unicodedata def match_key(s: str) -> str: return unicodedata.normalize('NFD', unicodedata.normalize('NFD', s).casefold()) ``` **It is not reversible.** You cannot recover ß from `ss`. Store the original string for display and compute the folded key separately — as a second column, an index, or a value derived on the fly. Overwriting the stored value with its folded form silently corrupts names. **It is not a security boundary by itself.** Folding makes `A` and `a` match; it does not make visually confusable characters from different scripts match, and it does not strip zero-width characters. ## Choosing between them The rule fits in one line: **`lower()` for humans, `casefold()` for comparisons.** Anything that decides whether two pieces of text are "the same" — a login lookup, a deduplication pass, a search index, a set membership test on labels — should fold. Anything that will be rendered — a heading, a slug shown to a user, a message — should use `lower()` or, more often, should keep the user's own casing. One practical consequence: because casefolding can change length, do not fold a string and then reuse offsets computed against the original. `len("Straße")` is 6 and `len("Straße".casefold())` is 7. Fold, then index the folded string, or index the original — never mix the two. Finally, both methods depend on the Unicode data bundled with the interpreter. CPython 3.14 ships Unicode 16.0.0 (`unicodedata.unidata_version` reports it), and that table advances between feature releases, so a folded key persisted years ago is not guaranteed byte-identical to one computed today for exotic characters.
- Is str.casefold() safe to store in place of the user's original text?No. Folding is deliberately lossy and irreversible — ß becomes ss, final sigma becomes sigma — so the original spelling cannot be recovered. Keep the user's string as the display value and treat the folded form as a derived comparison key computed on demand or stored alongside, never over, the original.
- Does str.casefold() make two visually identical names compare equal?Not on its own. Folding only removes case distinctions; it does not merge canonically equivalent sequences, so a precomposed accented letter and a base letter plus a combining accent still differ afterwards. A reliable key applies unicodedata.normalize around the fold, since folding can leave its own output unnormalized.
- Why does neither method handle Turkish casing correctly?Both implement Unicode's default, locale-independent mappings. Turkish needs dotted and dotless i kept apart, which is a locale-tailored rule; Python's str methods never consult a locale, so the Turkish capital İ folds to i plus a combining dot under both. Locale-specific casing needs a tailoring the string methods do not provide.
lower() is tidying a name for the nameplate; casefold() is squashing it into a filing code you would never hand back to the person.
saying these in an interview costs you the question
- Claiming casefold() is just a slower alias for lower()
- Using lower() as the key for a case-insensitive login lookup
- Storing the casefolded string back over the user's name
- Assuming casefolding preserves the string's length
- Believing casefold() also normalizes combining accents
- Expecting either method to follow the process locale