skip to content

Why does len() on a Python str mislead offset math that slices labels at fixed positions?

level: seniorimportance: should knowfreq 33%

answer

  1. The unit len() counts is not the visible one
  2. One mark on screen, several entries in the string
  3. Slicing can separate a letter from its accent
  4. Composed form makes the counts agree, mostly
  5. No grapheme segmentation in the standard library

basics

~20 s

len() counts code points, not the characters a reader sees. One visible mark can be several code points, so any offset computed from len() lands mid-character and slicing splits a combining mark from its base letter, shifting every boundary after it.

solid answer

~50 s

`len(s)` on a `str` counts **code points**, and `s[i]` indexes code points. A user-perceived character — a grapheme cluster — can be several: a base letter plus combining marks, an emoji joined by zero-width joiners, a regional-indicator pair. So a label written with a decomposed accent has `len()` one greater than its visible length, and a slice taken at a position computed from a smaller sample lands *between* a letter and its accent, producing an off-by-one boundary that reappears in every record after the first non-ASCII one. The practical fixes, in order: normalize input to NFC so precomposed characters count as one, refuse to derive offsets from length at all and parse by delimiter instead, and use `str.isascii()` to prove a fast path is safe. `unicodedata.category(ch) == 'Mn'` identifies a non-spacing mark when you must inspect a boundary. The standard library has no grapheme segmentation; NFC only shrinks the problem, it does not remove it.

code

python · 7 lines
python
import unicodedata

label = unicodedata.normalize('NFD', 'Café-01')
print(len(label))                      # 8, not 7
print(repr(label[:4]))                 # 'Cafe' -- accent orphaned
print(unicodedata.category(label[4]))  # Mn
print('Cafe-01'.isascii(), label.isascii())  # True False

go deeper

for a junior

Recall that len() on a str counts code points and that a single visible character can be more than one, so length is not a count of what you see on screen.

for a middle

Explain the mechanics: combining marks, why the decomposed spelling adds to len(), how a slice can orphan a mark, and how normalizing to NFC makes the counts agree for most text.

for a senior

Demonstrate the diagnosis — repr and len on the suspect value, unicodedata.category at the boundary — and the fix in priority order: parse by delimiter, normalize at input, guard an ASCII fast path explicitly.

for a principal

Own the design call: whether the system may count characters at all, whether a text-segmentation dependency is justified for a user-visible requirement, and how the input contract is stated so partner data cannot reintroduce the drift.

## Three different things called a "character" Text has at least three units, and Python only gives you one of them for free. - A **code point** is one entry in the Unicode codespace. A `str` is a sequence of these; `len()` counts them and `s[i]` indexes them. - A **code unit** is a piece of an encoding, and is what `bytes` holds after `encode`. Python's `str` hides this entirely, which is a genuine advantage over languages whose strings are UTF-16 code units. - A **grapheme cluster** is what a reader calls a character: the smallest unit you would delete with one press of backspace. This is the one nobody's `len()` returns. The gap between the first and the third is where offset bugs live. ## How the boundary goes wrong Consider a pipeline that annotates genomic records and derives display labels from sample names, slicing each label at positions computed from the length of a prefix. On ASCII sample names, code points and visible characters coincide and every boundary is right. The day a partner site contributes names carrying accents typed as base letter plus combining mark, `len()` for those names is one greater per accent, every derived offset shifts, and slices land between a letter and its accent. The symptom is subtle in the worst way: the accent, orphaned onto the following slice, renders on whatever character now precedes it, so the output looks *almost* right and the reviewer's eye slides over it. A small team reviewing each other's output will pass it several times before someone prints `repr()`. ```python import unicodedata label = unicodedata.normalize('NFD', 'Café-01') print(len(label)) # 8, not 7 print(repr(label[:4])) # 'Cafe' -- the accent was left behind print(unicodedata.category(label[4])) # 'Mn' ``` ## Diagnosing it The tell is that a length disagrees with a count of what you can see. `repr()` on the value shows the escapes; `len()` on both a suspect and a known-good value shows the drift; `unicodedata.category(ch)` on the character at the boundary returns `'Mn'` — a non-spacing mark — which proves the slice cut inside a cluster rather than between two letters. `unicodedata.combining(ch)` returns a non-zero combining class for the same reason. `str.isascii()` on the batch tells you instantly whether the input is even capable of the problem. ## Fixing it, in order of preference **Stop deriving offsets from length.** The real defect is usually not Unicode at all: it is positional parsing of data that has a delimiter. Split on the delimiter, or match a pattern, and the code point question disappears. **Normalize to NFC at input.** `unicodedata.normalize('NFC', s)` composes base-plus-mark sequences back into single precomposed code points wherever Unicode has one, which makes `len()` match the visible count for the large majority of European text. This shrinks the class of failures dramatically without eliminating it: many scripts have no precomposed forms, and emoji sequences never compose. **Use an ASCII fast path deliberately.** `if s.isascii():` is cheap and, when true, guarantees that code points, grapheme clusters and UTF-8 bytes are all in one-to-one correspondence. Making that assumption *explicit* is much better than making it implicitly and being wrong later. **Segment properly only if you must.** True grapheme-cluster segmentation follows a Unicode annex, and CPython does not implement it in the standard library. If a product genuinely needs a user-perceived character count — a truncation with an ellipsis, a fixed-width column, a cursor — that requires a text-segmentation library from outside the standard library, and the cost of taking on that dependency should be weighed against redesigning so no code counts characters. ## The cases beyond accents Emoji make the gap vivid and are worth knowing as examples: a family emoji built from three people joined by zero-width joiners is five code points and one visible glyph; a flag built from two regional indicators is two code points and one glyph; a skin-tone modifier adds another. None of these compose under NFC. Any code that truncates user-supplied text to "20 characters" using `len()` will eventually cut one of these in half and emit a broken sequence. The senior position to hold in the interview is not "use a segmentation library". It is: know which unit each API counts, normalize at the boundary so the common case behaves, avoid designs that depend on counting characters at all, and only reach for segmentation when a genuine user-visible requirement demands it.

  • Does normalizing to NFC make len() equal the number of visible characters?
    Only usually. NFC composes base-plus-mark sequences into single code points wherever Unicode defines a precomposed character, which covers most European text. Many scripts have no precomposed forms, and emoji built from zero-width joiners or regional-indicator pairs never compose, so the counts still diverge. NFC shrinks the failure class; it does not close it.
  • How would you prove a slice cut inside a grapheme cluster rather than between letters?
    Inspect the character at the boundary. unicodedata.category returns 'Mn' for a non-spacing mark and unicodedata.combining returns a non-zero combining class, either of which means the code point belongs to the preceding base character. Printing repr of both slices shows the orphaned mark directly, which is faster than reasoning about the rendered output.
  • When is it legitimate to rely on len() as a character count?
    When the string is ASCII and you have checked it. str.isascii() is cheap and, when true, guarantees code points, visible characters and UTF-8 bytes correspond one to one. Making that precondition explicit in a guard is fine engineering; assuming it silently is the defect. For anything else, prefer parsing by delimiter over counting positions.

Measuring a printed word by counting ink strokes rather than letters: fine for block capitals, wrong the moment a letter carries an accent drawn as a second stroke.

saying these in an interview costs you the question

  • Claiming len() returns the number of visible characters
  • Believing Python strings are UTF-16 code units
  • Assuming NFC guarantees one code point per glyph
  • Truncating user text by slicing at a fixed index
  • Thinking the standard library segments grapheme clusters
  • Blaming the display layer instead of printing repr

context