skip to content

A ticket-triage bot slices each Python str title to a fixed character budget and accents vanish - why, and how do you cut safely?

level: seniorimportance: should knowfreq 30%

answer

  1. The count and the eye disagree
  2. Indexing walks code points, not characters
  3. A mark is its own code point
  4. Normalize first, then guard the cut
  5. unicodedata.combining is nonzero on marks

basics

~20 s

len() and slicing count code points, not what a reader sees. If a title is decomposed, an accent is a separate combining code point after its base letter, so a cut can land between them and drop the accent. Normalize to NFC first and back off before a combining mark.

solid answer

~50 s

A Python `str` is indexed by code point, so `title[:40]` cuts after 40 code points, not 40 visible characters. When text arrives decomposed, an accented letter is two code points — base plus a combining mark whose `unicodedata.combining` value is nonzero — so a cut that lands between them either drops the accent or leaves a mark that the renderer glues onto whatever character precedes it. Emoji sequences and other multi-code-point clusters break the same way, only more visibly. The fix has two parts: normalize to NFC at the boundary, which collapses most Latin accents into one code point and removes the ordering variability between stacked marks; then, before slicing, check whether the character at the cut index is a combining mark and back the cut off until it is not. If exact visible-character counts matter — CJK width, emoji, Indic clusters — code points are the wrong unit entirely and you need grapheme clustering.

code

python · 5 lines
python
import unicodedata

title = unicodedata.normalize("NFD", "Caf\u00e9 outage")
print(len(title), len("Caf\u00e9 outage"))  # 12 11
print(repr(title[:4]))                       # 'Cafe'  - the accent is gone

go deeper

for a junior

Know that len() on a Python str counts code points and that an accented letter is sometimes two of them. Recognising that slicing can therefore cut inside what looks like one character is the takeaway.

for a middle

Explain the mechanism: a combining mark is a separate zero-width code point after its base, unicodedata.combining is nonzero for it, and normalizing to NFC composes most Latin accents into one code point.

for a senior

Demonstrate the production instinct: normalize once at the boundary, guard the cut index, and say out loud which unit the requirement really means - bytes for storage, code points for a loose cap, grapheme clusters for layout. Note that the failure is silent and data-dependent.

for a principal

Own the decision about the unit and the dependency: whether the platform needs real grapheme segmentation or an approximation is good enough, where normalization is enforced, and what the latency and correctness budget is for text handling across services.

## What `len()` actually measures `len(s)` on a `str` returns the number of **code points**, and `s[:n]` cuts after `n` of them. That is a precise, useful definition, and it is not the definition a product manager has in mind when they say "trim titles to 40 characters". Three units are in play and they routinely disagree: - **bytes** — what a database column, a wire protocol or a log budget usually limits - **code points** — what `len()` and slicing count - **grapheme clusters** — what a reader calls a character The gap between the second and third is where this bug lives. ## Why an accent can disappear Text that arrives decomposed spells an accented letter as a base character followed by a **combining mark** — a zero-width code point that the renderer draws over the character before it. `"Caf\u00e9 outage"` is 11 code points; the NFD form of the same title is 12, because the accent is its own code point. Cut that decomposed title at 4 code points and you get `"Cafe"`: the base `e` survived, the mark did not, and the accent is silently gone. Cut it at 5 and you keep the mark. Cut a *different* title so that the slice starts with an orphaned mark and the renderer stacks that accent onto whatever character precedes it in the surrounding text — an ellipsis, a quote mark, a space. Nothing raises; the output is just wrong, and it is wrong only for some inputs. There is an ordering assumption underneath the naive slice: the code assumes each index step advances one visible character, and that assumption holds for every ASCII fixture in the test suite. Two related ordering facts make hand-rolled fixes fragile — a base can carry several marks, and different sources stack them in different orders, which is exactly what normalization's canonical ordering exists to settle. ## The safe cut Two steps, in order: 1. **Normalize to NFC on the way in.** For Latin text this collapses base-plus-mark back into one precomposed code point, so most cuts stop being dangerous at all, and it puts any remaining stacked marks into canonical order. 2. **Do not cut immediately before a combining mark.** Look at the character at the cut index; while `unicodedata.combining(ch)` is nonzero, move the cut one character earlier. `unicodedata.category(ch) == "Mn"` identifies the same characters as nonspacing marks. ```python import unicodedata def truncate(text: str, limit: int) -> str: text = unicodedata.normalize("NFC", text) cut = text[:limit] while cut and len(text) > len(cut) and unicodedata.combining(text[len(cut)]): cut = cut[:-1] return cut ``` That is honest about its scope: it prevents a mark from being orphaned or a base from being stripped of its mark. It does not make the count equal to what a reader sees. ## When code points are the wrong unit entirely For scripts and sequences where one visible character is many code points by design, no amount of normalization helps: - an emoji with a skin-tone modifier, or a sequence joined by ZERO WIDTH JOINER, is one cluster of several code points - a flag is two regional indicator code points - Indic and Thai text has clusters that NFC does not collapse - East Asian characters are one code point but two display columns wide, which matters for fixed-width layout The standard's answer is **grapheme cluster** segmentation, which the standard library does not implement; `unicodedata.east_asian_width` tells you about display width but not clustering. So the senior judgement is to decide which unit the requirement actually means. If it is a storage or protocol limit, measure `len(s.encode("utf-8"))` and cut on an encoded boundary. If it is a layout limit, you need width or clustering, and that is a dependency decision, not a slicing trick. If it is a loose "keep titles short", code points plus the combining-mark guard is enough. ## Operating it Normalize once when the record is created, not on every classification pass: on a hot path with a 92nd-percentile latency budget, per-request Unicode work on every stored title is avoidable cost, and the normalized value is the thing every later comparison and cut should use. Then make the failure visible — assert in tests with one decomposed and one emoji fixture, because an ASCII-only suite passes with the broken version. ## The short answer "`len` counts code points. A decomposed accent is two of them, so a slice can land between a letter and its mark. Normalize to NFC at the boundary and back the cut off while `unicodedata.combining` of the next character is nonzero. And if the requirement really means visible characters, code points are the wrong unit and I would say so."

  • The requirement is a database column limit rather than a display limit. What changes?
    The unit changes from code points to bytes in the storage encoding. Measure `len(s.encode("utf-8"))` and cut on an encoded boundary, because one code point is one to four UTF-8 bytes and a code point budget can overflow a byte column. Cutting encoded bytes needs its own guard so you never split a multi-byte sequence.
  • Why does normalizing to NFC not make len() equal what a reader counts?
    NFC only composes sequences for which a single precomposed code point exists — largely Latin, Greek and Cyrillic accents. Emoji sequences joined by ZERO WIDTH JOINER, flags built from two regional indicators, and many Indic clusters have no precomposed form, so they remain several code points that render as one character no matter which form you normalize to.
  • How would you catch this class of bug before it reaches production?
    Put non-ASCII fixtures in the tests: one decomposed accented string, one emoji sequence, one CJK string. An ASCII-only suite exercises the same path for correct and broken implementations, so the defect ships silently and surfaces as mangled output in a user report rather than as a failing assertion or an exception.

Trimming decomposed text by code point is like cutting a page at a fixed number of pen strokes: sometimes you stop halfway through a letter, and the reader sees a different word rather than an error.

saying these in an interview costs you the question

  • Assumes one index step equals one visible character
  • Thinks the bug would raise an exception
  • Believes NFC makes len match what a reader counts
  • Confuses code point count with UTF-8 byte length
  • Strips all combining marks instead of moving the cut
  • Claims the standard library segments grapheme clusters

context