Why can truncating text at a fixed UTF-16 code-unit index emit an invalid character?
answer
- sixteen bits cannot address every code point
- some characters occupy two storage slots
- one reserved range holds only halves
- a cut can land between the halves
- half a pair is not a character
basics
~20 sCode points beyond the Basic Multilingual Plane are stored as two UTF-16 code units, a surrogate pair. Cutting between them leaves an unpaired surrogate, which is not a valid character and renders or serializes as a replacement symbol.
solid answer
~40 sUTF-16 stores code points above U+FFFF as a *surrogate pair*: a high unit in U+D800–U+DBFF followed by a low unit in U+DC00–U+DFFF. Neither half means anything alone. A preview that keeps the first 200 code units can cut exactly between the halves, and the survivor is an unpaired surrogate — invalid text that renders as a replacement glyph and cannot be encoded to UTF-8 at all, so the failure often surfaces later, at the serialization boundary, far from the slice. The minimum fix is to check whether the last kept unit is a high surrogate and drop it if so. The honest fix is to truncate by grapheme cluster, because even a well-formed cut can strip a combining accent or split a joined sequence and change what the user sees.
code
pseudocode · 11 lines// s is an array of UTF-16 code units; keep at most 200
limit = 200
if length(s) <= limit:
out = s
else:
out = s[0 .. limit-1] // hard cut at a fixed index
...
// if s[limit-1] is in 0xD800..0xDBFF it is a HIGH surrogate
// and its LOW partner at s[limit] was just discarded,
// so out now ends in half a character
// safe cut: if high_surrogate(s[limit-1]): out = s[0 .. limit-2]go deeper
Know that some characters occupy two storage slots in UTF-16 and that cutting between them leaves something that is not a character. Recognizing the replacement glyph as a symptom is enough here.
Explain the surrogate ranges, why they are reserved, and exactly what the naive slice keeps. Then give the constant-time repair: check the last kept unit and back off one slot if it is a high surrogate.
Separate validity from meaning. Show that the surrogate check stops invalid output while combining marks and joined sequences still need grapheme-aware iteration, and say where in the pipeline you would enforce each.
Decide the policy: which unit the product's limit is expressed in, whether previews are grapheme-aware everywhere, and how you stop this class of slice from being reintroduced across teams rather than fixing one call site.
## What a surrogate pair is Unicode defines code points from U+0000 to U+10FFFF — more than sixteen bits can address. UTF-16 stores each code point in 16-bit *code units*, so it needs an escape hatch for the roughly one million code points above U+FFFF. That hatch is the **surrogate pair**: a code point above U+FFFF is written as two units, a high (or leading) surrogate in the range U+D800–U+DBFF followed by a low (or trailing) surrogate in U+DC00–U+DFFF. That range was permanently reserved for the purpose, so no real character lives there and the two halves are unambiguous. The consequence is that a UTF-16 string is **variable-width**, exactly like UTF-8 — it is just less obvious, because everything an English-speaking team types in a test fixture lives in the Basic Multilingual Plane and occupies one unit. Emoji, many rare CJK characters, mathematical alphanumerics and most historic scripts do not. ## Why a fixed-index cut breaks A truncation that keeps the first `limit` code units is doing arithmetic on storage slots, not on characters. If unit `limit-1` happens to be a high surrogate, its partner sits at index `limit` and has just been thrown away. What remains is an **unpaired surrogate**: a value that is structurally legal to hold in memory but is not a character. Three things then go wrong, usually in this order: 1. **Rendering.** The text draws with a replacement glyph — the reader sees a box or a question-mark diamond at the end of every truncated preview containing an emoji. 2. **Serialization.** UTF-8 has no representation for a lone surrogate. Encoding the truncated text either throws, or silently substitutes U+FFFD, so the corruption becomes permanent the moment it is written to a store or a wire format. 3. **Comparison and hashing.** The stored value no longer round-trips, so equality checks against the original, or against a re-encoded copy, quietly fail. The distance between cause and symptom is what makes this expensive: the slice happens in a formatting helper, the exception surfaces in a serializer three layers away, and the stack trace names neither the message nor the user. ## The layered fix There are three levels of correctness here, and it is worth knowing which one you are buying. - **Surrogate-safe.** Before returning, check whether the last kept unit is a high surrogate; if it is, drop it. Cheap, O(1), and it removes the invalid-text class of bug entirely. This is the minimum bar for anything that will be serialized. - **Code-point-safe.** Iterate the text by code point and stop after `n` of them. Now the cut never lands mid-character, and the count you report matches the catalogue entries you kept. - **Grapheme-safe.** Iterate by grapheme cluster — the user-perceived character. This is the only level that survives a combining accent (a base letter followed by a mark), a flag built from two regional indicators, or a joined sequence held together by zero-width joiners. Cut inside any of those and the text stays *valid* while changing meaning: a family becomes two people, a flag becomes two letters, an accent jumps to the previous letter. Which level you need is a product decision, not a technical one. A preview shown to humans wants grapheme-safe. A hard storage limit wants byte-safe. Nothing wants unit-safe, which is what the naive slice gives you. ## Why the bug is so common Indexing a string by an integer is the single most natural operation in programming, and it is the one that variable-width encodings cannot honour. Mainstream runtimes disagree about what that integer even means: some index strings by UTF-16 code units, others by UTF-8 bytes, others by code points — so the same slice expression is a different operation depending on where it runs, and code ported between them inherits a boundary bug it never had before. On top of that, the failure is data-dependent. Every ASCII test passes. Every review passes, because the code looks like the textbook. The bug ships and then waits for the first user whose message ends in an emoji at exactly the wrong offset. That is the profile of a defect worth recognizing on sight in review: any expression that slices text at an integer computed from a limit deserves the question "what unit is that, and can it land inside a character?" ## What to say in an interview Name the mechanism (surrogate pair, reserved range, two units for one code point), name the artifact (unpaired surrogate, replacement character, failure to encode), and then separate *validity* from *meaning*: the surrogate check buys validity, and only grapheme-aware iteration buys the meaning the user expects.
- What happens if that truncated text is later encoded as UTF-8?UTF-8 has no encoding for an unpaired surrogate, so the encoder either raises an error or substitutes the replacement character U+FFFD. Either way the damage is discovered far from the slice that caused it, which is why this bug is usually reported as a serializer failure or as mojibake in stored data rather than as a truncation defect.
- Does UTF-8 have the same truncation hazard?Yes, at a different granularity. Cutting a UTF-8 byte array at a fixed offset can land between a lead byte and its continuation bytes, orphaning both halves. UTF-8 is easier to repair because it is self-synchronizing: continuation bytes are recognizable, so you can walk back at most three bytes to the nearest boundary in constant time.
- Is a surrogate-safe cut enough for a user-facing preview?No. It guarantees valid text, not sensible text. A cut can still separate a base letter from its combining accent, split a flag made of two regional indicators, or break a joined sequence into its parts. For anything a person reads, truncate on grapheme-cluster boundaries and treat the surrogate check as the floor, not the goal.
saying these in an interview costs you the question
- Says every character is one code unit in UTF-16
- Calls UTF-16 a fixed-width encoding
- Thinks a surrogate half is just an unusual character
- Believes the count reported by a slice equals characters
- Fixes rendering only and ignores the encoding failure