A service trims a UTF-8 text field so it fits a byte-length budget on the wire; what can go wrong, and why?
answer
- a length counts bytes, not letters
- one character, one to four bytes
- the cut can land mid-character
- continuation bytes are recognisable on sight
- back up at most three bytes
basics
~20 sA length counts bytes, but a UTF-8 character takes one to four of them, so a cut at an arbitrary byte position can land inside a character. The field then carries an incomplete sequence that a strict decoder rejects outright.
solid answer
~50 sIn UTF-8 a character occupies one to four bytes: a lead byte that announces the length, then continuation bytes, each of the form `10xxxxxx`. A byte budget cuts wherever the count runs out, which may be between a lead byte and its continuations. Decoders then diverge - a strict one rejects the field as malformed, a lenient one substitutes a replacement character - so the same bytes behave differently on different readers. The fix is to move the cut backwards to a character boundary, which the encoding makes cheap: continuation bytes are recognisable on sight, so backing up costs at most three steps. Two further traps follow: a byte budget admits far fewer characters in some scripts than in others, and even a code-point-clean cut can split a combining sequence and change what the user sees.
code
pseudocode · 7 linesfunction truncate_to_bytes(bytes, budget):
if length(bytes) <= budget:
return bytes
cut = budget // keep bytes[0 .. cut-1]
while cut > 0 and is_continuation(bytes[cut]):
cut = cut - 1 // bytes[cut] is 10xxxxxx, so a character straddles the cut
return bytes[0 .. cut - 1] // bytes[cut] is now a lead byte: the cut is on a boundarygo deeper
Remember that UTF-8 characters are one to four bytes long, so a limit counted in bytes is not a limit counted in characters, and cutting at a byte position can land in the middle of one.
Explain the lead-byte and continuation-byte patterns, and how the self-synchronising property lets a writer back a cut up to a character boundary in at most three steps.
Show what the bug looks like in production: failures that cluster by script and never reproduce on ASCII fixtures, strict readers rejecting while lenient ones substitute, and a round trip that makes the corruption permanent.
Make the unit of the limit an explicit contract term - bytes, code points or grapheme clusters - decide where truncation is allowed to happen at all, and weigh the fairness cost of a byte budget across the languages you serve.
## How UTF-8 lays a character out UTF-8 is a variable-length encoding of Unicode code points. The lead byte's high bits say how many bytes the character occupies; every following byte of the same character has the fixed prefix `10`. | Code point range | Bytes | Lead byte pattern | |---|---|---| | U+0000 - U+007F | 1 | `0xxxxxxx` | | U+0080 - U+07FF | 2 | `110xxxxx` | | U+0800 - U+FFFF | 3 | `1110xxxx` | | U+10000 - U+10FFFF | 4 | `11110xxx` | Continuation bytes are always `10xxxxxx`. Two properties follow, and both matter here: - **No byte of a multi-byte character is below 0x80.** A byte-oriented marker drawn from the ASCII range can therefore never be mistaken for part of a character, which is why byte-level scanning over UTF-8 is safe. - **The encoding is self-synchronising.** From any position, you can tell whether you are inside a character, because only continuation bytes look like `10xxxxxx`. Moving back at most three bytes always lands on a boundary. ## Why the length counts bytes A wire format's length prefix counts bytes because that is the only count the reader can act on directly: it takes that many bytes and moves on. A count of characters would force the reader to decode the text before it knew where the field ended - work it may not want to do at all, for a field it intends to skip. So the count is in bytes, and the mismatch with how people think about text length is structural, not accidental. ## What a mid-character cut produces Cutting at an arbitrary byte can leave a lead byte announcing three bytes with only one of them present. The result is not text with a missing letter, it is an **invalid byte sequence**, and what happens next is not uniform: - a strict decoder rejects the whole field, which typically surfaces as a decode failure on the far side rather than as a text problem where the truncation happened; - a lenient decoder substitutes a replacement character, so the data is quietly corrupted and only the affected users notice; - a re-encode round trip through a lenient decoder makes the corruption permanent, because the replacement character is itself valid text and the original bytes are gone. The operational signature is distinctive: failures cluster on records whose text is in a particular script, or on records near the size limit, and never reproduce with test fixtures written in plain ASCII. ## Cutting safely The repair is to move the cut backwards until it sits on a character boundary, which the self-synchronising property makes cheap and exact. Beyond that, two traps remain. 1. **A byte budget means different things to different users.** A 100-byte field holds 100 characters of ASCII, roughly 50 in a script whose characters take two bytes, about 33 where they take three, and 25 for a four-byte character. A limit stated as a byte count is a storage bound, and presenting it as a length limit silently discriminates by language. 2. **A code point is not what a user calls a character.** A base letter plus a combining mark, or a sequence of code points joined into one visible symbol, is several code points forming one **grapheme cluster**. A cut that is clean at the code-point level can still strip the mark off its base or split a joined sequence, changing or breaking the visible character even though the bytes are valid. ## What to settle in the contract - **Say which unit the limit is in** - bytes, code points, or grapheme clusters - and say it where the field is defined, not in the code that happens to enforce it. All three are defensible; leaving it implicit is not. - **Validate after truncating, not before.** Truncation is the step that can create an invalid sequence, so it is the step whose output needs checking. - **Reject rather than silently repair on input you control.** Substituting a replacement character turns a detectable error into data. - **Truncate once, at a boundary you chose.** The worst outcome is several layers each trimming to their own budget, because every one of them is an opportunity to cut inside a character.
- How does a reader find the start of a character from an arbitrary position in a UTF-8 buffer?It walks backwards while the byte matches the continuation pattern `10xxxxxx`, and stops at the first byte that does not - that byte is a lead byte or a single-byte ASCII character. Because no character exceeds four bytes, this costs at most three steps. That self-synchronising property is a deliberate design feature of the encoding, not a lucky accident.
- Why does a 100-byte limit on a name field mean different things to different users?Because the byte cost per character depends on the script. It admits 100 ASCII characters, around 50 where characters take two bytes, about 33 where they take three, and 25 for four-byte characters. A limit expressed in bytes is a storage bound; presenting it to people as a length limit quietly gives some languages a quarter of the room.
- Does cutting on a code-point boundary always preserve what the user sees?No. Several code points can combine into one visible symbol - a base character with a combining mark, or a joined sequence - which is a grapheme cluster. A cut between them is perfectly valid at the byte and code-point level but still strips a mark off its base or breaks a joined symbol, so a user-facing limit should be counted in grapheme clusters.
saying these in an interview costs you the question
- Assumes one character is always one byte in UTF-8
- Treats a byte limit and a character limit as interchangeable
- Says every decoder silently repairs a truncated sequence
- Believes a code-point-safe cut always preserves what the user sees
- Thinks a continuation byte cannot be told from a lead byte