skip to content

Why can a 10-character name occupy 34 bytes when stored as UTF-8?

level: juniorimportance: must knowfreq 70%

answer

  1. think about what one byte can hold
  2. 128 values do not cover every script
  3. the encoding spends one to four bytes
  4. emoji sit above the two-byte range
  5. ten code points, up to forty bytes

basics

~10 s

UTF-8 is variable-width: one code point takes one to four bytes. Unaccented Latin letters cost one byte, accented letters two, most emoji four. So ten characters can weigh anywhere from ten to forty bytes.

solid answer

~40 s

UTF-8 encodes each code point in one to four bytes. The ASCII range costs one byte, common accented Latin and Greek letters two, most CJK three, and code points beyond the Basic Multilingual Plane — which includes most emoji — four. A display name of seven emoji plus three precomposed accented letters is ten code points but `7*4 + 3*2 = 34` bytes. That is why a 255-byte column is not a 255-character column: the two numbers diverge the moment text leaves ASCII. If a product promises "up to 100 characters", decide which count the promise means, validate that count, and size storage for the worst case of roughly four bytes per code point instead of assuming the two limits are the same number.

go deeper

for a junior

Be ready to say that UTF-8 spends one to four bytes per code point and that a byte count and a character count are different numbers for anything beyond ASCII.

for a middle

Explain which ranges cost which widths and why ASCII costs one byte by design, then show the arithmetic on a mixed example instead of hand-waving that it is "more than one byte".

for a senior

Show where the mismatch bites: a client, an API and a store each counting a different unit, so the failure only appears for non-ASCII users. Say how you would pin one unit in the contract.

for a principal

Own the policy question: which unit the product promises, where it is enforced, and what the migration costs when a byte-sized column has to grow for text that was always legal.

## Two different questions "How long is this text?" hides at least three questions that give different answers: 1. **How many bytes** does it occupy in a given encoding? This is what storage limits, network payloads and quotas care about. 2. **How many code points** does it contain? A code point is one entry in the Unicode catalogue — a number like U+0041 (`A`) or U+1F600 (a grinning face). 3. **How many user-perceived characters** does a reader see? This is the *grapheme cluster* count, and it can be smaller still, because one visible character may be several code points. For pure ASCII all three numbers coincide, which is exactly why the confusion survives so long: every test fixture written by an English-speaking team agrees with every assumption. ## What UTF-8 actually spends UTF-8 is a variable-width encoding. It maps each code point to between one and four bytes: | code point range | bytes in UTF-8 | typical content | | --- | --- | --- | | U+0000–U+007F | 1 | ASCII letters, digits, punctuation | | U+0080–U+07FF | 2 | accented Latin, Greek, Cyrillic, Hebrew, Arabic | | U+0800–U+FFFF | 3 | most CJK, most other scripts, many symbols | | U+10000–U+10FFFF | 4 | emoji, rare CJK, historic scripts | The design is deliberate: ASCII text is byte-identical to its ASCII encoding, so the common case costs nothing extra, and the price of the rest of the world is paid only by text that uses it. So the 34 bytes in the question are not a mystery. Seven emoji, each a single code point above U+FFFF, cost four bytes each — 28. Three precomposed accented letters cost two bytes each — 6. Ten code points, 34 bytes. The same name is also 17 UTF-16 code units (each emoji is a surrogate pair of two units, each accented letter one unit) and, in UTF-16, also 34 bytes — a coincidence for this particular mix, not a rule. Fill the same name with ASCII and UTF-8 halves the UTF-16 cost; fill it with CJK and UTF-16 wins, three bytes against two. ## Where this bites in practice The classic bug report is "the form says 100 characters, the user typed far fewer, and the save failed." Three layers each counted something different: - The client counted code units, or code points, or grapheme clusters, depending on how it measured. - The API validated one of those counts. - The storage layer enforced a **byte** limit. Any mismatch between those three shows up only for non-ASCII input, which means it survives every review and every test until a real user with an accented surname or an emoji in their display name arrives. Note also that the failure is asymmetric: a byte-limited store silently accepts everything ASCII and rejects only the users whose names carry the most cultural weight, so the bug reads as discriminatory even though it is arithmetic. ## What to do about it - **Name the unit in the contract.** "At most 100 code points" and "at most 400 bytes" are both defensible; "at most 100 characters" is not, because it does not say which count. - **Size storage for the worst case.** If the contract is code points, budget four bytes per code point; if you truncate on the byte side, you also need a boundary-safe cut, or you corrupt the last character. - **Validate at one layer and derive the others.** Two independent limits drift apart the first time someone edits only one of them. - **Test with non-ASCII fixtures.** A single fixture containing an emoji, a combining accent and a CJK character catches nearly every instance of this family of bugs, and costs one line. ## Why encodings differ at all Unicode assigns numbers to characters; an encoding decides how those numbers become bytes. Fixed-width encodings make counting trivial and waste space; variable-width encodings are compact and make counting a scan. Mainstream ecosystems chose differently — some store text as UTF-8 bytes, others as UTF-16 code units — so the same string can honestly report different lengths in different systems without either being wrong. The number is only meaningful once you state its unit.

  • For a page of CJK text, which is more compact, UTF-8 or UTF-16?
    UTF-16, usually. Most CJK code points sit in the Basic Multilingual Plane, costing two bytes in UTF-16 but three in UTF-8. UTF-8 wins decisively on ASCII-heavy text such as markup, source code and JSON keys, which is why mixed documents — CJK prose inside ASCII tags — often still favour UTF-8 overall. Measure your real corpus rather than reasoning from the script alone.
  • How would you size storage for names of at most 50 characters?
    First decide what "character" means in the contract. If it means code points, the worst case is four bytes each, so 200 bytes covers everything. If the store counts bytes and the API counts code points, enforce the code-point limit in the API and make the byte limit strictly larger, so the store never rejects input the API accepted.
  • Is ASCII text larger when stored as UTF-8?
    No. UTF-8 was designed so that every code point in the ASCII range encodes to the identical single byte. Pure ASCII text is byte-for-byte the same in both, which is why UTF-8 could be adopted without rewriting existing files or protocols. The cost of the rest of Unicode is paid only by text that actually uses it.

saying these in an interview costs you the question

  • Says one character always equals one byte
  • Assumes every non-ASCII character costs exactly two bytes
  • Treats a byte limit and a character limit as the same check
  • Thinks emoji are ordinary two-byte characters
  • Believes ASCII text grows when encoded as UTF-8

context