skip to content

Given the UTF-16-based String API, what subtle bugs arise from char-based operations on text with supplementary characters, and how do you handle them correctly?

level: seniorimportance: should knowfreq 35%

answer

  1. char = 1 UTF-16 code unit, not always 1 character
  2. supplementary char (>0xFFFF) = surrogate pair = 2 chars
  3. high surrogate D800-DBFF, low DC00-DFFF
  4. length()/charAt()/substring count units -> can split
  5. fix: codePoints(), codePointCount, offsetByCodePoints; BreakIterator for graphemes

basics

~20 s

A Java char is one UTF-16 unit, but some characters (like many emoji) need two chars (a surrogate pair). So length() and charAt() count units, not real characters, and slicing or counting by char can split a character in half. Use code-point methods or codePoints() for correctness.

solid answer

~50 s

Java's String API is defined in terms of UTF-16 code units: length() returns the number of units and charAt(i) returns one unit. Characters outside the Basic Multilingual Plane - many emoji, some CJK extensions, historic scripts - are encoded as a surrogate pair: two chars, a high surrogate (0xD800-0xDBFF) followed by a low surrogate (0xDC00-0xDFFF). Iterating with charAt or substring at arbitrary indices can therefore split a surrogate pair, producing broken or replacement glyphs, and length() over-counts such 'characters'. To be correct, count with codePointCount, iterate with codePoints() or by stepping with Character.charCount, and reverse/slice on code-point boundaries. Even code points are not the final word for user-perceived characters (grapheme clusters like flag emoji or accented combinations), which need a grapheme/BreakIterator approach. The internal byte[]+coder layout does not change any of this; the API contract is still UTF-16.

code

java · 13 lines
java
String s = "A😀B"; // 'A', grinning-face emoji, 'B'

System.out.println(s.length());                       // 4  (emoji = 2 units)
System.out.println(s.codePointCount(0, s.length()));  // 3  (real characters)

System.out.println(s.substring(1, 2));   // BROKEN: lone high surrogate
// Safe slice on a code-point boundary:
int end = s.offsetByCodePoints(1, 1);    // step one code point from index 1
System.out.println(s.substring(1, end)); // the whole emoji

// Iterate by real characters:
s.codePoints().forEach(cp ->
    System.out.println(new String(Character.toChars(cp))));

go deeper

for a junior

Aware that some characters (emoji) behave oddly - length can be 2 - and that there are special methods for them.

for a middle

Explains that a char is a UTF-16 code unit and supplementary characters use surrogate pairs, so length/charAt can mislead; reaches for codePoints().

for a senior

Knows the surrogate ranges, which String methods are unit- vs code-point-based, how to slice/iterate/count correctly, and that StringBuilder.reverse is surrogate-aware.

for a principal

Distinguishes code units, code points, and grapheme clusters; chooses BreakIterator/ICU for user-facing text, and sets team conventions for safe text handling across i18n boundaries.

## The root cause: the API speaks UTF-16 code units No matter how a String is stored internally (char[] then; byte[]+coder now), its **public API is defined in UTF-16 code units**. A `char` is exactly one 16-bit code unit. `length()` returns the number of code units, and `charAt(i)` returns the i-th code unit. This is fine for the **Basic Multilingual Plane (BMP)** - Unicode code points 0x0000-0xFFFF - where one code point equals one code unit. ## Terms, defined - **Code point**: one logical Unicode character, a number from 0 to 0x10FFFF. - **Code unit**: one storage slot of the encoding; in UTF-16 it is 16 bits. - **BMP**: the first 65,536 code points; each fits in a single UTF-16 code unit. - **Supplementary character**: a code point above 0xFFFF (emoji, rare CJK, ancient scripts). It cannot fit in one 16-bit unit. - **Surrogate pair**: how UTF-16 encodes a supplementary character - **two** code units: a *high surrogate* in 0xD800-0xDBFF followed by a *low surrogate* in 0xDC00-0xDFFF. Neither half is a valid character alone. - **Grapheme cluster**: what a human perceives as one character, which may be several code points (e.g. a base letter plus combining accents, or a flag emoji built from two regional-indicator code points). ## The bugs 1. **Wrong length**: `"😀".length()` (a grinning-face emoji) returns **2**, not 1, because it is a surrogate pair. 2. **Splitting a character**: `substring(0, 1)` on that string keeps only the high surrogate - a broken, unpaired unit that renders as a replacement box. 3. **Broken iteration**: a `for (i; i<length(); i++) charAt(i)` loop sees two meaningless halves instead of one emoji. 4. **Broken reverse**: naively reversing chars swaps the surrogates into an invalid order, corrupting the character. (`StringBuilder.reverse()` is actually surrogate-aware and avoids this; manual char swaps are not.) ## Doing it correctly - **Count real characters**: `s.codePointCount(0, s.length())` instead of `length()`. - **Iterate by code point**: `s.codePoints().forEach(...)` (an `IntStream` of code points), or step manually with `Character.charCount(cp)` to advance 1 or 2 units. - **Index by code point**: `s.offsetByCodePoints(start, n)` to find safe boundaries before `substring`. - **Read a full code point**: `s.codePointAt(i)` returns the combined code point when `i` is a high surrogate. - **User-perceived characters**: for grapheme clusters (flags, skin-tone emoji, accented combinations) even code points are too granular - use a grapheme segmentation (`BreakIterator.getCharacterInstance` or a library) to split on what users see as one character. ## Why the internal layout is irrelevant here Compact strings changed *storage*, not the *contract*. A surrogate pair still occupies two `char`s at the API level (and, in UTF16 storage mode, four bytes). So none of these pitfalls were introduced or fixed by Java 9; they follow from UTF-16 being the API's unit. ## How to derive an answer at any level Start from "char = one UTF-16 unit, but supplementary characters take two (a surrogate pair)," conclude "length/charAt/substring count units, so they can over-count or split a character," then "use codePointCount / codePoints() / offsetByCodePoints, and BreakIterator for graphemes." That reconstructs the whole answer.

  • Why might reversing a String by swapping chars corrupt emoji, yet StringBuilder.reverse() not?
    Swapping raw chars reorders a surrogate pair into low-then-high, which is invalid and renders broken. StringBuilder.reverse() is surrogate-aware: it detects surrogate pairs and keeps each pair's two units together while reversing, so emoji survive.
  • If code points fix the count, why might counting still be 'wrong' for users?
    Because a user-perceived character (grapheme cluster) can be several code points: a flag emoji is two regional-indicator code points, and accented letters can be base + combining marks. To count what users see, segment with BreakIterator.getCharacterInstance (or an ICU grapheme library), not by code point.

A char is like a single playing card; a supplementary character is a two-card pair that only means something together. Counting cards (length) over-counts hands, and dealing one card from a pair (charAt/substring) leaves you holding nonsense.

saying these in an interview costs you the question

  • Assuming length() returns the number of human-visible characters.
  • Slicing/iterating by char index over arbitrary text without surrogate handling.
  • Thinking compact strings (Java 9) introduced or fixed surrogate issues; they are an API/UTF-16 matter, not storage.
  • Believing one code point always equals one user-perceived character (grapheme clusters break that).
  • Manually reversing chars and expecting emoji to stay intact.

context