skip to content

What is the difference between String.chars() and String.codePoints(), and why does it matter for characters like emoji?

level: seniorimportance: should knowfreq 40%

answer

  1. chars() = one int per UTF-16 code unit (== length())
  2. codePoints() = one int per Unicode code point (combines surrogate pairs)
  3. emoji/supplementary char = 2 code units, 1 code point
  4. both return IntStream because code points can exceed 0xFFFF
  5. grapheme clusters need BreakIterator, not codePoints

basics

~20 s

chars() gives a stream of UTF-16 code units, so a character stored as two units (like an emoji) shows up as two values. codePoints() gives a stream of full Unicode characters, so each emoji is one value.

solid answer

~50 s

Both String.chars() and String.codePoints() return an IntStream of int values, but they iterate at different granularities. chars() yields one int per UTF-16 code unit — exactly the same as String.length() counts — so a supplementary character (above U+FFFF, like most emoji) appears as two ints: its high and low surrogate. codePoints() decodes surrogate pairs and yields one int per actual Unicode code point, so each emoji is a single value. They return int (not char) precisely because a code point can exceed 0xFFFF, which a char cannot hold. Use codePoints() when counting or processing real characters (e.g. correct length, filtering, mapping), and chars() only when you genuinely work at the UTF-16 level. To turn code points back into text use StringBuilder.appendCodePoint or Character.toChars. Neither handles grapheme clusters (combining marks, flag/skin-tone sequences) — that needs java.text.BreakIterator.

code

java · 10 lines
java
String s = "a😀b"; // 'a', grinning-face emoji, 'b'
s.length();              // 4  (UTF-16 code units)
s.chars().count();       // 4
s.codePoints().count();  // 3  (a, emoji, b)

// rebuild text from code points (upper-cased):
String upper = s.codePoints()
    .map(Character::toUpperCase)
    .collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append)
    .toString();

go deeper

for a junior

Knows chars() iterates the individual char units and codePoints() iterates whole characters, and that both produce a stream of ints.

for a middle

Can explain that an emoji is two code units but one code point, so chars() over-counts compared to codePoints().

for a senior

Explains UTF-16 surrogate pairs, why the streams are IntStream, when to use each, and how to rebuild strings with appendCodePoint/Character.toChars.

for a principal

Sets text-handling policy across services (code-point vs grapheme correctness, BreakIterator for visual segmentation, encoding pitfalls in storage/transport) and reviews for these bugs.

## Background: how Java stores text Java strings are sequences of **`char`** values, and a `char` is a **16-bit UTF-16 code unit**. Unicode assigns each character a **code point** — an integer identity, e.g. 'A' is U+0041, the grinning-face emoji is U+1F600. Code points up to U+FFFF (the 'Basic Multilingual Plane', BMP) fit in one 16-bit char. Code points above that (**supplementary characters**, U+10000–U+10FFFF, including most emoji and many rare scripts) do **not** fit in 16 bits, so UTF-16 encodes them as a **surrogate pair**: two special char values, a 'high surrogate' (U+D800–DBFF) followed by a 'low surrogate' (U+DC00–DFFF). Together they represent one code point. This is why `"😀".length()` is **2**, not 1 — `length()` counts char/code units, not characters. ## chars() `chars()` returns an `IntStream` (a primitive stream of `int`) with **one element per char (UTF-16 code unit)**. Its count equals `length()`. For a supplementary character you get **two** ints — the surrogate values — each meaningless on its own. It returns `int` rather than `char` simply because `IntStream` is the primitive stream type Java offers; the values are still 16-bit char values widened to int. Example: `"a😀b".chars().count()` → 4 (a, high surrogate, low surrogate, b). ## codePoints() `codePoints()` returns an `IntStream` with **one element per Unicode code point**: it scans the string, and whenever it sees a valid surrogate pair it **combines** the two chars into the single code point integer. So each emoji is exactly one int, and that int can be larger than 0xFFFF — which is precisely why the stream is of `int`, not `char` (a char cannot hold a supplementary code point). Example: `"a😀b".codePoints().count()` → 3 (a, 😀, b). This is the correct way to count 'characters' for most purposes. ## Why it matters Using `chars()` or `length()` to count or slice user-visible text breaks on emoji and many non-Latin scripts: you over-count, and slicing can cut a surrogate pair in half, producing an invalid lone surrogate that renders as garbage. `codePoints()` (and code-point-aware methods like `codePointCount`, `offsetByCodePoints`) avoid that class of bug. ## Reconstructing strings from an int stream Since elements are ints, you cannot just `Stream<Character>`-collect them. To rebuild text from code points use, e.g., `codePoints().collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append).toString()`, or `Character.toChars(cp)` to get the char(s) for one code point. `appendCodePoint` correctly emits one or two chars as needed. ## The remaining limit: grapheme clusters Even a code point is not always a single **user-perceived character** (a 'grapheme cluster'). A base letter plus a combining accent, an emoji with a skin-tone modifier, or a flag (two regional-indicator code points) are multiple code points that display as one symbol. Neither `chars()` nor `codePoints()` groups these. For true visual-character segmentation use `java.text.BreakIterator.getCharacterInstance()`, which iterates grapheme boundaries. ## Practical guidance - Counting/processing characters → `codePoints()`. - Working at the encoding/buffer level → `chars()`. - Visual character count, cursor movement, truncation that must not split emoji → `BreakIterator`. - Both streams are lazy `IntStream`s, so you can `filter`, `map`, `mapToObj`, etc.

  • What do "a😀b".chars().count() and codePoints().count() each return, and why?
    chars().count() is 4 — the emoji is a surrogate pair of two code units, plus 'a' and 'b'. codePoints().count() is 3 — the surrogate pair is decoded into the single emoji code point.
  • If codePoints() handles emoji, why might you still need BreakIterator?
    A user-perceived character (grapheme cluster) can be several code points: a base letter plus combining accents, an emoji with a skin-tone modifier, or a flag built from two regional indicators. codePoints() counts each separately; BreakIterator groups them into one visual character.

saying these in an interview costs you the question

  • Believing chars().count() equals the number of visible characters
  • Slicing/length on UTF-16 units and corrupting emoji
  • Thinking codePoints() handles grapheme clusters (accents, flags, skin tones)
  • Assuming chars() returns char values you can collect directly
  • Using length() as a character count for international text

context