A Kotlin Char is a single 16-bit unit. What problems arise with characters outside the Basic Multilingual Plane (e.g. emoji), and how do String.length and indexing behave for them?
answer
- Char = 16-bit UTF-16 unit
- Supplementary char = surrogate pair (2 Chars)
- length counts code units → emoji = 2
- s[i] can be a lone surrogate
- codePointAt / codePointCount / codePoints() for real code points
basics
~20 sA Kotlin Char only holds 16 bits, but some characters (like many emoji) need more. Those are stored as two Chars (a 'surrogate pair'), so String.length counts them as 2 and a single index gives you only half a character.
solid answer
~40 sKotlin's Char is a UTF-16 code unit (16 bits), matching the JVM. Characters in the Basic Multilingual Plane (code points U+0000..U+FFFF) fit in one Char, but supplementary characters (U+10000 and above — many emoji, rare CJK, math symbols) are encoded as a surrogate pair: two Chars, a high surrogate (U+D800..U+DBFF) and a low surrogate (U+DC00..U+DFFF). Consequences: String.length returns the number of UTF-16 units, so an emoji contributes 2; s[i] may return a lone, meaningless surrogate Char; iterating with for (ch in s) walks half-characters. To work in real Unicode code points use codePointAt/codePointCount/offsetByCodePoints (JVM CharSequence APIs) or s.codePoints() and Character.charCount/toChars. Grapheme clusters (emoji with skin-tone/ZWJ sequences) are yet another level above code points. So 'length' is rarely the user-visible character count.
code
kotlin · 10 linesval s = "a\uD83D\uDE00b" // a 😀 b
println(s.length) // 4
println(s.codePointCount(0, s.length)) // 3
println(Character.charCount(0x1F600)) // 2 (surrogate pair)
// safe per-code-point walk:
var i = 0
while (i < s.length) {
val cp = s.codePointAt(i)
i += Character.charCount(cp)
}go deeper
Aware that length can differ from what you'd count by eye for emoji.
Explains surrogate pairs and that length counts UTF-16 units, so emoji count as 2.
Uses codePointAt/codePointCount/codePoints and writes a safe per-code-point walk avoiding lone surrogates.
Distinguishes code units vs code points vs graphemes, knows BreakIterator for grapheme-aware limits, and designs APIs/storage around it.
## Char = one UTF-16 code unit Kotlin `Char` is 16 bits, mirroring the JVM `char`. UTF-16 encodes Unicode like this: - **BMP** (Basic Multilingual Plane, U+0000..U+FFFF): one `Char` each. Covers Latin, most CJK, etc. - **Supplementary** (U+10000..U+10FFFF): cannot fit in 16 bits, so they're stored as a **surrogate pair** — two `Char` values: a *high surrogate* (U+D800..U+DBFF) followed by a *low surrogate* (U+DC00..U+DFFF). Many emoji, some historic scripts, and math alphanumerics live here. ## Why `length` and `[]` surprise you ```kotlin val s = "a\uD83D\uDE00b" // 'a' + 😀 + 'b' println(s.length) // 4, not 3 (emoji = 2 units) println(s[1]) // a lone high surrogate, not '😀' for (ch in s) print("[$ch]") // [a][?][?][b] ``` - `String.length` = number of UTF-16 **code units**, so a single emoji counts as 2. - `s[i]` returns one code unit; for a supplementary char that's a **lone surrogate**, which is not a valid standalone character. - `for (ch in s)` / `s.forEach` iterate code units, splitting supplementary characters. ## Working at the code-point level Use the JVM `Character`/`CharSequence` APIs (available on Kotlin String): ```kotlin val cp = s.codePointAt(1) // full code point of the emoji val count = s.codePointCount(0, s.length) // 3 — real characters s.codePoints().forEach { /* Int code points */ } val chars = Character.toChars(0x1F600) // CharArray of size 2 val n = Character.charCount(0x1F600) // 2 ``` `offsetByCodePoints(index, n)` lets you step by whole code points instead of units. ## Even code points aren't 'characters' to a user A **grapheme cluster** — what a human perceives as one symbol — can span several code points: e.g. an emoji with a skin-tone modifier, or a family emoji joined by Zero-Width Joiners (ZWJ, U+200D). Counting those correctly needs a grapheme/locale-aware library (`java.text.BreakIterator`), not `length` or even `codePointCount`. ## Practical guidance - Don't use `s.length` as a user-facing character count for arbitrary text. - Don't index/truncate arbitrary strings by `Char` — you can split a surrogate pair and produce mojibake. - For validation/limits on user input with emoji, count code points (or graphemes) deliberately. ## Key takeaways - `Char` is a UTF-16 unit; supplementary chars = 2 Chars (surrogate pair). - `length`/`[]`/iteration are code-unit level → emoji break naively. - Code-point APIs: `codePointAt`, `codePointCount`, `codePoints()`, `Character.charCount/toChars`. - Graphemes (ZWJ/modifiers) are a further layer above code points.
- Why might truncating a String to 10 chars corrupt the text?The 10th index could fall between the two surrogates of one emoji, leaving a lone surrogate and producing a broken/replacement character.
- What's the difference between a code point and a grapheme cluster?A code point is one Unicode scalar (an Int); a grapheme is what a user sees as one symbol and may combine several code points (e.g. emoji + skin tone via ZWJ).
Char is like a syllable, not a word: some 'characters' (emoji) take two syllables, so counting syllables (length) overcounts the words a person sees.
saying these in an interview costs you the question
- Assuming String.length equals the visible character count
- Indexing/truncating arbitrary text by Char without surrogate awareness
- Believing every Char is a complete character
- Not knowing codePointAt/codePointCount exist
- Conflating code points with grapheme clusters