Why does `'๐'.length` evaluate to 2 in JavaScript, and how do you count what a user would call one character?
answer
- length is not characters
- sixteen bits is not enough
- two reserved ranges, one code point
- spread iterates, split slices
- index methods can cut a pair
basics
~20 sJavaScript's .length counts UTF-16 code units, not characters. The ๐ emoji sits above the Basic Multilingual Plane, so it is stored as a two-unit surrogate pair. Iterating the string, for example with [...str], walks whole code points and yields 1.
solid answer
~40 sA JavaScript string is a sequence of UTF-16 code units, and `.length` reports how many code units it holds. Code points up to U+FFFF fit in one unit; anything above that โ emoji, many CJK extensions, historic scripts โ is encoded as a *surrogate pair*, a high unit in U+D800โU+DBFF followed by a low unit in U+DC00โU+DFFF. `'๐'` is U+1F642, so it costs two units and `.length` is 2. To count code points instead, use the string iterator: `[...'๐'].length` and `Array.from('๐').length` are both 1, because iteration is code-point aware (ES2015). Likewise `codePointAt(0)` returns the full 128578 while `charCodeAt(0)` returns only the high surrogate 55357. Anything index-based โ `slice`, `at`, `split('')` โ works in code units and can cut a pair in half.
code
javascript ยท 10 linesconst text = 'a๐b';
console.log(text.length); // 4 โ UTF-16 code units
console.log([...text].length); // 3 โ code points
console.log(Array.from(text).length); // 3
console.log(text.split('').length); // 4 โ split('') is unit-based
console.log(text.charCodeAt(1)); // 55357 โ high surrogate only
console.log(text.codePointAt(1)); // 128578 โ the whole code point
console.log(String.fromCodePoint(128578) === '๐'); // truego deeper
Know that .length counts UTF-16 code units, which is why an emoji can report 2, and that spreading with [...str] or using Array.from gives a code-point count instead.
Explain surrogate pairs concretely: code points above U+FFFF are stored as a high unit in D800โDBFF plus a low unit in DC00โDFFF, which is exactly why charCodeAt returns half of one and codePointAt returns the whole value.
Demonstrate the downstream damage: slice, at(-1) and split('') can leave an unpaired surrogate that renders as a replacement character and makes encodeURIComponent throw, and show how you detect it with isWellFormed rather than by hand-checking surrogate ranges.
Be able to justify why the language is stuck with UTF-16 โ indexes and length are observable, so the representation cannot change โ and argue for treating user text as opaque in most of the system, measuring it only at the few boundaries that genuinely need a count.
## Strings are sequences of UTF-16 code units JavaScript fixed its string representation early: a string is an ordered sequence of 16-bit values called *code units*, interpreted as UTF-16. That decision is baked into `length`, into every index, and into every method that takes a position. It is not "a sequence of characters", and the difference becomes visible the moment your text leaves the ASCII range. Unicode assigns each character a *code point*, a number from U+0000 to U+10FFFF. Code points up to U+FFFF โ the Basic Multilingual Plane, which covers Latin, Cyrillic, Greek, Arabic, Hebrew, and common CJK โ fit in a single 16-bit unit, so for that text code units and code points line up and nobody notices the distinction. ## Surrogate pairs Above U+FFFF, one 16-bit unit is not enough. UTF-16 encodes such a code point as two units: a *high surrogate* in the range U+D800โU+DBFF followed by a *low surrogate* in U+DC00โU+DFFF. Those two ranges are permanently reserved, so a pair is unambiguous. `'๐'` is U+1F642, well above the BMP, so it is stored as the pair U+D83D U+DE42 โ two code units, one code point: ```javascript const face = '๐'; console.log(face.length); // 2 โ code units console.log(face.charCodeAt(0)); // 55357 (0xD83D) โ high surrogate only console.log(face.charCodeAt(1)); // 56898 (0xDE42) โ low surrogate only console.log(face.codePointAt(0)); // 128578 (0x1F642) โ the whole code point ``` `codePointAt` (ES2015) reads a full code point when the index lands on a high surrogate that is followed by a low one; `charCodeAt` always returns exactly one unit. The inverse pair is `String.fromCharCode` (units) versus `String.fromCodePoint` (code points, ES2015): `String.fromCodePoint(0x1F642) === '๐'`. ## Counting code points The string iterator โ added in ES2015 and used by spread, `Array.from`, and `for...of` โ is code-point aware. It yields a whole surrogate pair as one string: ```javascript const text = 'a๐b'; console.log(text.length); // 4 code units console.log([...text].length); // 3 code points console.log(Array.from(text).length); // 3 console.log(text.split('').length); // 4 โ split('') is unit-based, not iteration for (const ch of text) console.log(ch); // "a", "๐", "b" ``` The `split('')` case is the one that catches people: it looks like the same thing as spreading, but it slices by code unit and tears the emoji in half. ## Where the distinction bites Every index-taking operation is in code units: ```javascript console.log('๐'.slice(0, 1)); // an unpaired high surrogate โ usually renders as "๏ฟฝ" console.log('a๐'.at(-1)); // the low surrogate alone, not the emoji (String.prototype.at is ES2022) console.log([...'a๐'].reverse().join('')); // "๐a" โ correct, because spreading was code-point aware ``` An *unpaired surrogate* is a valid JavaScript string but not valid Unicode text. It typically renders as the replacement character U+FFFD, and `encodeURIComponent` throws a `URIError` on it, so a naive truncation in the browser can turn into a request-building failure two layers away. Since ES2024, `String.prototype.isWellFormed()` and `toWellFormed()` let you detect and repair such strings without hand-rolling surrogate-range checks. Regular expressions have the same split: without the `u` flag, `.` matches one code unit, so `'๐'.match(/./gu).length` is 1 while `'๐'.match(/./g).length` is 2. ## Code points are still not "characters" Counting code points fixes emoji like `'๐'`, but a single thing the user perceives as one character can be several code points: a base letter plus combining marks, a flag built from two regional indicators, or an emoji joined with zero-width joiners. Those user-perceived units are called grapheme clusters, and counting them requires segmentation rather than iteration โ worth naming in an interview so you show you know `[...str].length` is a *better* count, not a final one. ## How to answer crisply Say: `.length` is UTF-16 code units; U+1F642 needs a surrogate pair, hence 2; use spread or `Array.from` for code points, `codePointAt` instead of `charCodeAt`; and remember index-based methods can split a pair and produce an unpaired surrogate.
- What is the difference between charCodeAt and codePointAt?`charCodeAt(i)` always returns the single UTF-16 code unit at that index, so on an astral character it hands back a bare surrogate. `codePointAt(i)` returns the full code point when the index lands on a high surrogate followed by a low one, and otherwise behaves like `charCodeAt`. Their inverses are `String.fromCharCode` and `String.fromCodePoint` respectively.
- What actually happens if you slice a string in the middle of a surrogate pair?You get a valid JavaScript string containing an unpaired surrogate, which is not valid Unicode text. It usually renders as the replacement character, comparisons against properly formed text fail, and `encodeURIComponent` throws a `URIError`. Since ES2024 you can detect it with `isWellFormed()` and repair it with `toWellFormed()`.
- Is counting code points enough to count what a user calls a character?No. A base letter plus a combining accent, a regional-indicator flag, or a zero-width-joiner emoji sequence are each several code points but one user-perceived character. Code-point counting is a strict improvement over `.length`, but only grapheme-cluster segmentation matches human intuition.
saying these in an interview costs you the question
- Says length counts characters or letters
- Assumes every emoji has length 2
- Treats split('') as equivalent to spreading the string
- Uses charCodeAt and expects the full emoji code point
- Claims [...str].length always equals the visible character count