skip to content

Why does `'๐Ÿ™‚'.length` evaluate to 2 in JavaScript, and how do you count what a user would call one character?

level: middleimportance: must knowfreq 58%

answer

  1. length is not characters
  2. sixteen bits is not enough
  3. two reserved ranges, one code point
  4. spread iterates, split slices
  5. index methods can cut a pair

basics

~20 s

JavaScript's .length counts UTF-16 code units, not characters. The ๐Ÿ™‚ emoji sits above the Basic Multilingual Plane, so it is stored as a two-unit surrogate pair. Iterating the string, for example with [...str], walks whole code points and yields 1.

solid answer

~40 s

A JavaScript string is a sequence of UTF-16 code units, and `.length` reports how many code units it holds. Code points up to U+FFFF fit in one unit; anything above that โ€” emoji, many CJK extensions, historic scripts โ€” is encoded as a *surrogate pair*, a high unit in U+D800โ€“U+DBFF followed by a low unit in U+DC00โ€“U+DFFF. `'๐Ÿ™‚'` is U+1F642, so it costs two units and `.length` is 2. To count code points instead, use the string iterator: `[...'๐Ÿ™‚'].length` and `Array.from('๐Ÿ™‚').length` are both 1, because iteration is code-point aware (ES2015). Likewise `codePointAt(0)` returns the full 128578 while `charCodeAt(0)` returns only the high surrogate 55357. Anything index-based โ€” `slice`, `at`, `split('')` โ€” works in code units and can cut a pair in half.

code

javascript ยท 10 lines
javascript
const text = 'a๐Ÿ™‚b';

console.log(text.length);             // 4 โ€” UTF-16 code units
console.log([...text].length);        // 3 โ€” code points
console.log(Array.from(text).length); // 3
console.log(text.split('').length);   // 4 โ€” split('') is unit-based

console.log(text.charCodeAt(1));      // 55357 โ€” high surrogate only
console.log(text.codePointAt(1));     // 128578 โ€” the whole code point
console.log(String.fromCodePoint(128578) === '๐Ÿ™‚'); // true

go deeper

for a junior

Know that .length counts UTF-16 code units, which is why an emoji can report 2, and that spreading with [...str] or using Array.from gives a code-point count instead.

for a middle

Explain surrogate pairs concretely: code points above U+FFFF are stored as a high unit in D800โ€“DBFF plus a low unit in DC00โ€“DFFF, which is exactly why charCodeAt returns half of one and codePointAt returns the whole value.

for a senior

Demonstrate the downstream damage: slice, at(-1) and split('') can leave an unpaired surrogate that renders as a replacement character and makes encodeURIComponent throw, and show how you detect it with isWellFormed rather than by hand-checking surrogate ranges.

for a principal

Be able to justify why the language is stuck with UTF-16 โ€” indexes and length are observable, so the representation cannot change โ€” and argue for treating user text as opaque in most of the system, measuring it only at the few boundaries that genuinely need a count.

## Strings are sequences of UTF-16 code units JavaScript fixed its string representation early: a string is an ordered sequence of 16-bit values called *code units*, interpreted as UTF-16. That decision is baked into `length`, into every index, and into every method that takes a position. It is not "a sequence of characters", and the difference becomes visible the moment your text leaves the ASCII range. Unicode assigns each character a *code point*, a number from U+0000 to U+10FFFF. Code points up to U+FFFF โ€” the Basic Multilingual Plane, which covers Latin, Cyrillic, Greek, Arabic, Hebrew, and common CJK โ€” fit in a single 16-bit unit, so for that text code units and code points line up and nobody notices the distinction. ## Surrogate pairs Above U+FFFF, one 16-bit unit is not enough. UTF-16 encodes such a code point as two units: a *high surrogate* in the range U+D800โ€“U+DBFF followed by a *low surrogate* in U+DC00โ€“U+DFFF. Those two ranges are permanently reserved, so a pair is unambiguous. `'๐Ÿ™‚'` is U+1F642, well above the BMP, so it is stored as the pair U+D83D U+DE42 โ€” two code units, one code point: ```javascript const face = '๐Ÿ™‚'; console.log(face.length); // 2 โ€” code units console.log(face.charCodeAt(0)); // 55357 (0xD83D) โ€” high surrogate only console.log(face.charCodeAt(1)); // 56898 (0xDE42) โ€” low surrogate only console.log(face.codePointAt(0)); // 128578 (0x1F642) โ€” the whole code point ``` `codePointAt` (ES2015) reads a full code point when the index lands on a high surrogate that is followed by a low one; `charCodeAt` always returns exactly one unit. The inverse pair is `String.fromCharCode` (units) versus `String.fromCodePoint` (code points, ES2015): `String.fromCodePoint(0x1F642) === '๐Ÿ™‚'`. ## Counting code points The string iterator โ€” added in ES2015 and used by spread, `Array.from`, and `for...of` โ€” is code-point aware. It yields a whole surrogate pair as one string: ```javascript const text = 'a๐Ÿ™‚b'; console.log(text.length); // 4 code units console.log([...text].length); // 3 code points console.log(Array.from(text).length); // 3 console.log(text.split('').length); // 4 โ€” split('') is unit-based, not iteration for (const ch of text) console.log(ch); // "a", "๐Ÿ™‚", "b" ``` The `split('')` case is the one that catches people: it looks like the same thing as spreading, but it slices by code unit and tears the emoji in half. ## Where the distinction bites Every index-taking operation is in code units: ```javascript console.log('๐Ÿ™‚'.slice(0, 1)); // an unpaired high surrogate โ€” usually renders as "๏ฟฝ" console.log('a๐Ÿ™‚'.at(-1)); // the low surrogate alone, not the emoji (String.prototype.at is ES2022) console.log([...'a๐Ÿ™‚'].reverse().join('')); // "๐Ÿ™‚a" โ€” correct, because spreading was code-point aware ``` An *unpaired surrogate* is a valid JavaScript string but not valid Unicode text. It typically renders as the replacement character U+FFFD, and `encodeURIComponent` throws a `URIError` on it, so a naive truncation in the browser can turn into a request-building failure two layers away. Since ES2024, `String.prototype.isWellFormed()` and `toWellFormed()` let you detect and repair such strings without hand-rolling surrogate-range checks. Regular expressions have the same split: without the `u` flag, `.` matches one code unit, so `'๐Ÿ™‚'.match(/./gu).length` is 1 while `'๐Ÿ™‚'.match(/./g).length` is 2. ## Code points are still not "characters" Counting code points fixes emoji like `'๐Ÿ™‚'`, but a single thing the user perceives as one character can be several code points: a base letter plus combining marks, a flag built from two regional indicators, or an emoji joined with zero-width joiners. Those user-perceived units are called grapheme clusters, and counting them requires segmentation rather than iteration โ€” worth naming in an interview so you show you know `[...str].length` is a *better* count, not a final one. ## How to answer crisply Say: `.length` is UTF-16 code units; U+1F642 needs a surrogate pair, hence 2; use spread or `Array.from` for code points, `codePointAt` instead of `charCodeAt`; and remember index-based methods can split a pair and produce an unpaired surrogate.

  • What is the difference between charCodeAt and codePointAt?
    `charCodeAt(i)` always returns the single UTF-16 code unit at that index, so on an astral character it hands back a bare surrogate. `codePointAt(i)` returns the full code point when the index lands on a high surrogate followed by a low one, and otherwise behaves like `charCodeAt`. Their inverses are `String.fromCharCode` and `String.fromCodePoint` respectively.
  • What actually happens if you slice a string in the middle of a surrogate pair?
    You get a valid JavaScript string containing an unpaired surrogate, which is not valid Unicode text. It usually renders as the replacement character, comparisons against properly formed text fail, and `encodeURIComponent` throws a `URIError`. Since ES2024 you can detect it with `isWellFormed()` and repair it with `toWellFormed()`.
  • Is counting code points enough to count what a user calls a character?
    No. A base letter plus a combining accent, a regional-indicator flag, or a zero-width-joiner emoji sequence are each several code points but one user-perceived character. Code-point counting is a strict improvement over `.length`, but only grapheme-cluster segmentation matches human intuition.

saying these in an interview costs you the question

  • Says length counts characters or letters
  • Assumes every emoji has length 2
  • Treats split('') as equivalent to spreading the string
  • Uses charCodeAt and expects the full emoji code point
  • Claims [...str].length always equals the visible character count

context