skip to content

Truncating user-supplied text keeps producing broken emoji and stray replacement characters in production. Working in JavaScript, how do you cut a string safely, and what unit should you cut on?

level: seniorimportance: should knowfreq 30%

answer

  1. three units, one of them human
  2. surrogate pair is the first trap
  3. joiners and combining marks are the second
  4. segmentation, not iteration
  5. cap bytes separately anyway

basics

~20 s

Cut on grapheme clusters, not code units or code points. Use Intl.Segmenter with granularity 'grapheme' to walk user-perceived characters and join the first N. Slicing by index splits surrogate pairs, and spreading splits joined emoji and combining marks.

solid answer

~50 s

There are three candidate units and only one of them matches what a user sees. `slice` and `substring` work in UTF-16 code units, so they can cut a surrogate pair in half and leave an unpaired surrogate that renders as `�` and makes `encodeURIComponent` throw. Spreading with `[...str]` fixes that but still cuts inside a grapheme cluster — a zero-width-joiner family emoji, a regional-indicator flag, a skin-tone modifier, or a letter plus a combining accent are all several code points that display as one character. The correct unit is the grapheme cluster, which you get from `Intl.Segmenter` with `granularity: 'grapheme'`: iterate its segments and join the first N. Normalise to NFC first so the same visible text measures the same way, and enforce a separate hard byte cap so a single grapheme padded with combining marks cannot blow up storage.

code

javascript · 16 lines
javascript
const graphemes = new Intl.Segmenter('en', { granularity: 'grapheme' });

function truncate(str, maxGraphemes) {
  const out = [];
  for (const { segment } of graphemes.segment(str.normalize('NFC'))) {
    if (out.length === maxGraphemes) break;
    out.push(segment);
  }
  return out.join('');
}

const text = 'ok 👨‍👩‍👧‍👦!';
console.log(text.length);                     // 15 code units
console.log([...text].length);                // 11 code points
console.log([...graphemes.segment(text)].length); // 5 grapheme clusters
console.log(truncate(text, 4));               // "ok 👨‍👩‍👧‍👦"

go deeper

for a junior

Know that cutting a string with slice can split an emoji and leave a broken character, and that iterating with spread is safer than indexing by number.

for a middle

Explain the three units side by side — code units, code points, grapheme clusters — and show with a concrete emoji why each successive unit fixes one class of breakage and leaves another.

for a senior

Bring the production judgment: reach for Intl.Segmenter with grapheme granularity, construct it once outside the hot path, normalise before measuring, and name the downstream symptom you were chasing, such as encodeURIComponent throwing on an unpaired surrogate.

for a principal

Own the policy question — where truncation is allowed at all, whether input is rejected rather than silently shortened, and how a user-facing grapheme limit is paired with a hard byte cap so unbounded combining marks cannot become a resource problem.

## Three units, three different answers Ask "how long is this string" and JavaScript can give you three legitimate numbers: ```javascript const text = 'ok 👨‍👩‍👧‍👦!'; console.log(text.length); // 15 — UTF-16 code units console.log([...text].length); // 11 — code points // grapheme clusters: 5 — "o", "k", " ", the family emoji, "!" ``` The family emoji is four people joined by three zero-width joiners (U+200D): seven code points, eleven code units, one thing the user sees. Any truncation rule has to pick a unit, and picking the wrong one is what produces the bug in the question. ## Why `slice` is the worst choice `slice`, `substring` and `substr` index by code unit. Cutting between the halves of a surrogate pair leaves an *unpaired surrogate*: a legal JavaScript string that is not legal Unicode text. It usually renders as the replacement character U+FFFD, it breaks equality against the properly formed original, and `encodeURIComponent` throws a `URIError` on it — which is how a display-layer truncation turns into a failed request further down the stack. ```javascript const half = '🙂'.slice(0, 1); console.log(half.length); // 1 — an unpaired high surrogate ``` Since ES2024 you can at least detect the damage: `half.isWellFormed()` is `false`, and `half.toWellFormed()` replaces the lone surrogate with U+FFFD. That is a repair, not a fix — better not to create the problem. ## Why code points are still not enough Spreading fixes surrogate pairs, and for a lot of text that is where people stop. But plenty of single user-perceived characters are multi-code-point sequences: - **combining marks** — `'é'` is `e` plus a combining acute, displayed as `é`; - **regional indicator flags** — `'🇺🇸'` is two regional indicator symbols, four code units, two code points; - **skin-tone modifiers** — a hand emoji plus a modifier; - **ZWJ sequences** — the family emoji above. Cut a code-point array in the middle of any of those and you get a headless combining accent, half a flag, or a stray joiner. So `[...str].slice(0, n).join('')` is an improvement, not a solution. ## Segmenting on grapheme clusters `Intl.Segmenter` (part of ECMA-402, widely available in modern browsers and in Node 16+) does the segmentation defined by the Unicode text-segmentation rules. With `granularity: 'grapheme'` it yields user-perceived characters: ```javascript const segmenter = new Intl.Segmenter('en', { granularity: 'grapheme' }); function truncate(str, maxGraphemes) { const out = []; for (const { segment } of segmenter.segment(str)) { if (out.length === maxGraphemes) break; out.push(segment); } return out.join(''); } ``` Each yielded object carries `segment` (the text), `index` (its starting code-unit offset) and `input`. Construct the segmenter **once** at module scope: creating one per call is a real cost in a hot path. `granularity` also accepts `'word'` and `'sentence'`, which is what you want if the product rule is "trim on a word boundary and add an ellipsis" rather than a hard character cap. ## The rest of a robust truncation **Normalise first.** Run `str.normalize('NFC')` before measuring so text that arrived decomposed does not count differently from visually identical text that arrived precomposed. **Keep a separate byte ceiling.** Grapheme clusters are unbounded: a single base letter can carry hundreds of combining marks, so "at most 30 graphemes" is not a bound on storage or bandwidth. Pair the user-facing limit with a generous hard cap in code units or UTF-8 bytes, and reject beyond it rather than truncating. **Decide reject vs truncate.** Silently truncating a name is a data-quality problem; rejecting with a clear message is usually better for input, while truncation belongs in display code. **Do not confuse count with width.** Grapheme count is not rendered width — CJK characters are typically double-width and emoji wider still. If the requirement is "fits on one line", that is a layout problem, not a string-length problem. ## What the interviewer is listening for The strong answer names all three units, explains what breaks at each level with a concrete example, reaches for `Intl.Segmenter` rather than a hand-rolled surrogate-range check, and then adds the operational caveats: normalise before measuring, cap bytes separately, and be explicit about reject versus truncate.

  • Why is `[...str].slice(0, n).join('')` not a safe truncation?
    It cuts on code points, which handles surrogate pairs but not grapheme clusters. A zero-width-joiner emoji, a regional-indicator flag, a skin-tone modifier and a letter-plus-combining-accent are each several code points rendered as one character, so the cut can leave a bare joiner, half a flag, or an orphaned combining mark.
  • You cap display names at 30 graphemes. Why is that not a bound on storage?
    Because a single grapheme cluster can be arbitrarily long: one base letter may carry an unlimited run of combining marks. Thirty graphemes can therefore be many kilobytes. Enforce the user-facing limit in graphemes for fairness, and a separate hard cap in code units or UTF-8 bytes to bound resources, rejecting anything past it.
  • What would you do differently if the requirement were 'truncate at a word boundary'?
    Use the same API with `granularity: 'word'`. Segmenter yields word-like and non-word segments, with `isWordLike` on each, so you can accumulate until you would exceed the budget and stop at the last complete word. Locale matters here far more than for graphemes, since word breaking differs by language.
  • How do you detect that some upstream code already broke a string?
    Since ES2024, `String.prototype.isWellFormed()` returns false when the string contains an unpaired surrogate, and `toWellFormed()` replaces each one with U+FFFD. Checking at an ingest boundary tells you the damage happened before you, which is far easier to act on than debugging a URIError from `encodeURIComponent` deep in a request builder.

saying these in an interview costs you the question

  • Uses slice or substring on user text and calls it fixed
  • Thinks spreading the string solves every truncation case
  • Hand-rolls surrogate-range checks instead of segmenting
  • Assumes grapheme count bounds storage size
  • Equates character count with rendered width

context