In Dart, why can a menu item name containing emoji have a String.length larger than its visible character count, and how do you count characters correctly?
answer
- UTF-16 code units
- surrogate pairs for emoji
- runes are code points
- grapheme clusters via characters
- substring can split a pair
basics
~20 sA Dart String is a sequence of UTF-16 code units, and length counts those units. Most emoji need two (a surrogate pair), and flags or family emoji combine several code points. Count user-perceived characters with the characters package, e.g. name.characters.length.
solid answer
~50 sDart strings store **UTF-16 code units**, and `length`, indexing and `substring` all work in code units. Characters outside the Basic Multilingual Plane, including most emoji, take a **surrogate pair** of two code units, so `'🍕'.length` is 2. `runes` combines surrogate pairs and yields **Unicode code points**, so `'🍕'.runes.length` is 1. Even code points are not what users see: a flag like 🇮🇹 is two regional-indicator code points, and family emoji join several code points with zero-width joiners. What users perceive as one character is a **grapheme cluster**, which the `characters` package exposes as `string.characters`; Flutter re-exports it from `package:flutter/widgets.dart`. So `'Pizza 🍕🇮🇹'` has length 12, 9 runes and 8 characters. Truncating with `substring` can cut a surrogate pair and leave a broken symbol; `characters.take(n)` does not. String `==` compares code units, so differently composed accents are not equal.
code
dart · 12 linesimport 'package:characters/characters.dart';
void main() {
const name = 'Pizza \u{1F355}\u{1F1EE}\u{1F1F9}'; // Pizza, pizza emoji, Italian flag
print(name.length); // 12: UTF-16 code units
print(name.runes.length); // 9: code points
print(name.characters.length); // 8: user-perceived characters
final broken = name.substring(0, 7); // cuts the pizza's surrogate pair
final safe = name.characters.take(7).join(); // 'Pizza ' plus the pizza emoji
print('${broken.length} ${safe.length}'); // 7 8
}go deeper
Recall that length counts UTF-16 code units, so emoji can count as two or more.
Explain code units, runes and grapheme clusters, and use characters for counting and truncating visible text.
Audit validation, truncation and comparison code that handles user-entered names so emoji and accents do not corrupt data.
Define text-handling rules for shared models, deciding where grapheme-aware counting and normalization are required.
## Three ways to count Text has several layers, and Dart exposes each: | Layer | Dart API | 'Pizza 🍕🇮🇹' | |---|---|---| | UTF-16 code units | `length`, `codeUnits`, `codeUnitAt`, `substring` | 12 | | Unicode code points (runes) | `runes` | 9 | | Grapheme clusters (what users see) | `characters` from package:characters | 8 | - **Code unit**: a 16-bit value. The SDK documents `String` as "a sequence of UTF-16 code units". - **Code point (rune)**: a Unicode scalar value. Code points above U+FFFF, such as U+1F355 🍕, are stored as a **surrogate pair** of two code units. The SDK's own example: the G-clef 𝄞 has `length` 2 and `runes.length` 1. - **Grapheme cluster**: what a person calls one character. A flag is two regional-indicator code points; 👨👩👧 is five code points (three people joined by two zero-width joiners) and eight code units, yet one visible character. ## Where it bites 1. **Length limits.** A menu importer that rejects names longer than 30 `length` units rejects emoji-heavy names that look short. 2. **Truncation.** `name.substring(0, 7)` on `'Pizza 🍕🇮🇹'` ends between the two code units of 🍕, producing a lone surrogate that renders as a replacement symbol. dart.dev shows the same effect: the last code unit of `'Hi 🇩🇰'` prints as garbage, while `characters.last` prints the flag. 3. **Reversal and indexing.** Reversing code units or picking `name[i]` breaks pairs and clusters. 4. **Equality.** `==` compares code units and "does not check for Unicode equivalence", so `'Am\xe9lie' == 'Ame\u{301}lie'` is `false` even though both render as Amélie. ## Doing it right ```dart import 'package:characters/characters.dart'; final visible = name.characters.length; final short = name.characters.take(7).join(); ``` In Flutter code you usually already have `characters`, because `package:flutter/widgets.dart` re-exports it; Flutter's own `TextField` counts its `maxLength` with `characters.length` so an emoji counts as one. ## Writing Unicode in source - `\uXXXX` writes a code point with exactly four hex digits, for example `\u2665` for ♥. - `\u{...}` takes more or fewer than four hex digits in braces, for example `\u{1F355}` for 🍕. - Emoji can also be typed directly; the source file is UTF-8, and the string still stores UTF-16 code units. ## Rules for a menu importer 1. **Validate** name limits with `characters.length`, so a two-character emoji name is not rejected as four. 2. **Truncate** for receipts or previews with `characters.take(n)`, never with `substring`, so no emoji is cut in half. 3. **Compare** names knowing that `==` works on code units and that `dart:core`'s `String` offers no Unicode normalization method; if names from different sources must match, normalize them before they reach Dart or with a dedicated package. 4. **Store** the original string unchanged; counting rules belong in validation and display, not in the data. 5. **Test** with emoji, flags, family sequences and combining accents, because ASCII-only fixtures hide every one of these bugs. ## When code units are fine Code units are correct when you need storage size in UTF-16, when you work with ASCII-only data such as currency codes, or when positions come from another code-unit-based API. They are wrong whenever a person will see or count the result.
- What does runes give you that length does not, and why is it still not enough?`runes` combines surrogate pairs into Unicode code points, so an emoji like 🍕 counts once. But one visible character can be several code points, such as a flag (two regional indicators) or a family emoji joined with zero-width joiners. Only grapheme clusters, via `characters`, match what users see.
- Why are 'Am\xe9lie' and 'Ame\u{301}lie' not equal in Dart?String `==` compares code units one by one and does not apply Unicode normalization. The first uses a precomposed é, the second an e followed by a combining accent, so the code units differ even though both render the same.
Counting code units is like counting parcels instead of furniture: a chair ships in one box, a table in two, and a sofa set in four. The box count is exact for the courier but overstates what the customer ordered, and opening only the first box of the table gives the customer half a table.
saying these in an interview costs you the question
- String.length counts the characters a user sees
- runes always equals the number of visible characters
- Dart strings are stored as UTF-8 bytes
- substring always cuts between whole characters
- == treats differently composed accents as equal