A PHP marketplace truncates emoji-rich product names with mb_substr() and leaves broken flags and family emoji; why, and what does intl's grapheme API fix?
answer
- code point is not a visible character
- regional indicator pairs and ZWJ sequences
- grapheme_strlen and grapheme_substr
- grapheme_str_split since PHP 8.4
- intl needs valid UTF-8
basics
~20 smb_substr() counts code points, but a flag is two code points and a family emoji is several joined by zero-width joiners, so a cut can split one visible character. grapheme_substr() and grapheme_strlen() from intl count grapheme clusters instead.
solid answer
~40 smbstring works in **code points**, while readers see **grapheme clusters**: user-perceived characters that may span several code points. A flag is two regional-indicator symbols, a thumbs-up with a skin tone is two code points, and a family emoji is several people joined by U+200D zero-width joiners. `mb_substr($name, 0, 20)` counts code points, so it can keep half a flag or the first person of a family, which renders as a stray letter or a single face. The intl extension's grapheme functions follow Unicode's cluster rules via ICU: `grapheme_strlen()` counts clusters, `grapheme_substr()` cuts on cluster boundaries, and `grapheme_str_split()`, added in PHP 8.4, splits into clusters. They need valid UTF-8 (on invalid input `grapheme_strlen()` returns `null` and `grapheme_substr()` returns `false`) and cost more than mbstring, so use them where limits are user-facing.
code
php · 17 lines<?php
declare(strict_types=1);
$flag = "\u{1F1EF}\u{1F1F5}"; // Japan flag
$family = "\u{1F468}\u{200D}\u{1F469}\u{200D}\u{1F467}"; // man, ZWJ, woman, ZWJ, girl
$title = $flag . ' Kimono ' . $family;
var_dump(mb_strlen($title, 'UTF-8')); // int(15)
var_dump(grapheme_strlen($title)); // int(10)
var_dump(mb_substr($flag, 0, 1, 'UTF-8') === "\u{1F1EF}"); // bool(true): half a flag
var_dump(grapheme_substr($title, 0, 1) === $flag); // bool(true): the whole flag
var_dump(count(grapheme_str_split($family))); // int(1), PHP 8.4+
$bad = "\xE6\x97"; // truncated UTF-8
var_dump(grapheme_strlen($bad)); // NULL: validate firstgo deeper
Recall that some visible characters, such as flags and emoji with skin tones, are made of several code points.
Explain code points versus grapheme clusters with flag and ZWJ examples, and name grapheme_strlen, grapheme_substr and grapheme_str_split.
Diagnose a split-emoji bug, move user-facing limits to grapheme functions, validate UTF-8 before calling them, and declare ext-intl with a known ICU version.
Align the unit of every text limit across form, API and storage, and own the decision of which layer enforces the user-facing grapheme limit.
## Code points versus what the reader sees Unicode assigns numbers, **code points**, to abstract characters. mbstring functions such as `mb_strlen()` and `mb_substr()` count and cut in code points. A reader, however, perceives **grapheme clusters**: sequences of one or more code points that display as a single character. Unicode defines where cluster boundaries fall, and most text has one code point per cluster, which is why the difference goes unnoticed until emoji or combining marks appear. Common multi-code-point clusters in product names: | Visible character | Code points | UTF-8 bytes | |---|---|---| | `é` written as `e` + U+0301 combining acute | 2 | 3 | | 🇯🇵 flag: regional indicators J and P | 2 | 8 | | 👍🏽 thumbs up + skin-tone modifier | 2 | 8 | | 👨👩👧 man, ZWJ, woman, ZWJ, girl | 5 | 18 | `ZWJ` is U+200D ZERO WIDTH JOINER, which glues emoji into one glyph when the font supports the sequence. ## What goes wrong with mb_substr() A marketplace that limits listing titles to 20 characters with `mb_substr($title, 0, 20)` is safe for encoding, because it never splits a byte sequence, but it can split a cluster: - cutting after the first regional indicator leaves a lone letter-like symbol instead of a flag; - cutting inside a ZWJ sequence leaves one person of a family, or a trailing invisible joiner; - cutting between a letter and its combining accent drops the accent. Counting has the same problem: `mb_strlen()` says a single family emoji is 5 characters, so a title that looks like 18 characters can be rejected as 30 long. ## The grapheme functions The **intl** extension wraps the ICU library and provides grapheme-aware counterparts: 1. `grapheme_strlen(string $string): int|false|null` returns the number of clusters. 2. `grapheme_substr(string $string, int $offset, ?int $length = null, string $locale = ""): string|false` cuts on cluster boundaries. 3. `grapheme_str_split(string $string, int $length = 1): array|false`, added in **PHP 8.4**, splits into clusters. 4. `grapheme_strpos()`, `grapheme_stripos()`, `grapheme_strrpos()` and `grapheme_strripos()` return offsets in clusters. 5. `grapheme_extract()` takes up to a number of clusters, bytes or code points starting at a byte offset, useful for byte budgets that must not split clusters. PHP 8.5 added a `$locale` parameter to `grapheme_substr()`, `grapheme_strpos()` and related search functions, and a new `grapheme_levenshtein()`. ## Operational constraints - **Valid UTF-8 only.** The grapheme functions require UTF-8 input. On invalid input `grapheme_strlen()` returns `null`, and functions such as `grapheme_substr()` return `false`. Validate with `mb_check_encoding($s, 'UTF-8')` first, and handle the failure value rather than passing it on. - **An extra extension.** intl depends on ICU. Declare `"ext-intl": "*"` in `composer.json` so a server without it fails at install time, not at runtime. - **ICU version matters.** Segmentation uses ICU's Unicode character data; emoji code points added in a newer Unicode version than a server's ICU knows may be segmented differently, so results can differ between servers. - **Cost.** Segmenting text is more work than counting code points. For a listing page it is negligible; for bulk processing of millions of rows, measure. ## A worked fix For a listing that shows at most 20 visible characters of a title: 1. Validate once on input: reject titles for which `mb_check_encoding($title, 'UTF-8')` is `false`. 2. Measure with `grapheme_strlen($title)` and check the result is an `int`. 3. If it exceeds 20, take `grapheme_substr($title, 0, 19)` and append `'…'`. 4. Enforce the same grapheme limit in the seller form's validation, so the preview and the stored listing agree. 5. Size the storage column in bytes with generous headroom, since a single family emoji is 18 bytes. ## Choosing the unit per rule - **Storage and transport limits**: bytes, with `strlen()`, or `mb_strcut()` when you must cut. - **Protocol or database limits defined in code points**: `mb_strlen()` and `mb_substr()`. - **Anything a user reads or counts**, such as a title limit or a character counter: `grapheme_strlen()` and `grapheme_substr()`. A frequent production bug is a mismatch between these units across layers: the form counts one way, the API validates another way, and the database stores a third. Agree on graphemes for user-facing rules and derive the byte budget for storage from them with headroom.
- What does grapheme_strlen() return for invalid UTF-8, and how should calling code handle it?It returns `null` rather than a count, because the input cannot be converted for ICU; its `int|false|null` return type allows both failure values. Validate input with `mb_check_encoding($s, 'UTF-8')` first, and treat a non-int result as an error; passing it into a comparison or arithmetic silently hides the problem.
- Why can the same emoji-heavy title have different grapheme counts on two servers?Cluster boundaries come from ICU's implementation of Unicode segmentation rules. A server whose ICU data predates recently added emoji code points may not classify them as emoji and can split a sequence that a newer ICU keeps together. Pinning ICU through the platform image, or at least recording the version, keeps behaviour consistent.
- When is mb_substr() still the right choice over grapheme_substr()?When the limit is defined in code points by another system, or the text is known not to contain combining marks or emoji sequences, such as generated identifiers. It is cheaper and does not require intl, while still never producing invalid UTF-8.
saying these in an interview costs you the question
- mb_substr never splits anything a reader sees as one character.
- A family emoji is a single code point.
- grapheme_strlen quietly repairs invalid UTF-8 and returns a count.
- The grapheme functions come from mbstring, so no extra extension is needed.
- Grapheme clusters only matter for emoji, never for accented letters.