In PHP, why does strlen('日本語') return 9, and how do you count the characters of a UTF-8 string instead?
answer
- strings are byte sequences
- UTF-8 uses 1 to 4 bytes
- mb_strlen counts code points
- default_charset feeds the mbstring default
- bytes still matter for storage
basics
~20 sA PHP string is a sequence of bytes with no encoding attached, and strlen() counts bytes; each of those three kanji takes three bytes in UTF-8. mb_strlen($s, 'UTF-8') decodes the bytes and counts characters, returning 3.
solid answer
~40 sPHP strings are plain byte arrays: the engine does not know or store an encoding. `strlen()` returns the number of **bytes**, and UTF-8 encodes a code point in one to four bytes, so `'日本語'` is 9 bytes and `'👍'` is 4. The mbstring extension decodes the bytes: `mb_strlen($s, 'UTF-8')` returns the number of **code points**, 3 for `'日本語'`. Without the encoding argument mbstring uses its internal encoding, which comes from `default_charset`, `UTF-8` out of the box; passing it explicitly removes the dependency on configuration. Byte counts are still the right measure for storage, network payloads and binary data. For what a reader perceives as one character, such as a flag or an emoji with a skin tone, even `mb_strlen()` over-counts, and `grapheme_strlen()` from intl is needed.
code
php · 16 lines<?php
declare(strict_types=1);
$title = '日本語';
var_dump(strlen($title)); // int(9): bytes
var_dump(mb_strlen($title, 'UTF-8')); // int(3): code points
var_dump(bin2hex($title[0])); // string(2) "e6": one byte, not a character
var_dump(mb_substr($title, 0, 1, 'UTF-8')); // string(3) "日"
$thumb = "\u{1F44D}\u{1F3FD}"; // thumbs up + medium skin tone
var_dump(strlen($thumb)); // int(8)
var_dump(mb_strlen($thumb, 'UTF-8')); // int(2)
var_dump(grapheme_strlen($thumb)); // int(1), needs ext-intl
var_dump(mb_internal_encoding()); // string(5) "UTF-8" with default settingsgo deeper
Recall that strlen counts bytes, that UTF-8 characters take one to four bytes, and that mb_strlen with 'UTF-8' counts characters.
Explain the byte model of PHP strings, where mbstring's default encoding comes from, and the difference between code points and grapheme clusters.
Show you would pick bytes, code points or graphemes deliberately per rule, declare ext-mbstring and ext-intl, and pass encodings explicitly in shared code.
Set a text-handling policy for the codebase: UTF-8 everywhere, which unit each limit is defined in, and which layer validates encoding on input.
## A PHP string is bytes In PHP a `string` is an ordered sequence of **bytes** plus a length. The engine attaches no encoding to it: the same variable can hold ASCII text, UTF-8 text, a Shift_JIS file or a PNG image. Functions from the core string library, such as `strlen()`, `substr()` and `strtoupper()`, work on those bytes and nothing else. **UTF-8** encodes each Unicode **code point** in one to four bytes: - ASCII letters and digits: 1 byte (`a` is `0x61`). - Latin letters with accents, Greek, Cyrillic: 2 bytes (`é` is `0xC3 0xA9`). - Most CJK characters, including Japanese kana and kanji: 3 bytes (`日` is `0xE6 0x97 0xA5`). - Emoji and other characters outside the Basic Multilingual Plane: 4 bytes. So `strlen('日本語')` returns `9`: three characters of three bytes each. Nothing is wrong with the function; it answers a different question. ## Three ways to count | String | `strlen()` bytes | `mb_strlen()` code points | `grapheme_strlen()` perceived characters | |---|---|---|---| | `'abc'` | 3 | 3 | 3 | | `'café'` (precomposed é) | 5 | 4 | 4 | | `'日本語'` | 9 | 3 | 3 | | `'👍'` | 4 | 1 | 1 | | `'👍🏽'` (thumb plus skin tone) | 8 | 2 | 1 | The first two columns come from the core and mbstring; the third comes from the intl extension and counts **grapheme clusters**, the units a reader sees as one character. ## mb_strlen() and the internal encoding `mb_strlen(string $string, ?string $encoding = null): int` decodes the bytes in the given encoding and returns the number of code points. When `$encoding` is `null`, mbstring uses its **internal encoding**: 1. `mb_internal_encoding()` returns it, and `mb_internal_encoding('UTF-8')` sets it for the rest of the request. 2. If nothing sets it, it follows the `default_charset` ini directive, whose built-in default is `UTF-8`. 3. The old `mbstring.internal_encoding` directive is deprecated; configure `default_charset` instead. Because the result depends on configuration, many teams pass `'UTF-8'` explicitly in library code. An unknown encoding name throws a `ValueError`. mbstring is an **extension**. Distribution packages usually ship it, but it is not compiled in by default when PHP is built from source, so a project that relies on it should declare the dependency, for example `"ext-mbstring": "*"` in `composer.json`, instead of discovering it on a server without it. ## When bytes are the right answer Counting characters is not always the goal. `strlen()` is correct when the limit is physical: - the size of a payload, a file or a cache entry; - a binary protocol field or a hash, which are bytes by definition; - a storage limit defined in bytes rather than characters. The mistake is using a byte count for a **human** rule, such as "a product title may have at most 40 characters". A Japanese title of 14 characters is already 42 bytes and would be rejected by a `strlen()` check, while an English title of 40 letters passes. Validation, truncation and padding of visible text need a character-aware function. ## Byte offsets The same byte model applies to offsets. `$s[0]` and `substr($s, 0, 1)` return the first **byte**, which for `'日本語'` is `0xE6`, an incomplete sequence that is not valid UTF-8 on its own. `mb_substr($s, 0, 1)` returns `'日'`. Any code that slices text written by users should use the `mb_*` family, or the `grapheme_*` family when emoji and combining marks must stay intact. ## Symptoms in a real codebase Byte-counting bugs rarely announce themselves. Typical symptoms in a marketplace: - a title validator that accepts 40 English letters but rejects a 14-character Japanese title as too long; - a character counter in the seller form that disagrees with the server by a factor of three for kana and kanji; - a column of product names, padded or truncated with byte functions, whose layout breaks only for non-Latin rows; - a `strrev()` or a byte-by-byte loop that turns Japanese text into invalid UTF-8. All four disappear once the code states the unit it means and uses the matching function. ## What an interviewer wants to hear A solid answer states that PHP strings are bytes, that `strlen()` counts bytes, that UTF-8 is variable-width, and that `mb_strlen()` with an explicit `'UTF-8'` counts code points. A better one adds that code points are still not what users see, names `grapheme_strlen()`, and says when a byte count is exactly what you want.
- Where does mb_strlen() get its encoding when you do not pass one?From mbstring's internal encoding, readable and settable with `mb_internal_encoding()`. If nothing sets it, it follows `default_charset`, which is `UTF-8` by default. The old `mbstring.internal_encoding` directive is deprecated. Passing `'UTF-8'` explicitly makes library code independent of that configuration.
- Why can mb_strlen() still disagree with what a user counts on screen?It counts code points, and one visible character can be several: an emoji with a skin-tone modifier is two, a flag is two regional-indicator symbols, and a letter plus a combining accent is two. `grapheme_strlen()` from intl counts grapheme clusters, which match what the reader perceives.
- When is strlen() the correct function even for UTF-8 text?When the limit is about bytes: payload and file sizes, binary protocol fields, hashes, or a storage column whose limit is defined in bytes. The bug is using a byte count for a human rule such as a maximum title length.
saying these in an interview costs you the question
- PHP strings are UTF-8 internally, so strlen counts characters.
- strlen returns the number of characters for any text in PHP 8.
- mb_strlen counts what users see, including emoji with skin tones as one.
- $s[0] returns the first character of a UTF-8 string.
- mbstring is always compiled into PHP, so no dependency needs declaring.