skip to content

Multibyte & UTF-8

strlen counts bytes, so UTF-8 text needs mbstring: mb_strlen, mb_substr and mb_str_pad count characters, and intl normalizes text. Interviewers probe the byte-versus-character bug.

part ofPHPoverview, primer and where to startread it →
on this pageshow

explore

questions

6

In PHP, why does strlen('日本語') return 9, and how do you count the characters of a UTF-8 string instead?

level: juniorimportance: must knowfreq 72%

answer

  1. strings are byte sequences
  2. UTF-8 uses 1 to 4 bytes
  3. mb_strlen counts code points
  4. default_charset feeds the mbstring default
  5. bytes still matter for storage

basics

~20 s

A PHP string is a sequence of bytes with no encoding attached, and strlen() counts bytes; each of those three kanji takes three bytes in UTF-8. mb_strlen($s, 'UTF-8') decodes the bytes and counts characters, returning 3.

solid answer

~40 s

PHP strings are plain byte arrays: the engine does not know or store an encoding. `strlen()` returns the number of **bytes**, and UTF-8 encodes a code point in one to four bytes, so `'日本語'` is 9 bytes and `'👍'` is 4. The mbstring extension decodes the bytes: `mb_strlen($s, 'UTF-8')` returns the number of **code points**, 3 for `'日本語'`. Without the encoding argument mbstring uses its internal encoding, which comes from `default_charset`, `UTF-8` out of the box; passing it explicitly removes the dependency on configuration. Byte counts are still the right measure for storage, network payloads and binary data. For what a reader perceives as one character, such as a flag or an emoji with a skin tone, even `mb_strlen()` over-counts, and `grapheme_strlen()` from intl is needed.

code

php · 16 lines
php
<?php
declare(strict_types=1);

$title = '日本語';

var_dump(strlen($title));                    // int(9): bytes
var_dump(mb_strlen($title, 'UTF-8'));        // int(3): code points
var_dump(bin2hex($title[0]));                // string(2) "e6": one byte, not a character
var_dump(mb_substr($title, 0, 1, 'UTF-8'));  // string(3) "日"

$thumb = "\u{1F44D}\u{1F3FD}";                // thumbs up + medium skin tone
var_dump(strlen($thumb));                    // int(8)
var_dump(mb_strlen($thumb, 'UTF-8'));        // int(2)
var_dump(grapheme_strlen($thumb));           // int(1), needs ext-intl

var_dump(mb_internal_encoding());            // string(5) "UTF-8" with default settings

go deeper

for a junior

Recall that strlen counts bytes, that UTF-8 characters take one to four bytes, and that mb_strlen with 'UTF-8' counts characters.

for a middle

Explain the byte model of PHP strings, where mbstring's default encoding comes from, and the difference between code points and grapheme clusters.

for a senior

Show you would pick bytes, code points or graphemes deliberately per rule, declare ext-mbstring and ext-intl, and pass encodings explicitly in shared code.

for a principal

Set a text-handling policy for the codebase: UTF-8 everywhere, which unit each limit is defined in, and which layer validates encoding on input.

## A PHP string is bytes In PHP a `string` is an ordered sequence of **bytes** plus a length. The engine attaches no encoding to it: the same variable can hold ASCII text, UTF-8 text, a Shift_JIS file or a PNG image. Functions from the core string library, such as `strlen()`, `substr()` and `strtoupper()`, work on those bytes and nothing else. **UTF-8** encodes each Unicode **code point** in one to four bytes: - ASCII letters and digits: 1 byte (`a` is `0x61`). - Latin letters with accents, Greek, Cyrillic: 2 bytes (`é` is `0xC3 0xA9`). - Most CJK characters, including Japanese kana and kanji: 3 bytes (`日` is `0xE6 0x97 0xA5`). - Emoji and other characters outside the Basic Multilingual Plane: 4 bytes. So `strlen('日本語')` returns `9`: three characters of three bytes each. Nothing is wrong with the function; it answers a different question. ## Three ways to count | String | `strlen()` bytes | `mb_strlen()` code points | `grapheme_strlen()` perceived characters | |---|---|---|---| | `'abc'` | 3 | 3 | 3 | | `'café'` (precomposed é) | 5 | 4 | 4 | | `'日本語'` | 9 | 3 | 3 | | `'👍'` | 4 | 1 | 1 | | `'👍🏽'` (thumb plus skin tone) | 8 | 2 | 1 | The first two columns come from the core and mbstring; the third comes from the intl extension and counts **grapheme clusters**, the units a reader sees as one character. ## mb_strlen() and the internal encoding `mb_strlen(string $string, ?string $encoding = null): int` decodes the bytes in the given encoding and returns the number of code points. When `$encoding` is `null`, mbstring uses its **internal encoding**: 1. `mb_internal_encoding()` returns it, and `mb_internal_encoding('UTF-8')` sets it for the rest of the request. 2. If nothing sets it, it follows the `default_charset` ini directive, whose built-in default is `UTF-8`. 3. The old `mbstring.internal_encoding` directive is deprecated; configure `default_charset` instead. Because the result depends on configuration, many teams pass `'UTF-8'` explicitly in library code. An unknown encoding name throws a `ValueError`. mbstring is an **extension**. Distribution packages usually ship it, but it is not compiled in by default when PHP is built from source, so a project that relies on it should declare the dependency, for example `"ext-mbstring": "*"` in `composer.json`, instead of discovering it on a server without it. ## When bytes are the right answer Counting characters is not always the goal. `strlen()` is correct when the limit is physical: - the size of a payload, a file or a cache entry; - a binary protocol field or a hash, which are bytes by definition; - a storage limit defined in bytes rather than characters. The mistake is using a byte count for a **human** rule, such as "a product title may have at most 40 characters". A Japanese title of 14 characters is already 42 bytes and would be rejected by a `strlen()` check, while an English title of 40 letters passes. Validation, truncation and padding of visible text need a character-aware function. ## Byte offsets The same byte model applies to offsets. `$s[0]` and `substr($s, 0, 1)` return the first **byte**, which for `'日本語'` is `0xE6`, an incomplete sequence that is not valid UTF-8 on its own. `mb_substr($s, 0, 1)` returns `'日'`. Any code that slices text written by users should use the `mb_*` family, or the `grapheme_*` family when emoji and combining marks must stay intact. ## Symptoms in a real codebase Byte-counting bugs rarely announce themselves. Typical symptoms in a marketplace: - a title validator that accepts 40 English letters but rejects a 14-character Japanese title as too long; - a character counter in the seller form that disagrees with the server by a factor of three for kana and kanji; - a column of product names, padded or truncated with byte functions, whose layout breaks only for non-Latin rows; - a `strrev()` or a byte-by-byte loop that turns Japanese text into invalid UTF-8. All four disappear once the code states the unit it means and uses the matching function. ## What an interviewer wants to hear A solid answer states that PHP strings are bytes, that `strlen()` counts bytes, that UTF-8 is variable-width, and that `mb_strlen()` with an explicit `'UTF-8'` counts code points. A better one adds that code points are still not what users see, names `grapheme_strlen()`, and says when a byte count is exactly what you want.

  • Where does mb_strlen() get its encoding when you do not pass one?
    From mbstring's internal encoding, readable and settable with `mb_internal_encoding()`. If nothing sets it, it follows `default_charset`, which is `UTF-8` by default. The old `mbstring.internal_encoding` directive is deprecated. Passing `'UTF-8'` explicitly makes library code independent of that configuration.
  • Why can mb_strlen() still disagree with what a user counts on screen?
    It counts code points, and one visible character can be several: an emoji with a skin-tone modifier is two, a flag is two regional-indicator symbols, and a letter plus a combining accent is two. `grapheme_strlen()` from intl counts grapheme clusters, which match what the reader perceives.
  • When is strlen() the correct function even for UTF-8 text?
    When the limit is about bytes: payload and file sizes, binary protocol fields, hashes, or a storage column whose limit is defined in bytes. The bug is using a byte count for a human rule such as a maximum title length.

saying these in an interview costs you the question

  • PHP strings are UTF-8 internally, so strlen counts characters.
  • strlen returns the number of characters for any text in PHP 8.
  • mb_strlen counts what users see, including emoji with skin tones as one.
  • $s[0] returns the first character of a UTF-8 string.
  • mbstring is always compiled into PHP, so no dependency needs declaring.
open as a page

In a PHP marketplace, why does substr($name, 0, 20) corrupt Japanese product names, and how do you truncate them safely for a listing?

level: middleimportance: must knowfreq 55%

basics

~20 s

substr() cuts at a byte offset, so it can split a three-byte Japanese character and leave invalid UTF-8. mb_substr() cuts at character boundaries, and mb_strimwidth() truncates to a display width, counting full-width characters as two columns.

open as a page

In PHP 8.4 and later, why do strtolower(), str_pad(), trim() and ucfirst() mishandle UTF-8 text, and which mb_* functions replace them?

level: middleimportance: should knowfreq 40%

basics

~10 s

The byte functions only understand ASCII: strtolower() leaves 'Ä' alone, str_pad() counts bytes, trim() ignores the ideographic space U+3000, and ucfirst() cannot uppercase 'é'. Use mb_strtolower(), mb_str_pad() (8.3), mb_trim() and mb_ucfirst() (8.4).

open as a page

A PHP marketplace truncates emoji-rich product names with mb_substr() and leaves broken flags and family emoji; why, and what does intl's grapheme API fix?

level: seniorimportance: should knowfreq 30%

basics

~20 s

mb_substr() counts code points, but a flag is two code points and a family emoji is several joined by zero-width joiners, so a cut can split one visible character. grapheme_substr() and grapheme_strlen() from intl count grapheme clusters instead.

open as a page

A PHP marketplace imports supplier product feeds in Shift_JIS and sometimes broken UTF-8; how do you validate and convert them with mb_check_encoding(), mb_convert_encoding() and iconv()?

level: seniorimportance: should knowfreq 32%

basics

~10 s

Convert each feed once, from its declared encoding, with mb_convert_encoding($s, 'UTF-8', 'SJIS'), then check the result with mb_check_encoding($s, 'UTF-8'). mbstring replaces invalid bytes with a substitute character; iconv() returns false unless //IGNORE is appended.

open as a page

In PHP, why can two identical-looking UTF-8 strings such as 'café' fail an === comparison, and how does intl's Normalizer fix it?

level: middleimportance: nice to knowfreq 22%

basics

~10 s

'é' can be one code point, U+00E9, or 'e' plus a combining accent U+0301; both render the same but differ in bytes, so === fails. Normalizer::normalize($s, Normalizer::FORM_C) converts both to the same composed form.

open as a page