In PHP 8.4 and later, why do strtolower(), str_pad(), trim() and ucfirst() mishandle UTF-8 text, and which mb_* functions replace them?
answer
- byte functions: ASCII case only
- str_pad measures bytes
- trim misses U+3000 and NBSP
- mb_str_pad in 8.3
- mb_trim and mb_ucfirst in 8.4
basics
~10 sThe byte functions only understand ASCII: strtolower() leaves 'Ä' alone, str_pad() counts bytes, trim() ignores the ideographic space U+3000, and ucfirst() cannot uppercase 'é'. Use mb_strtolower(), mb_str_pad() (8.3), mb_trim() and mb_ucfirst() (8.4).
solid answer
~40 sThe classic string functions work on bytes. Since PHP 8.2 `strtolower()`, `strtoupper()`, `ucfirst()` and friends perform **ASCII-only** case conversion regardless of locale, so `strtolower('ÄPFEL')` returns `'Äpfel'`. `mb_strtolower()` and `mb_strtoupper()` apply full Unicode case mapping, so `mb_strtoupper('Straße')` is `'STRASSE'`. `str_pad()` pads to a **byte** length: `str_pad('日本', 6, '*')` adds nothing because the string is already six bytes, while `mb_str_pad()`, added in PHP 8.3, pads to a character count. `trim()` strips only ASCII whitespace and NUL, missing the non-breaking space and the Japanese ideographic space U+3000; `mb_trim()`, `mb_ltrim()` and `mb_rtrim()` from PHP 8.4 strip Unicode whitespace. `mb_ucfirst()` and `mb_lcfirst()`, also 8.4, replace `ucfirst()` and `lcfirst()`.
code
php · 16 lines<?php
declare(strict_types=1);
var_dump(strtolower('ÄPFEL')); // string(6) "Äpfel"
var_dump(mb_strtolower('ÄPFEL', 'UTF-8')); // string(6) "äpfel"
var_dump(mb_strtoupper('Straße', 'UTF-8')); // string(7) "STRASSE"
var_dump(str_pad('日本', 6, '*')); // string(6) "日本": already 6 bytes
var_dump(mb_str_pad('日本', 6, '*')); // string(10) "日本****"
$pasted = "\u{3000}限定モデル\u{00A0}"; // ideographic space, NBSP
var_dump(trim($pasted) === $pasted); // bool(true): nothing stripped
var_dump(mb_trim($pasted)); // string(15) "限定モデル"
var_dump(ucfirst('élan')); // string(5) "élan"
var_dump(mb_ucfirst('élan')); // string(5) "Élan"go deeper
Recall that the byte helpers only handle ASCII letters and whitespace, and that mb_strtolower, mb_str_pad and mb_trim are the UTF-8 versions.
Explain the PHP 8.2 ASCII-only case change, byte versus character padding, mb_trim's default list, and which release added each replacement.
Show you would audit input cleanup for trim and strtolower on user text, raise the platform requirement to 8.4 when adopting mb_trim, and pad by width where alignment matters.
Decide whether to standardise on mbstring helpers everywhere or behind a small text utility, and how to enforce that choice in review and static analysis.
## Why the byte functions fail The core string functions treat a string as bytes and know the letters of ASCII only. For English text that is invisible. For product names in Japanese, German or French, each of the everyday helpers has a failure mode, and mbstring has provided or gained a replacement for each. | Byte function | Problem with UTF-8 | Replacement | Available since | |---|---|---|---| | `strtolower()` / `strtoupper()` | ASCII letters only | `mb_strtolower()` / `mb_strtoupper()` | long-standing | | `ucfirst()` / `lcfirst()` | ASCII first byte only | `mb_ucfirst()` / `mb_lcfirst()` | PHP 8.4 | | `ucwords()` | ASCII letters only | `mb_convert_case($s, MB_CASE_TITLE)` | long-standing | | `str_pad()` | length in bytes | `mb_str_pad()` | PHP 8.3 | | `trim()` / `ltrim()` / `rtrim()` | ASCII whitespace only | `mb_trim()` / `mb_ltrim()` / `mb_rtrim()` | PHP 8.4 | | `str_split()` | splits bytes | `mb_str_split()` | PHP 7.4 | ## Case conversion Since **PHP 8.2**, `strtolower()`, `strtoupper()`, `ucfirst()`, `lcfirst()`, `ucwords()` and `str_ireplace()` no longer consult the locale; they convert ASCII letters as if the locale were `C`. So: - `strtolower('ÄPFEL')` gives `'Äpfel'`: the `Ä` is two bytes that are not ASCII letters. - `ucfirst('élan')` returns `'élan'` unchanged. `mb_strtolower(string $string, ?string $encoding = null)` and `mb_strtoupper()` use Unicode case mapping, including **full** mappings that change length: `mb_strtoupper('Straße')` is `'STRASSE'`. `mb_convert_case()` offers more modes through constants: - `MB_CASE_UPPER`, `MB_CASE_LOWER`, `MB_CASE_TITLE` for full mappings; - `MB_CASE_FOLD` for case folding, intended for comparisons; - `*_SIMPLE` variants that keep one code point per code point. `mb_ucfirst()` and `mb_lcfirst()`, added in **PHP 8.4**, change only the first character; `mb_ucfirst()` uses the **title-case** mapping, which matters for a few characters whose title case differs from their uppercase. Japanese kana and kanji have no case, so these functions leave them alone; the issue is mixed catalogues with Latin, Greek or Cyrillic names. ## Padding `str_pad(string $string, int $length, string $pad_string = " ", int $pad_type = STR_PAD_RIGHT)` compares `$length` with the **byte** length. `str_pad('日本', 6, '*')` returns `'日本'` unchanged, because the two kanji are already six bytes, and a padding string with multi-byte characters can be cut mid-character. `mb_str_pad()`, added in **PHP 8.3**, has the same parameters plus `$encoding` and measures in code points: `mb_str_pad('日本', 6, '*')` returns `'日本****'`. The padding string may itself be multi-byte, and is truncated at character boundaries if it does not fit evenly. Neither function measures display width. For aligned plain-text columns that mix full-width and half-width characters, compute the padding from `mb_strwidth()` instead. ## Trimming `trim()` strips only space, tab, newline, carriage return, vertical tab and NUL. Text pasted from spreadsheets and typed with Japanese input methods often carries: - U+00A0 NO-BREAK SPACE; - U+3000 IDEOGRAPHIC SPACE, the full-width space of Japanese input; - other Unicode spaces such as U+2002 to U+200A, and line separators U+2028 and U+2029. `mb_trim(string $string, ?string $characters = null, ?string $encoding = null)`, added in **PHP 8.4** with `mb_ltrim()` and `mb_rtrim()`, strips those by default. Its default list does **not** include U+200B ZERO WIDTH SPACE or U+FEFF, the byte-order mark; pass them in `$characters` if your data contains them. ## Putting it together For a seller-entered product name, a typical cleanup is: 1. validate UTF-8 with `mb_check_encoding()`; 2. `mb_trim()` the value; 3. normalize it to NFC; 4. use `mb_*` functions for every later length, case and padding operation. The replacements need the mbstring extension, and `mb_str_pad()`, `mb_trim()` and `mb_ucfirst()` need PHP 8.3 or 8.4, so the `composer.json` platform requirement should say so. ## What interviewers listen for A good answer does not just list the `mb_*` names. It explains **why** each byte function fails: case conversion limited to ASCII since 8.2, padding and splitting by bytes, trimming only ASCII whitespace. It also knows the release that introduced each replacement, because adopting `mb_trim()` silently raises the minimum PHP version of a library. Strong candidates add the edge cases: full case mapping changes string length, `mb_trim()` does not strip a byte-order mark, and character-based padding still misaligns full-width text. Those points show the answer comes from handling real multilingual input rather than from a cheat sheet.
- Why does mb_strtoupper('Straße') return a longer string?Since PHP 7.3 mbstring applies full Unicode case mapping, in which the uppercase of `ß` is the two letters `SS`. Full mappings can change the length of the string. The `MB_CASE_UPPER_SIMPLE` mode of `mb_convert_case()` keeps a one-to-one mapping instead.
- Which whitespace does mb_trim() leave in place by default?Its default list covers ASCII whitespace, NUL, the no-break space, the ideographic space and the other Unicode space and line-separator characters, but not U+200B ZERO WIDTH SPACE or U+FEFF, the byte-order mark. Pass those in the `$characters` argument when your data contains them.
- Does mb_str_pad() align columns that mix Japanese and Latin text?Not visually. It pads to a number of code points, and a full-width character occupies two columns on screen. For aligned plain-text output, compute the padding from `mb_strwidth()`.
saying these in an interview costs you the question
- strtolower handles accented letters correctly when a UTF-8 locale is set in PHP 8.
- str_pad counts characters, so it pads Japanese text correctly.
- trim removes the Japanese ideographic space by default.
- mb_trim strips the byte-order mark U+FEFF by default.
- mb_str_pad aligns full-width and half-width text visually.