skip to content

In a PHP marketplace, why does substr($name, 0, 20) corrupt Japanese product names, and how do you truncate them safely for a listing?

level: middleimportance: must knowfreq 55%

answer

  1. cuts through a multi-byte sequence
  2. invalid UTF-8 downstream
  3. mb_substr counts characters
  4. full-width characters are two columns
  5. mb_strimwidth includes the marker

basics

~20 s

substr() cuts at a byte offset, so it can split a three-byte Japanese character and leave invalid UTF-8. mb_substr() cuts at character boundaries, and mb_strimwidth() truncates to a display width, counting full-width characters as two columns.

solid answer

~40 s

`substr()` counts bytes, and a Japanese character is three bytes in UTF-8, so a cut at byte 20 usually lands inside a character. The result is **invalid UTF-8**: browsers show a replacement glyph, `mb_check_encoding()` returns `false`, and `json_encode()` fails unless told to substitute. `mb_substr($name, 0, 20, 'UTF-8')` counts characters and never splits one. For a listing that must fit a visual width, `mb_strimwidth($name, 0, $width, '…', 'UTF-8')` treats full-width characters, such as kana and kanji, as two columns and includes the trim marker in the width. Append an ellipsis only when the text was actually shortened. `mb_strcut()` exists for byte limits: it cuts by bytes but backs off to a character boundary. Emoji sequences need the grapheme functions, because `mb_substr()` can still split a flag.

code

php · 22 lines
php
<?php
declare(strict_types=1);

$name = '東京タワー限定モデル';                    // 10 characters, 30 bytes

$broken = substr($name, 0, 10);                   // 3 characters + 1 stray byte
var_dump(mb_check_encoding($broken, 'UTF-8'));   // bool(false)

echo mb_substr($name, 0, 5, 'UTF-8'), "\n";       // 東京タワー
echo mb_strwidth($name, 'UTF-8'), "\n";           // 20
echo mb_strimwidth($name, 0, 10, '…', 'UTF-8'), "\n"; // 東京タワ…

function truncateTitle(string $title, int $max): string
{
    if (mb_strlen($title, 'UTF-8') <= $max) {
        return $title;
    }

    return mb_substr($title, 0, $max - 1, 'UTF-8') . '…';
}

echo truncateTitle($name, 6), "\n";                // 東京タワー…

go deeper

for a junior

Recall that substr counts bytes, that Japanese characters are three bytes in UTF-8, and that mb_substr counts characters.

for a middle

Explain why a byte cut produces invalid UTF-8, how mb_strimwidth measures full-width characters and the marker, and when mb_strcut fits.

for a senior

Show you would define each limit in a unit, validate encoding after byte cuts, test with CJK and emoji fixtures, and switch to grapheme functions for emoji-rich names.

for a principal

Decide where truncation happens, in the API or the view, so every client gets consistent, valid text and limits stay defined in one place.

## Why substr() breaks Japanese text A PHP string is a sequence of bytes, and `substr(string $string, int $offset, ?int $length = null): string` counts bytes. In UTF-8, Japanese kana and kanji take **three bytes** each. A listing that truncates names with `substr($name, 0, 20)` therefore keeps six full characters (18 bytes) and the first two bytes of the seventh. Those two dangling bytes are not a character. The string is now **invalid UTF-8**, and the damage shows up away from the truncation: - browsers render a replacement glyph or garbage at the end of the title; - `mb_check_encoding($title, 'UTF-8')` returns `false`; - `json_encode()` fails with a UTF-8 error for an API response, unless a substitution flag is passed. Latin-only test data never reveals the bug, because every ASCII character is one byte. ## Truncating by characters `mb_substr(string $string, int $start, ?int $length = null, ?string $encoding = null): string` counts **code points** in the given encoding, so it cannot split a multi-byte sequence: ```php $short = mb_substr($name, 0, 20, 'UTF-8'); ``` A listing usually adds an ellipsis, and only when something was removed: 1. Measure with `mb_strlen($name, 'UTF-8')`. 2. If it fits, return the name unchanged. 3. Otherwise take `mb_substr($name, 0, $max - 1)` and append `'…'`, so the result still has `$max` characters. ## Truncating by display width Character counts do not match the space text takes on screen. In a monospaced or grid layout, **full-width** characters, including kana, kanji and many emoji, occupy two columns, while Latin letters occupy one. mbstring models this East Asian width: - `mb_strwidth($s)` returns the width, counting full-width characters as 2. - `mb_strimwidth(string $string, int $start, int $width, string $trim_marker = "", ?string $encoding = null): string` returns a string whose width, **including the trim marker**, does not exceed `$width`. `mb_strimwidth('Hello World', 0, 10, '…')` keeps nine columns of text plus the one-column marker. For `'東京タワー限定モデル'` (width 20), a width of 10 yields `'東京タワ…'`: four two-column characters plus the marker, because a fifth character would exceed the limit. Mixed titles such as `'Tokyo タワー'` are where width-based truncation beats a character count. Since PHP 8.3, passing a **negative** width to `mb_strimwidth()` is deprecated. ## Byte limits without broken characters Sometimes the limit really is in bytes, for example a legacy column or a message field. `mb_strcut()` has the same parameters as `mb_substr()` but interprets `$start` and `$length` as **bytes**; if a cut falls inside a character it moves to that character's first byte, so the result stays valid. | Need | Function | Unit | |---|---|---| | Maximum characters | `mb_substr()` | code points | | Maximum visual width | `mb_strimwidth()` | columns, full-width = 2 | | Maximum bytes, valid output | `mb_strcut()` | bytes | | Maximum perceived characters | `grapheme_substr()` | grapheme clusters | | Never, for user text | `substr()` | bytes, may split | ## Where mb_substr() is still not enough `mb_substr()` is safe for encoding, but a code point is not always what the reader sees. Flags are pairs of code points, an emoji with a skin tone is two, and family emoji are sequences joined by zero-width joiners. Cutting between them leaves valid UTF-8 that looks wrong, such as a lone regional-indicator letter. When product names contain emoji, truncate with `grapheme_substr()` from the intl extension. ## A defensive checklist - Validate incoming names with `mb_check_encoding($name, 'UTF-8')` before storing them. - Define every length rule in a unit: characters, columns or bytes. - Test truncation with Japanese, accented Latin and emoji fixtures, not only ASCII. - Pass `'UTF-8'` explicitly or make sure `default_charset` is `UTF-8`. ## What interviewers listen for The first half of a good answer is the diagnosis: bytes versus characters, and invalid UTF-8 as the visible consequence. The second half is choosing the unit. Candidates who stop at `mb_substr()` answer the junior version; stronger candidates ask whether the listing limit is about characters, screen columns or bytes, name `mb_strimwidth()` and `mb_strcut()` for the last two, and point out that emoji sequences need grapheme functions. Mentioning that the ellipsis must be counted inside the limit, and appended only when text was removed, shows attention to the details that reviewers catch.

  • What does mb_strimwidth() count toward the width, and what does the trim marker cost?
    It counts East Asian width: full-width characters such as kana and kanji are 2 columns, most others 1. The trim marker is part of the result, so its width is subtracted from the budget; with `'…'` and a width of 10, at most nine columns of the original text remain.
  • When would you use mb_strcut() instead of mb_substr()?
    When the limit is in bytes, for example a storage field sized in bytes. `mb_strcut()` takes byte offsets and lengths, but if a cut falls inside a multi-byte character it moves to that character's start, so the output is always valid in the given encoding.
  • Why does a truncation bug in product names often show up as a failed API response?
    A byte cut leaves invalid UTF-8, and `json_encode()` refuses invalid UTF-8 unless a substitution flag is passed. The HTML page may merely show a replacement glyph, while the JSON endpoint returns an error. Validating with `mb_check_encoding()` after any byte-level cut catches it early.

saying these in an interview costs you the question

  • substr is safe for UTF-8 as long as the length is large enough.
  • mb_substr counts bytes but avoids splitting characters.
  • mb_strimwidth counts every character as one column.
  • The trim marker of mb_strimwidth is appended on top of the requested width.
  • mb_substr never splits anything a user would see as one character.