skip to content

A PHP marketplace imports supplier product feeds in Shift_JIS and sometimes broken UTF-8; how do you validate and convert them with mb_check_encoding(), mb_convert_encoding() and iconv()?

level: seniorimportance: should knowfreq 32%

answer

  1. never guess, use the declared encoding
  2. mb_check_encoding with an explicit encoding
  3. invalid bytes become the substitute character
  4. iconv returns false without //IGNORE
  5. unknown encoding name: ValueError

basics

~10 s

Convert each feed once, from its declared encoding, with mb_convert_encoding($s, 'UTF-8', 'SJIS'), then check the result with mb_check_encoding($s, 'UTF-8'). mbstring replaces invalid bytes with a substitute character; iconv() returns false unless //IGNORE is appended.

solid answer

~40 s

Treat encoding as a boundary concern: every byte stream entering the system is converted to UTF-8 once, from an encoding you **know**, not one you guess. `mb_convert_encoding($feed, 'UTF-8', 'SJIS')` (or `'CP932'` for Windows-produced files) converts, replacing bytes it cannot decode with the substitute character, `?` by default and configurable with `mb_substitute_character()`. For data that claims to be UTF-8, `mb_check_encoding($s, 'UTF-8')` validates, recursively for arrays, and `mb_scrub()` replaces invalid sequences. `iconv($from, $to, $s)` returns `false` with a notice on illegal input unless `//IGNORE` is appended, and `//TRANSLIT` depends on the platform's iconv implementation. `mb_detect_encoding()` is a heuristic, so use it only as a last resort with `strict` set. Unknown encoding names throw `ValueError`, and calling `mb_check_encoding()` without arguments is deprecated since 8.1.

code

php · 22 lines
php
<?php
declare(strict_types=1);

mb_substitute_character(0xFFFD);                      // make damage visible

$raw = file_get_contents('supplier-feed.csv');         // declared as Shift_JIS
if ($raw === false) {
    throw new RuntimeException('Feed not readable');
}

$utf8 = mb_convert_encoding($raw, 'UTF-8', 'SJIS');
if (!mb_check_encoding($utf8, 'UTF-8')) {
    throw new UnexpectedValueException('Conversion produced invalid UTF-8');
}

$replaced = mb_substr_count($utf8, "\u{FFFD}");
if ($replaced > 0) {
    error_log("Feed had {$replaced} undecodable sequences");
}

$latin1 = iconv('UTF-8', 'ISO-8859-1', 'Price: 100 €');
var_dump($latin1);                                     // bool(false) plus an E_NOTICE

go deeper

for a junior

Recall that text from outside may not be UTF-8 and that mb_convert_encoding converts it when you know the source encoding.

for a middle

Explain mb_check_encoding with an explicit encoding, the substitute character, and iconv's false return with the //IGNORE and //TRANSLIT suffixes.

for a senior

Design an import that converts from the declared encoding once, makes substitutions visible, measures them per record, and rejects feeds above a threshold.

for a principal

Own the contract with data suppliers: required encoding declarations, rejection rules and the single boundary where conversion happens.

## The boundary rule Inside a modern PHP application all text should be **UTF-8**. Anything that enters from outside, such as supplier feeds, uploaded CSV files, legacy databases or email, gets converted **once, at the boundary**, from an encoding you know. Afterwards every `mb_*` call, every template and every JSON response can assume UTF-8. The encoding should come from a declaration, in order of reliability: 1. the supplier's contract or feed specification; 2. the `charset` of an HTTP `Content-Type` header; 3. an encoding declaration inside the document, such as an XML prolog; 4. only then, detection as a fallback that you log. ## Validating UTF-8 `mb_check_encoding(array|string|null $value = null, ?string $encoding = null): bool` returns `true` when the bytes are valid in the given encoding. Details that matter: - Pass the encoding explicitly: `mb_check_encoding($s, 'UTF-8')`. - An **array** is validated recursively, keys and values, which suits decoded payloads. - Calling it without arguments, which used to check the whole request input, is **deprecated since PHP 8.1**. When invalid data must be kept rather than rejected, `mb_scrub($s, 'UTF-8')` converts the string to the same encoding and replaces invalid sequences with the substitute character. ## Converting with mbstring `mb_convert_encoding(array|string $string, string $to_encoding, array|string|null $from_encoding = null): array|string|false` converts between encodings mbstring knows. For Japanese suppliers the relevant names are: | Name | Meaning | |---|---| | `SJIS` | Shift_JIS as standardised | | `CP932` | Microsoft's Shift_JIS variant, typical of Windows-produced files | | `EUC-JP` | the Unix-world Japanese encoding | | `UTF-8` | the target | Bytes that are invalid in the source encoding are replaced by the **substitute character**, `?` by default. `mb_substitute_character(0xFFFD)` switches to U+FFFD REPLACEMENT CHARACTER, which makes damage visible and searchable instead of hiding it among real question marks. An unknown encoding name throws a `ValueError` since PHP 8.0. ## Converting with iconv `iconv(string $from_encoding, string $to_encoding, string $string): string|false` uses the system's iconv library. Its behaviour on problem input differs from mbstring: - By default, a character it cannot convert raises an `E_NOTICE` and the function returns **`false`**. - Appending `//IGNORE` to the target, as in `'ISO-8859-1//IGNORE'`, silently drops such characters. - Appending `//TRANSLIT` approximates them, for example `€` as `EUR`, but whether and how that works depends on the implementation reported by the `ICONV_IMPL` constant; some implementations ignore it. `false` is the dangerous result: stored as a string it becomes empty. Always check `iconv()` against `false` strictly before using its output. ## Detection is a heuristic `mb_detect_encoding(string $string, array|string|null $encodings = null, bool $strict = false): string|false` guesses from a candidate list. Since PHP 8.1 it uses heuristics to pick the most plausible candidate, which means short or ambiguous strings can be misidentified: many byte sequences are valid in several encodings at once. With `$strict = false` it may return the closest match even when the string is not valid in it. If you must detect, pass a short, ordered list, set `$strict` to `true`, and log every guess. ## A defensive import pipeline 1. Read the declared encoding for the feed. 2. Convert with `mb_convert_encoding($raw, 'UTF-8', $declared)`, with U+FFFD as the substitute character. 3. Verify with `mb_check_encoding($converted, 'UTF-8')`. 4. Count replacement characters per record; above a threshold, reject the record and report it to the supplier. 5. Normalize and trim the text fields, then store. ## What interviewers listen for The question separates candidates who have only ever seen UTF-8 from those who have integrated real data sources. Strong answers: - insist on a declared source encoding and treat detection as a logged fallback; - know that mbstring substitutes while `iconv()` fails with `false`, and handle both outcomes; - make substitutions visible with U+FFFD and measure them instead of shipping silent question marks; - mention the PHP 8 `ValueError` for unknown encoding names, which turns a typo such as `'SHIFTJIS'` into an immediate failure rather than garbage. Those details show the candidate has debugged a mojibake incident, not just read the function list.

  • What is the difference between mb_convert_encoding() and iconv() when the input contains bytes that cannot be converted?
    mbstring substitutes them with its substitute character, `?` unless changed with `mb_substitute_character()`, and returns a string. `iconv()` raises a notice and returns `false`, unless `//IGNORE` drops the characters or `//TRANSLIT` approximates them, and translit support depends on the platform's iconv implementation.
  • Why is mb_detect_encoding() a poor primary strategy for supplier feeds?
    Many byte sequences are valid in several encodings, so detection is a guess; since PHP 8.1 it ranks candidates heuristically, and without `strict` it can return an encoding the string is not even valid in. Use the declared encoding, and treat detection as a logged fallback.
  • What does calling mb_check_encoding() without arguments do in PHP 8.5?
    It is deprecated since PHP 8.1. It used to validate the whole request input. Pass the value and the encoding explicitly, for example `mb_check_encoding($_POST, 'UTF-8')`, which validates arrays recursively.

saying these in an interview costs you the question

  • mb_detect_encoding reliably identifies the encoding of any string.
  • iconv returns the partially converted string when it meets an illegal character.
  • //TRANSLIT behaves identically on every platform.
  • mb_convert_encoding throws an exception for undecodable bytes.
  • Converting text back and forth in each layer is safer than converting once at the boundary.