In PHP, why does json_encode() return false for some database rows, and which flags control how slashes and non-ASCII text are escaped?
answer
- input strings must be valid UTF-8
- Malformed UTF-8 characters, possibly incorrectly encoded
- INVALID_UTF8_SUBSTITUTE writes U+FFFD
- default output escapes / and non-ASCII
- UNESCAPED_SLASHES, UNESCAPED_UNICODE, PRETTY_PRINT
basics
~10 sjson_encode() requires valid UTF-8; a Latin-1 string makes it fail with JSON_ERROR_UTF8 and return false. By default it escapes / as / and non-ASCII as \uXXXX; JSON_UNESCAPED_SLASHES and JSON_UNESCAPED_UNICODE turn that off.
solid answer
~40 sJSON text is Unicode, so `json_encode()` requires every PHP string it encodes to be valid **UTF-8**. A row read over a connection using another character set, or a truncated multibyte string, contains invalid bytes, and the whole call returns `false` with `JSON_ERROR_UTF8` ("Malformed UTF-8 characters, possibly incorrectly encoded"); with `JSON_THROW_ON_ERROR` it throws instead. The real fix is correct encoding at the source; `JSON_INVALID_UTF8_SUBSTITUTE` (replace with U+FFFD) or `JSON_INVALID_UTF8_IGNORE` (drop the bytes) are fallbacks. Output escaping is separate: by default `/` becomes `\/` and every non-ASCII character becomes a `\uXXXX` escape. `JSON_UNESCAPED_SLASHES` and `JSON_UNESCAPED_UNICODE` write them literally, and `JSON_PRETTY_PRINT` adds newlines and four-space indentation. Flags combine with `|`.
code
php · 14 lines<?php
declare(strict_types=1);
$row = ['city' => "S\xE3o Paulo", 'url' => 'https://example.org/a']; // Latin-1 byte 0xE3
var_dump(json_encode($row)); // bool(false)
echo json_last_error_msg(), "\n"; // Malformed UTF-8 characters, possibly incorrectly encoded
echo json_encode($row, JSON_INVALID_UTF8_SUBSTITUTE), "\n";
// {"city":"S\ufffdo Paulo","url":"https:\/\/example.org\/a"}
$fixed = ['city' => 'São Paulo', 'url' => 'https://example.org/a'];
echo json_encode($fixed, JSON_UNESCAPED_SLASHES | JSON_UNESCAPED_UNICODE), "\n";
// {"city":"São Paulo","url":"https://example.org/a"}go deeper
Recall that json_encode() needs UTF-8 strings and returns false otherwise, and name the flags for unescaped slashes, unescaped Unicode and pretty printing.
Explain why one bad byte fails the whole call, how the INVALID_UTF8 flags differ, and why escaped and unescaped forms decode the same.
Trace empty API bodies to unchecked false returns and non-UTF-8 sources, and set a project-wide flag default including JSON_THROW_ON_ERROR.
Own the end-to-end encoding policy, from database connection charset to response flags, so encoding failures are prevented rather than masked.
## Why one bad byte fails the whole call `json_encode()` builds a JSON **text**, and JSON text is Unicode. PHP strings are byte strings, so the encoder checks each one as UTF-8 while escaping it. When it finds an invalid sequence, it records `JSON_ERROR_UTF8`, whose message is "Malformed UTF-8 characters, possibly incorrectly encoded", and: - without `JSON_THROW_ON_ERROR`, returns `false` for the **entire** value; - with `JSON_THROW_ON_ERROR`, throws a `JsonException` with that code. Typical sources of invalid bytes: 1. a database connection whose character set is not UTF-8, so `é` arrives as the single Latin-1 byte `0xE9`; 2. a string cut with a byte-based function in the middle of a multibyte character; 3. file contents in a legacy encoding read without conversion. A frequent secondary bug is `echo json_encode($rows);` with no error handling: `false` echoes as an empty string, and the client receives an empty body. ## Fixing it | Approach | Effect | When to use | |---|---|---| | fix the source encoding | strings are valid UTF-8 | always the real fix | | `JSON_INVALID_UTF8_SUBSTITUTE` | invalid bytes become U+FFFD (the replacement character) | logs, diagnostics, best-effort output | | `JSON_INVALID_UTF8_IGNORE` | invalid bytes are dropped | rarely; silently alters data | | `JSON_PARTIAL_OUTPUT_ON_ERROR` | an unencodable value becomes `null` | debugging only; loses data without an exception | Converting between encodings with the multibyte functions, and setting a database connection's charset, are separate topics; the JSON-level flags only decide what happens when bad bytes reach the encoder. Other values that fail encoding for the same reason: - `INF` and `NAN` floats: `JSON_ERROR_INF_OR_NAN`; - recursive structures: `JSON_ERROR_RECURSION`; - resources: `JSON_ERROR_UNSUPPORTED_TYPE`. ## Output escaping flags By default `json_encode()` is conservative about which characters it writes literally: | Flag | Default output for `/` and `é` | With the flag | |---|---|---| | none | `"a\/b"`, `"\u00e9"` | | | `JSON_UNESCAPED_SLASHES` | | `"a/b"` | | `JSON_UNESCAPED_UNICODE` | | `"é"` as UTF-8 bytes | | `JSON_PRETTY_PRINT` | one line | newlines, four-space indent | Both escaped and unescaped forms are valid JSON that decodes to the same string; the flags change size and readability, not meaning. Two details: - with `JSON_UNESCAPED_UNICODE`, the line terminators U+2028 and U+2029 are **still** escaped unless you also pass `JSON_UNESCAPED_LINE_TERMINATORS`; - when JSON is embedded inside an HTML `<script>` block, the escaped slash keeps `</script>` from appearing literally, and the `JSON_HEX_TAG`, `JSON_HEX_AMP`, `JSON_HEX_APOS` and `JSON_HEX_QUOT` flags escape `<`, `>`, `&`, `'` and `"` as `\u003C`-style sequences. Dropping escapes is fine for API bodies. ## Combining flags Flags are bit masks combined with `|`: ```php $json = json_encode( $payload, JSON_THROW_ON_ERROR | JSON_UNESCAPED_SLASHES | JSON_UNESCAPED_UNICODE | JSON_PRETTY_PRINT, ); ``` A named-argument call reads well when only flags change: `json_encode($payload, flags: JSON_THROW_ON_ERROR)`. ## A practical default - API bodies: `JSON_THROW_ON_ERROR | JSON_UNESCAPED_SLASHES | JSON_UNESCAPED_UNICODE`. - Human-read files and fixtures: add `JSON_PRETTY_PRINT`. - Inline `<script>` data: keep slashes escaped and add the `JSON_HEX_*` flags. ## Diagnosing an empty JSON response 1. Log `json_last_error_msg()` next to the failing `json_encode()` call, or switch it to `JSON_THROW_ON_ERROR` so the handler records the message. 2. If the message is the UTF-8 one, find the row: run `mb_check_encoding($value, 'UTF-8')` over the fields, or bisect the payload. 3. Fix the source encoding, then keep the throw flag so the next bad byte is loud rather than blank.
- Why is JSON_INVALID_UTF8_IGNORE a poor default for API output?It deletes the invalid bytes silently, so `São` from a Latin-1 source becomes `So` and nobody is told. The data is changed without an error. Fix the source encoding; if a fallback is needed, `JSON_INVALID_UTF8_SUBSTITUTE` at least leaves a visible replacement character.
- Does JSON_UNESCAPED_UNICODE change what a client decodes?No. `"\u00e9"` and a literal `é` decode to the same string. The flag makes the output shorter and readable, and it keeps U+2028 and U+2029 escaped unless `JSON_UNESCAPED_LINE_TERMINATORS` is also passed.
saying these in an interview costs you the question
- Believing json_encode() converts Latin-1 strings to UTF-8 automatically
- Echoing json_encode() output without checking for false or using the throw flag
- Thinking the escaped \/ form is invalid JSON
- Treating JSON_INVALID_UTF8_IGNORE as a safe fix for bad data
- Assuming JSON_UNESCAPED_UNICODE changes the decoded value