skip to content

In PHP, what is the difference between htmlspecialchars() and htmlentities(), and when, if ever, do you need htmlentities()?

level: juniorimportance: should knowfreq 50%

answer

  1. five characters vs every named entity
  2. same safety for markup characters
  3. é becomes é
  4. characters without names left as is
  5. pointless on a UTF-8 page

basics

~20 s

htmlspecialchars() converts only &, <, >, and the quotes; htmlentities() also converts every character that has a named HTML entity, such as é into é. Both protect markup equally, so on a UTF-8 page htmlspecialchars() is the standard choice.

solid answer

~40 s

Both functions take the same parameters and defaults and neutralise markup the same way: `&`, `<`, `>`, `"` and `'` become entities. `htmlentities()` additionally replaces every character that has a **named** HTML entity: `é` becomes `&eacute;`, `©` becomes `&copy;`, `€` becomes `&euro;`. That adds no security; it only changes how non-ASCII text is written. On a page served as UTF-8 the browser displays `é` directly, so `htmlspecialchars()` gives shorter, readable output with the same protection. `htmlentities()` also depends on the correct input charset to recognise those characters, and it does not make output pure ASCII, because characters without a named entity, most CJK text for example, pass through unchanged. It survives mainly in legacy code for pages that were not UTF-8.

code

php · 10 lines
php
<?php
declare(strict_types=1);

$title = "Café "Sol" & Mar <3";

echo htmlspecialchars($title, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8'), PHP_EOL;
// Café &quot;Sol&quot; &amp; Mar &lt;3

echo htmlentities($title, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8'), PHP_EOL;
// Caf&eacute; &quot;Sol&quot; &amp; Mar &lt;3

go deeper

for a junior

Recall that htmlspecialchars() encodes the five markup characters and htmlentities() also encodes characters with named entities, and that both protect markup equally.

for a middle

Explain why the extra conversions add no security, why htmlspecialchars() suits UTF-8 pages, and the charset sensitivity and named-entity gap of htmlentities().

for a senior

Show how you migrate legacy htmlentities() usage without breaking stored, already-encoded data, and standardise one escaping helper.

for a principal

Frame the choice as an encoding policy, UTF-8 end to end, so escaping only ever handles markup characters and never character-set conversion.

## Same signature, different scope The two functions share a signature and defaults: - `htmlspecialchars(string $string, int $flags = ENT_QUOTES | ENT_SUBSTITUTE | ENT_HTML401, ?string $encoding = null, bool $double_encode = true): string` - `htmlentities(string $string, int $flags = ENT_QUOTES | ENT_SUBSTITUTE | ENT_HTML401, ?string $encoding = null, bool $double_encode = true): string` They differ in how many characters they replace: | Input | `htmlspecialchars()` | `htmlentities()` | |---|---|---| | `<b>` | `&lt;b&gt;` | `&lt;b&gt;` | | `Tom & Jerry's` | `Tom &amp; Jerry&#039;s` | `Tom &amp; Jerry&#039;s` | | `Café São Paulo` | `Café São Paulo` | `Caf&eacute; S&atilde;o Paulo` | | `5 €` | `5 €` | `5 &euro;` | | `東京` | `東京` | `東京` (no named entity) | ## Security is identical What makes HTML output safe is encoding the characters that the parser treats as markup: `&`, `<`, `>` and the quotes. Both functions do exactly that. The extra conversions `htmlentities()` performs, accented letters, currency signs, typographic symbols, have no special meaning to the HTML parser, so converting them neither adds nor removes protection. A candidate who says `htmlentities()` is "more secure" is describing a difference that does not exist. ## Why htmlspecialchars() is the default choice - **UTF-8 pages display the characters directly.** A travel site served as UTF-8 shows `São Paulo` correctly without any entity, and a review written in Portuguese stays readable in the page source. - **Smaller, simpler output.** Every converted character costs several bytes, and the page source becomes harder to read and debug. - **Less dependence on the charset parameter.** `htmlentities()` must decode the input correctly to know which characters have names. With a wrong `$encoding`, multibyte characters are misread and turned into the wrong entities, while `htmlspecialchars()` only rewrites ASCII characters, which sit at the same byte positions in UTF-8 and the common single-byte encodings. - **No false promise of ASCII output.** Only characters with a *named* entity are converted. Emoji, most Asian scripts and many symbols pass through unchanged, so `htmlentities()` does not make arbitrary text safe for an ASCII-only channel. ## When htmlentities() still appears - **Legacy pages in a single-byte charset** such as ISO-8859-1, where characters outside that charset could not be written directly. Named entities let such pages show a limited set of extra characters. - **Old code bases** that used it by habit. Replacing it with `htmlspecialchars()` changes no security property, only the bytes on the page, but stored data that was entity-encoded on input must be handled separately. - **Output channels that dislike raw non-ASCII**, where named entities are preferred for the characters that have them; the CJK gap above still applies. ## Decoding counterparts Each function has an inverse: `htmlspecialchars_decode()` reverses only the special characters, and `html_entity_decode()` reverses all named and numeric entities. Needing to call either on your own output is usually a sign that text was escaped at input and is now being un-escaped for a non-HTML use, which is the pattern to fix at its source. ## Migrating from htmlentities() Replacing `htmlentities()` with `htmlspecialchars()` is safe for security, but it changes bytes on the page, so do it deliberately: 1. Confirm the pages are served as UTF-8 and the templates are saved as UTF-8, so characters like `é` display correctly without entities. 2. Replace the calls through one helper, keeping `ENT_QUOTES | ENT_SUBSTITUTE` and `'UTF-8'`. 3. Check stored data: if any column was entity-encoded on input with `htmlentities()`, decode it once with `html_entity_decode()` in a migration rather than leaving it double-encoded on output. 4. Compare rendered pages before and after for a sample of non-ASCII content, such as city names and reviews in several languages. ## The interview answer State the scope difference, stress that security is the same because markup characters are handled identically, and recommend `htmlspecialchars()` with `ENT_QUOTES | ENT_SUBSTITUTE` and `'UTF-8'` for UTF-8 pages. Mention the charset sensitivity and the "only named entities" limit as the reasons `htmlentities()` is not an upgrade.

  • Does htmlentities() turn any UTF-8 text into pure ASCII?
    No. It only converts characters that have a named HTML entity, such as accented Latin letters and common symbols. Characters without a name, including most Chinese, Japanese and Korean text and emoji, are left as they are, so the output can still contain non-ASCII bytes.
  • What goes wrong if htmlentities() is given the wrong encoding argument?
    It interprets the bytes in the wrong charset, so multibyte characters are split and converted into the wrong entities, producing visible mojibake. `htmlspecialchars()` is far less sensitive because it only rewrites ASCII characters, which sit at the same byte positions in UTF-8 and the common single-byte encodings.

saying these in an interview costs you the question

  • htmlentities() is more secure than htmlspecialchars()
  • htmlspecialchars() leaves < and > unescaped
  • htmlentities() converts every non-ASCII character to an entity
  • UTF-8 pages need htmlentities() to show accented letters
  • The two functions have different default flags