skip to content

Why is PHP's strip_tags() not a safe way to clean user reviews for display, and what goes wrong with its allowed_tags argument?

level: middleimportance: should knowfreq 40%

answer

  1. the manual says not for XSS
  2. allowed tags keep every attribute
  3. a stray < eats the rest of the text
  4. quotes and & left unescaped
  5. allowed_tags as an array since 7.4

basics

~20 s

strip_tags() removes things that look like tags but does not escape what remains, keeps every attribute on allowed tags, and can delete ordinary text after a stray <. The PHP manual says not to use it against XSS.

solid answer

~50 s

`strip_tags($review)` deletes anything that looks like a tag, plus HTML comments and PHP tags. It is not an encoder: the remaining text still contains `&`, quotes and possibly `>`, so it still needs `htmlspecialchars()` on output. Its parser does not validate HTML, so a `<` followed by a non-space character starts a "tag" that runs to the next `>`: `Loved it <3 would book again` becomes `Loved it `. The second argument, a string like `'<b><i>'` or, since PHP 7.4, an array like `['b', 'i']`, lets tags through **with all their attributes**, so `<b onmouseover="...">` survives intact. The PHP manual warns against using it to prevent XSS. For plain-text reviews, store the raw text and escape it; for reviews that really need formatting, use a real HTML sanitizer with an element and attribute allow-list, or a markup language such as Markdown rendered by a safe library.

code

php · 12 lines
php
<?php
declare(strict_types=1);

$review = 'Loved it <3 would book again';
echo strip_tags($review), PHP_EOL;                 // "Loved it "

$review = '<b onmouseover="alert(1)">Great</b> pool';
echo strip_tags($review, ['b']), PHP_EOL;          // attribute survives intact

// Plain-text reviews: escape, then turn newlines into <br>
$safe = nl2br(htmlspecialchars($review, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8'));
echo $safe, PHP_EOL;                               // markup shown as text

go deeper

for a junior

Recall that strip_tags() is not an escaping function, and that user reviews are shown with htmlspecialchars() at output.

for a middle

Explain how its parser treats a < before a non-space character, why allowed tags keep all attributes, and the string and array forms of allowed_tags.

for a senior

Replace strip_tags()-based cleaning with escape-at-output for plain text and a real sanitizer for rich text, and plan how to handle data already mangled by it.

for a principal

Decide product-wide whether user content may carry formatting at all, since that choice determines whether an HTML sanitizer must be owned and maintained.

## What strip_tags() does `strip_tags(string $string, array|string|null $allowed_tags = null): string` walks through the string and removes: - anything that looks like an HTML tag, from `<` to the matching `>`; - HTML comments; - PHP tags. The optional second argument lists tags to keep, either as a string, `'<b><i>'`, or since PHP 7.4 as an array, `['b', 'i']`. HTML comments and PHP tags are always removed, whatever you allow. ## Pitfall 1: it is not escaping After stripping, the text still contains characters that matter in HTML: `&`, both quote characters and sometimes a bare `>`. If the result is echoed into an attribute, a quote in a review can still close the value. If it is echoed as text, `&lt;` typed by the user becomes a literal `<` when rendered. So a correct pipeline still ends in `htmlspecialchars()` at output, at which point `strip_tags()` has added nothing but data loss. The PHP manual is explicit: the function should not be used to try to prevent XSS; use `htmlspecialchars()` or other context-appropriate functions instead. ## Pitfall 2: it eats legitimate text `strip_tags()` does not validate HTML, and the manual warns that partial or broken tags can remove more text than expected. The rule its parser follows: a `<` immediately followed by a non-space character starts a tag, which lasts until the next `>`, or to the end of the string if there is none. | Review text | Result | |---|---| | `Rooms were 5 < 10 minutes from the beach` | unchanged: `<` followed by a space is kept | | `Loved it <3 would book again` | `Loved it ` | | `Price was <100 EUR, view was > expected` | `Price was expected` | On a travel site, where reviews compare prices and distances, this silently mangles honest content, and the author never learns why. ## Pitfall 3: allowed tags keep their attributes The allow-list works on tag names only. Every attribute of an allowed tag passes through untouched. The manual warns specifically about `style` and `onmouseover`: - `<b onmouseover="stealCookies()">great</b>` survives `strip_tags($s, ['b'])` intact; - `<a href="javascript:...">` survives if `a` is allowed; - `<i style="position:fixed;top:0;left:0;width:100%;height:100%">` lets a review cover the page. Allowing "just a few harmless tags" therefore allows arbitrary script through their attributes. ## Pitfall 4: allowed_tags syntax traps - In the string form each tag is written with angle brackets, `'<b><i>'`; `'b,i'` does not work as intended. - Tag names longer than 1023 bytes are treated as invalid regardless of the allow-list. ## What to use instead The right tool depends on whether reviews may contain formatting at all: 1. **Plain text (the usual case).** Store the review exactly as typed. Echo it with `htmlspecialchars()`, and convert newlines to `<br>` *after* escaping, for example with `nl2br(e($review))`, so line breaks show without allowing any markup. 2. **Limited formatting.** Accept a lightweight markup such as Markdown and render it with a library configured to disallow raw HTML, then pass the output through an HTML sanitizer. 3. **Rich HTML.** Use a dedicated HTML sanitizer that parses the document and applies an allow-list of elements **and** attributes, with URL scheme checks for `href` and `src`. ## Cleaning up data that was already stripped Applications that ran `strip_tags()` on input for years hold reviews that were silently truncated. Fixing the code does not restore them, so plan the migration: - stop stripping on input and start escaping on output in the same change, so no window exists with neither; - keep the stored text as it is: it contains no tags, so escaping it on output is harmless; - if the original submissions survive elsewhere, such as in request logs or a moderation queue, they cannot simply be re-imported, because they may contain the markup that was stripped; they need the new escaping or sanitizing path. In none of these does `strip_tags()` play a security role. It remains useful for non-security jobs, such as producing a plain-text excerpt of trusted HTML for an email subject or a search index, and even then its output needs escaping wherever it is shown.

  • Why is nl2br() applied after htmlspecialchars() and not before?
    `nl2br()` inserts real `<br />` tags. If it ran first, `htmlspecialchars()` would turn those tags into visible `&lt;br /&gt;` text. Escaping first and then adding the line breaks means the only markup in the output is the markup you added deliberately.
  • Is strip_tags() ever the right tool?
    Yes, for non-security tasks on text you trust, such as a plain-text excerpt of your own HTML for an email subject or a search index. Even then its output is text that needs escaping wherever it is displayed, and it should not be run on user input as a safety measure.

saying these in an interview costs you the question

  • strip_tags() makes user input safe to echo into HTML
  • Tags in allowed_tags have their dangerous attributes removed
  • strip_tags() only removes text between well-formed tags
  • strip_tags() also escapes quotes and ampersands
  • Allowing <b> and <i> in strip_tags() is harmless