skip to content

Your app must render user-submitted rich text as real HTML, so textContent is not an option. What does a sanitizer such as DOMPurify actually do, and at what point in the flow must it run?

level: middleimportance: should knowfreq 42%

answer

  1. parse the tree, don't match strings
  2. allowlist elements, attributes, schemes
  3. clean on the way out
  4. serialise-then-reparse is the trap
  5. Unsafe in the name means unsafe

basics

~20 s

A sanitizer parses the markup into an inert DOM, walks every node, and removes anything outside an allowlist of elements, attributes and URL schemes — event handlers first. Run it at the sink, immediately before insertion, not once when the value is stored.

solid answer

~50 s

`DOMPurify.sanitize(dirty)` does not pattern-match strings; it parses the input into an inert document, walks the resulting tree, and deletes every element and attribute that is not on its allowlist — all `on*` handlers, `javascript:` values in `href`/`src`, and unknown elements go. You constrain it further with `ALLOWED_TAGS` and `ALLOWED_ATTR`. The critical rule is *where* it runs: at the sink, right before the string reaches `innerHTML`, on the same page load. Sanitizing at input time and storing the result is wrong, because the same stored value may later be rendered in a different context, and because sanitizer fixes for newly discovered bypasses only help data sanitized after the upgrade. There is a standard alternative, `Element.setHTML()` from the Sanitizer API, but cross-engine availability is still uneven as of 2026 — note that the similarly named `setHTMLUnsafe()` does no sanitizing at all.

code

javascript · 13 lines
javascript
import DOMPurify from 'dompurify';

// Store the raw value; clean it here, at the sink, on every render.
function renderBio(rawHtml) {
  const clean = DOMPurify.sanitize(rawHtml, {
    ALLOWED_TAGS: ['b', 'i', 'em', 'strong', 'a', 'p', 'ul', 'ol', 'li'],
    ALLOWED_ATTR: ['href', 'title'],
  });
  document.querySelector('#bio').innerHTML = clean;
}

renderBio('<p onclick="steal()">hi <a href="javascript:alert(1)">x</a></p>');
// renders: <p>hi <a>x</a></p>

go deeper

for a junior

Know that you only reach for a sanitizer when the value must remain real markup, and that DOMPurify.sanitize() with an allowlist runs right before the string is assigned to innerHTML.

for a middle

Explain the mechanism — parse into an inert DOM, walk it, drop everything off the allowlist including on* handlers and javascript: URLs — and why cleaning at render time rather than at storage time is what makes sanitizer upgrades retroactive.

for a senior

Show judgment about the allowlist itself: which elements you refuse to re-admit under product pressure, how you handle mutation XSS by avoiding the serialise-reparse round trip, and how you keep the library current as bypasses are published.

for a principal

Own the position that per-call-site discipline decays: argue for a single rendering chokepoint plus an enforcement layer that makes unsanitized strings impossible at any sink, and be able to state the migration cost that buys.

## Escaping versus sanitizing These are different jobs and the choice is decided by what the value *is*. - If the value is **text** that happens to contain angle brackets, you do not sanitize — you assign `textContent` and there is no parser in the path at all. - If the value is **markup you must preserve** (a comment with bold and links, an email body, editor output), you cannot escape it without destroying it. You sanitize: keep the safe subset of markup, delete the rest. Reaching for a sanitizer when `textContent` would do is a smell; it adds a dependency and a bypass surface to a problem that had a structural fix. ## What a sanitizer actually does A credible HTML sanitizer is not a regular expression over a string. It: 1. **Parses** the input with the browser's own HTML parser into an inert container — a `<template>` content fragment or a document created by `DOMParser`, where nothing loads and no handler fires. 2. **Walks** the resulting tree node by node. 3. **Applies an allowlist**: an element not on the list is dropped (or unwrapped, keeping its children); an attribute not on the list is removed. Every `on*` attribute goes unconditionally. 4. **Checks URL-bearing attributes** — `href`, `src`, `action`, `xlink:href` — against permitted schemes, so `javascript:` and unexpected `data:` values are stripped. 5. **Serialises** the cleaned tree back to a string (or hands you the DOM node directly). The reason it parses rather than pattern-matches is that only a parser knows what the browser will conclude the markup means. Any filter working on raw text is guessing at another implementation's behaviour, which is the same losing position as a blocklist. ## Configuring it honestly ```js const clean = DOMPurify.sanitize(userHtml, { ALLOWED_TAGS: ['b', 'i', 'em', 'strong', 'a', 'p', 'ul', 'ol', 'li'], ALLOWED_ATTR: ['href', 'title'], }); ``` A narrow allowlist is the whole point: the smaller the accepted subset, the less parser surface an attacker can aim at. Loosening the configuration to make a rendering bug go away is how these deployments decay. Watch particularly for re-admitting `style` (CSS is its own injection surface), `iframe`, `svg`, `math`, and any `data-*` your own framework interprets as instructions. ## Why the sink, and not the input Sanitizing once on write and storing the cleaned string looks efficient and is a recurring mistake: - **Context is not fixed at write time.** The same record may later be rendered into an email, a PDF, an attribute, or a different framework — each with different rules. Cleaning for one destination bakes in an assumption the storage layer cannot see. - **Bypasses are found continuously.** Sanitizers ship fixes; a value cleaned by last year's version stays in your database in its last-year-clean form. Sanitizing at render means every upgrade retroactively protects all stored data. - **You lose the original.** Once you have destroyed the user's input you can never re-render it correctly under a new policy. So store the raw value (with ordinary input validation for size and shape), and sanitize on the way out, next to the sink. ## Mutation XSS The subtle failure mode is **mXSS**: a string that is genuinely clean as the sanitizer parsed it, but which the browser re-parses *differently* when you assign it to `innerHTML`. This arises from serialise-then-reparse round trips, especially around `<template>`, `<noscript>`, and namespace switches into `<svg>` or `<math>` where HTML's parsing rules change mid-tree. Real sanitizers have shipped fixes for exactly this class. Two practical mitigations: keep the allowlist away from the elements that trigger namespace switches unless you need them, and prefer handing the sanitized **DOM node** to the page instead of a string, so there is no second parse. DOMPurify's `RETURN_DOM`/`RETURN_DOM_FRAGMENT` options exist for this. ## The platform alternative The Sanitizer API standardises this in the browser: `Element.setHTML(input, options)` parses and sanitizes in one step with a safe default configuration, with no library to keep updated. As of 2026 cross-engine availability is still uneven, so a library remains the portable choice for most products. Be careful with the adjacent names. `Element.setHTMLUnsafe()` and `Document.parseHTMLUnsafe()` exist to parse markup *including* declarative shadow roots — they perform **no sanitization whatsoever**. Their `Unsafe` suffix is the entire API documentation on that point, and treating them as the modern `innerHTML` replacement is a real bug seen in review. ## What sanitizing does not do It cleans the string you hand it. It does nothing for a `javascript:` URL your own code assigns to `href`, nothing for `eval`, and nothing for a sink you forgot. That is the argument for pairing it with an enforcement layer that makes every sink refuse plain strings, rather than trusting that every call site remembered to call the sanitizer.

  • What is mutation XSS and why does returning a DOM node help?
    mXSS is markup that is clean as the sanitizer parsed it but means something different when the browser re-parses the serialised string — typically around template, noscript, or a namespace switch into svg or math. Returning a DOM node with `RETURN_DOM` and inserting it directly removes the second parse, so there is no gap between the two interpretations.
  • Why not sanitize once when the comment is saved and store the clean version?
    Because the destination is unknown at write time — the same record may later be rendered into an attribute, an email, or another framework with different rules — and because sanitizer bypass fixes only protect data cleaned after the upgrade. Store the raw value and sanitize on every render, so an upgrade retroactively covers everything.
  • Is Element.setHTMLUnsafe() a safer modern replacement for innerHTML?
    No, and the name says so. `setHTMLUnsafe()` parses markup without sanitizing; it exists so that declarative shadow roots in the string are honoured, which plain `innerHTML` ignores. The sanitizing member of that family is `Element.setHTML()`. Confusing the two swaps a known sink for an equally dangerous one.

saying these in an interview costs you the question

  • "A regex that strips <script> tags is a sanitizer"
  • "Sanitize on input and store the clean copy"
  • "setHTMLUnsafe is the modern safe innerHTML"
  • "Allowing style attributes is harmless, it is only CSS"
  • "If we sanitize, we no longer need to check URLs we assign"

context