skip to content

In the browser DOM, which APIs turn a plain string into live markup, and why does assigning "<script>alert(1)</script>" to an element's innerHTML not run the script while "<img src=x onerror=alert(1)>" does?

level: juniorimportance: must knowfreq 72%

answer

  1. string in, parser out
  2. two families: markup and source
  3. script tags arrive already started
  4. the onerror attribute is unaffected
  5. textContent creates one text node

basics

~20 s

innerHTML, outerHTML, insertAdjacentHTML and document.write parse a string as HTML, so untrusted data becomes markup. HTML fragment parsing marks script elements inserted that way as already-started, so they never run, but event-handler attributes such as onerror still fire. Use textContent instead.

solid answer

~40 s

A DOM XSS sink is any API that hands a string to a parser that can produce executable constructs. The HTML-parsing sinks are `Element.innerHTML`, `Element.outerHTML`, `Element.insertAdjacentHTML()`, `document.write()`/`writeln()`, plus the newer `Element.setHTMLUnsafe()` and `Document.parseHTMLUnsafe()`. `<script>` from innerHTML is inert because the HTML fragment-parsing algorithm marks those script elements as already-started, so inserting them never triggers execution — but that rule covers only script elements. Event-handler content attributes are compiled normally, so `<img src=x onerror=...>` fires when the bogus image fails to load, and `<svg onload=...>` fires immediately. That is why "innerHTML is safe because script tags don't run" is wrong. When the value is text, assign `textContent`, which creates one text node and invokes no parser at all.

code

javascript · 10 lines
javascript
const el = document.querySelector('#out');

// Script elements inserted by innerHTML are marked already-started: nothing runs.
el.innerHTML = '<script>alert(1)<\/script>';

// Event-handler attributes are unaffected by that rule: this one fires.
el.innerHTML = '<img src=x onerror="alert(1)">';

// textContent invokes no parser: the markup is displayed as literal text.
el.textContent = '<img src=x onerror="alert(1)">';

go deeper

for a junior

Be able to name innerHTML, outerHTML, insertAdjacentHTML and document.write as the APIs that parse markup, and say that textContent is the default for displaying user text.

for a middle

Explain the mechanics: fragment parsing marks inserted script elements already-started, while event-handler attributes such as onerror and onload are compiled normally, so the inertness of script tags proves nothing about safety.

for a senior

Show how you would find these sinks across a real codebase and decide per call site whether the value is text, a URL, or markup that genuinely needs a sanitizer — and where you would place enforcement so new sinks cannot appear unnoticed.

for a principal

Own the argument that a sink audit does not scale: articulate why the goal is an architecture where raw strings cannot reach a parser at all, and what that costs teams in templating conventions and review load.

## What makes an API a sink A *sink* is a browser API that accepts a string and passes it to a parser capable of producing executable constructs. The DOM has two families: APIs that parse a string as **HTML markup**, and APIs that parse a string as **JavaScript source**. When attacker-influenced data reaches either one, script runs in your page's origin — with full access to the DOM, cookies readable from script, and any token your app holds. The defect lives entirely in client-side code, so the server may never see the offending value. ## The HTML-parsing sinks These all run the HTML parser over the string you give them and splice the resulting nodes into the document: - `Element.innerHTML` (setter) - `Element.outerHTML` (setter) - `Element.insertAdjacentHTML(position, html)` - `document.write()` and `document.writeln()` - `Element.setHTMLUnsafe()` and `Document.parseHTMLUnsafe()` — the newer additions whose names carry the warning `DOMParser.prototype.parseFromString(str, "text/html")` also parses markup, into an inert document; nothing executes there, but adopting those nodes into the live document can fire handlers, so it is not a laundering step. Attributes are a second doorway. `element.setAttribute("onclick", value)` compiles `value` into a real event handler, because event-handler *content attributes* are code by definition. (Assigning a string to the IDL property, `el.onclick = "alert(1)"`, does nothing — the property expects a function.) Separately there are the JavaScript-source sinks — `eval()`, the `Function` constructor, and `setTimeout`/`setInterval` when handed a string instead of a function. ## Why the script tag stays silent When you set `innerHTML`, the browser uses the HTML **fragment parsing algorithm**. Script elements produced by that algorithm are created with their "already started" flag set, so when they are inserted into a document the normal "prepare the script element" path is skipped and nothing executes. Build the same element imperatively and it runs fine: ```js const s = document.createElement('script'); s.textContent = 'alert(1)'; document.body.appendChild(s); // this DOES run ``` So the inertness is a narrow rule about one element type, not a property of `innerHTML`. ## Why the img still fires Nothing suppresses event-handler attributes. `<img src=x onerror=alert(1)>` asks the browser to load a resource named `x`, the load fails, an `error` event fires at the image, and the compiled handler runs. `<svg onload=alert(1)>` needs no failed load at all. Other shapes in the same family include `<body onload>`, `<iframe onload>`, `<details ontoggle>`, and `<input autofocus onfocus=...>`. This is why payload lists for DOM XSS almost never contain a `<script>` tag. ```js const el = document.querySelector('#out'); el.innerHTML = '<script>alert(1)<\/script>'; // silent el.innerHTML = '<img src=x onerror="alert(1)">'; // alerts el.textContent = '<img src=x onerror="alert(1)">'; // shows the text literally ``` ## The safe defaults - **Text is text.** `Node.textContent` sets a single text node; there is no parser in the path, so `<`, `>` and `&` are displayed rather than interpreted. `append()` and `prepend()` with a string do the same. - **Build nodes, don't build markup.** `document.createElement()` plus `textContent` and validated attribute values is verbose but has no sink in it. - **Attributes need their own thought.** `setAttribute` is fine for `class` or `data-*`, but URL-bearing attributes (`href`, `src`, `formaction`) and any `on*` attribute are sinks in their own right. - **If you genuinely need markup**, that is what a sanitizer is for — the string still has to be cleaned before it reaches the parser, and the enforcement layer that guarantees this at every sink is Trusted Types. ## Where this bites in practice The realistic bug is not a demo payload; it is a template helper that interpolates a display name, an error message echoed from an API response, or a "render markdown" path that concatenates HTML. Grepping for `innerHTML` finds most of it, and that grep is the cheap first audit anyone can run on a frontend codebase.

  • If innerHTML is a sink, why is textContent not one?
    Because `textContent` does not parse. Assigning to it replaces the node's children with a single text node whose data is exactly your string, so `<`, `>` and `&` are characters to display rather than syntax. No HTML parser and no script compiler is ever invoked, which is what removes the whole bug class rather than filtering it.
  • Are there sinks that take JavaScript source rather than markup?
    Yes. `eval()`, the `Function` constructor, and `setTimeout`/`setInterval` when passed a string all compile their argument as script. `element.setAttribute("onclick", value)` belongs there too, since an event-handler content attribute is compiled into a handler. Passing a real function to the timers, and never building handler attributes from data, closes those.
  • Does document.write behave differently from innerHTML?
    It is worse in practice. `document.write()` streams into the parser at the current insertion point, so a `<script>` it writes during initial parsing *does* execute, and calling it after load implicitly wipes the document. It is a sink with fewer guardrails than `innerHTML`, which is why modern code avoids it entirely.

saying these in an interview costs you the question

  • "innerHTML is safe because script tags don't execute"
  • "Escaping only quotes and angle brackets covers every case"
  • "There is no XSS if the server never echoes the value"
  • "setHTMLUnsafe sanitizes because it is the newer API"
  • "Only a <script> tag can execute injected code"

context