Cross-site scripting is commonly split into reflected, stored and DOM-based forms. Explain what actually distinguishes them at the level of where untrusted data becomes code, and why that distinction changes which defences and which detection can possibly work.
answer
- axis 1: transition on server or in browser
- axis 2: request-carried or persisted
- stored → no delivery step, hits admins, persists after the fix
- fragment never sent → server logs blind
- mXSS = serialise + reparse creates a tree the sanitizer never saw
basics
~20 sThe three differ by where the data-to-code transition happens and how the data got there. Reflected and stored transitions occur while the server builds the response; the DOM-based transition occurs in the browser, in client code writing to a script-executing sink. Server-side encoding cannot fix the third, and server-side logs and filters never see it.
solid answer
~60 sFrame it as **source → sink**, plus the storage hop. - **Reflected**: the source is the current request, the sink is the server's response construction. One victim per delivered link. - **Stored**: the same sink, but the source was persisted earlier and is now served to everyone who views it — so it is self-propagating and the blast radius includes privileged viewers such as administrators. - **DOM-based**: the transition happens entirely in the browser. Client code reads a source (`location.search`, `location.hash`, `postMessage` data, a value the server returned safely) and writes it into an execution sink such as `innerHTML`, `document.write`, `eval`, or a framework's raw-HTML escape hatch. The consequences are asymmetric. Server-side output encoding is decisive for the first two and irrelevant to the third. If the source is the URL fragment, it is never transmitted, so server logs, request filters and network detection are structurally blind — the only evidence is client-side. Stored form additionally requires that you encode at render time and keep the stored value raw, so the same record is safe in every sink.
go deeper
Say reflected comes from the current request, stored comes from saved data served to everyone, and DOM-based happens in browser code writing untrusted data into an executing sink.
Explain that server-side encoding fixes the first two and not the third, and that stored issues persist in the data after the code is fixed.
Use the source-to-sink frame, name real client sources and sinks, note that fragment-sourced payloads never reach the server so detection is blind, and raise the sanitizer round-trip problem.
Turn it into coverage strategy: which layer owns each transition, how client-side sinks are constrained and reviewed, and how stored-data remediation is planned as a separate workstream from the code fix.
## The distinction that matters The taxonomy is often taught as three unrelated attack types. It is better understood as two independent axes: - **Where does data become code?** In the server's response construction, or in the browser after the page has loaded. - **How did the data get there?** Carried in the request that triggers the render, or persisted earlier and served to other users. Reflected = server-side transition, request-carried. Stored = server-side transition, persisted. DOM-based = client-side transition (and it can be either request-carried or persisted, which is why it is really the other axis). ## Reflected The server takes a value from the current request — a query parameter, a form field, sometimes a header — and interpolates it into the response without encoding for the position it lands in. Exploitation needs delivery: the victim must be induced to issue the crafted request. Defence is ordinary context-correct encoding at the sink. Detection is comparatively easy because the payload traverses the network to the server and appears in the response body. ## Stored The same server-side transition, but the value was written earlier — a profile field, a comment, a filename, a support-ticket body, even a value harvested from a log or an imported file. Three consequences follow. It fires without any delivery step, so every viewer is a victim. It reaches viewers the attacker could not otherwise target, notably administrators looking at moderation or support screens, which is why stored XSS so often becomes privilege escalation. And it persists after the code is fixed: remediation includes finding and cleaning the stored records, which is a data problem, not just a code one. The stored case also settles the encode-when argument. The record may later be rendered into a page, serialised into JSON, exported to a spreadsheet, put in an email, or written to a log viewer. Each is a different grammar. So the record must be stored **raw** and encoded at each sink; encoding it on the way in picks one grammar for a value that will meet several, and produces double-encoded, unsearchable data. ## DOM-based Here the server may have done everything right. The page ships correct markup, and then client code reads a **source** and writes it into a **sink** that executes. - Typical sources: `location.search`, `location.hash`, `document.referrer`, `window.name`, `postMessage` payloads, values from an API response, values from storage. - Typical sinks: `innerHTML` / `outerHTML`, `document.write`, `insertAdjacentHTML`, `eval` and its relatives, assigning to `location`, setting an event-handler property from a string, injecting a `<script src>`, and framework escape hatches that render raw HTML. Two structural consequences. First, server-side output encoding cannot help, because the server is not building the dangerous string. Second, when the source is the **URL fragment**, the value is never sent to the server at all — so server logs, request-inspecting filters and any network-side detection are blind by construction. Investigating a DOM-based issue means client-side taint analysis: enumerate sources, enumerate sinks, and prove no flow between them. The fix is the same ladder applied client-side — prefer structural APIs (`textContent`, attribute setters, creating nodes) over building markup strings, and treat every raw-HTML escape hatch as a reviewed exception. ## The sanitizer trap When an application must accept rich text, it sanitizes HTML: parse, keep an allow-list of elements and attributes, drop the rest. This works only if the parser used to sanitize agrees exactly with the parser that will later consume the output. Where they disagree — because the sanitized tree is serialised back to a string and re-parsed by a browser that applies different quirks, or because the markup is placed into a context with different parsing rules such as inside SVG or a foreign-content element — the round trip can produce a tree that was not in the sanitized one. That is **mutation XSS**: nothing in the sanitizer's model was violated; the serialise-and-reparse step created new structure. The lesson is a ladder lesson. Sanitizing is the escaping rung on a grammar that is exceptionally hard to model, so it must be done by a maintained library operating on the consumer's own parser and applied as close to the consumption point as possible — never by a hand-rolled tag filter, and never with the sanitized result re-parsed in a different context than the one it was checked for. The structurally superior answer, where the product allows it, is to accept a restricted non-HTML input format and render it to markup yourself, so no attacker-authored markup is ever parsed. ## What the distinction buys you in practice It tells you where to look and what evidence can exist. A report with no matching server log entry does not mean the report is false — it may mean the payload lived in the fragment. A fix in the template layer does not close a client-side sink. And a stored issue is not remediated when the code is deployed; it is remediated when the data is clean.
- Your template layer encodes correctly everywhere and you still have XSS. Where do you look?At client-side flows. Enumerate sources — `location.search` and `location.hash`, `document.referrer`, `window.name`, `postMessage` handlers, API responses, storage — and the sinks they reach: `innerHTML`, `document.write`, `eval`, location assignment, `<script src>` injection and raw-HTML escape hatches. Also check anywhere the server emits data into a script literal or a JSON blob inline in the page, since that nests grammars and defeats the ordinary HTML escaper.
- Why is stored XSS often rated higher than reflected even with an identical payload?There is no delivery step, so every viewer is exposed rather than only those who follow a crafted link; it commonly reaches privileged viewers such as administrators in moderation or support screens, turning it into privilege escalation; and it survives the code fix, because the payload is in the data and must be found and cleaned separately.
saying these in an interview costs you the question
- Believing server-side output encoding covers DOM-based issues.
- Expecting to find every payload in server logs, when fragment-sourced values are never transmitted.
- Treating the three forms as unrelated bug types rather than as a source/sink and storage distinction.
- Sanitizing HTML on input and storing the result, which fixes one sink and corrupts the value for the others.
- Trusting a hand-rolled tag filter, which cannot account for the consumer's parser and mutation on re-parse.