A team strips HTML from every fetched page — why does an attacker's planted span still arrive intact?
answer
- who was the sanitiser protecting?
- tags out, text nodes untouched
- the payload is prose, not markup
- browser-safe is not model-safe
basics
~20 sStripping HTML removes tags, scripts and attributes and keeps the text — and the text is the injection. A sanitiser built to stop a browser executing markup is a no-op against natural-language sentences a model reads as directive.
solid answer
~50 sAn HTML sanitiser carries a browser threat model: it exists so a page cannot execute script, fire an event handler, or point at a dangerous URL inside somebody's browser. It works by discarding markup and preserving text content, because the text is what the page was for. A span an attacker plants on a page they published is ordinary prose sitting in an ordinary text node, so the scraper, the DOM parse, the sanitiser and the HTML-to-text converter each hand it forward unchanged. What sanitising establishes is that markup a browser could have executed was removed; it establishes nothing about what the surviving sentences ask a model to do. The error is treating two consumers as one — the browser is attacked through structure, the model is addressed through content — and the control was written for the first.
go deeper
Be ready to say in one sentence what an HTML sanitiser removes and what it keeps, and to notice that a planted instruction lives in the part it keeps. Do not offer fixes; the interviewer is checking whether you can name the consumer each control was written for.
Expect to walk the fetch chain stage by stage — parse, boilerplate removal, sanitise, convert — and say for each what it discards and what it necessarily preserves. Explain why a prose span is invariant under all of them rather than merely lucky.
Show you can state what a clean sanitise result licenses and what it does not, and resist both errors: that stripping solves injection, and that stripping is therefore worthless. Name the real limits — a truncation budget, a dropped boilerplate region — instead of the sanitiser.
Own the framing that one word covering two threat models is how this survives review after review. Be able to say what evidence a team would need before claiming a text-cleaning stage bears on injection at all, and what class of finding it can never speak to.
## Two consumers, two threat models A browsing or research agent that fetches pages nobody showed it — following links off a search step, reading whatever it lands on, returning a cited brief — puts a chain of text-processing stages between a stranger's web server and a model's context window. A typical chain is: fetch the bytes, parse them into a DOM, drop boilerplate regions (navigation, footer, sidebars), sanitise, convert HTML to plain text or markdown, then chunk. Teams often describe this chain as *cleaning* the page, and from there it is a short step to `we strip HTML, so injection cannot survive`. That conclusion is wrong, and the reason is worth stating precisely rather than as a slogan. An HTML sanitiser was designed for a browser. Its adversary is a page that makes a browser do something: run a script, fire an event handler, load a resource, navigate somewhere. It defends by removing the *structural* carriers of that behaviour — script elements, style elements, event-handler attributes, dangerous URL schemes — and by keeping everything else, above all the text, because the text is the page's actual content and removing it would defeat the purpose of fetching the page at all. ## The payload was never markup A prompt injection planted for an agent to read is a sentence. It sits in a paragraph, in a list item, in a caption — a normal text node in a normal element. Nothing about it is structural. So every transform in the chain that is defined as *discard markup, preserve prose* passes it through by construction: | stage | what it removes | what it keeps | | --- | --- | --- | | DOM parse | nothing; it re-represents | all text nodes | | boilerplate removal | whole low-content regions | the main content region | | sanitiser | executable and unsafe markup | text node contents | | HTML-to-text | remaining tags | the concatenated prose | The stage that would have noticed *structure* is exactly the stage that hands the *text* on intact. This is not a bug in the sanitiser. The sanitiser met its specification; its specification was never about meaning. ## Getting the direction of the claim right This is where interviewers separate people. A page passing a sanitiser establishes one fact: markup a browser could have executed was removed. It does not establish that the remaining text is inert, that a model will treat it as data, or that it scored below anything — a sanitiser emits no score and makes no judgment about content. Nor does a model later obeying a fetched sentence prove anything about the sanitiser; it proves the sentence reached the context window with no marker separating it from the analyst's own request. The converse trap is just as common: because stripping is a no-op here, people conclude it is worthless. It is not. It still does its own job against its own adversary, and it still matters for what the agent's *output* is rendered into later. Two different problems, one control, and only one of them is solved. ## What this costs the attacker, and where it stops The striking property of this construction is how cheap it is. It needs no concealment, no encoding trick, no bet on a particular parser's quirk. Plain prose is invariant under whitespace collapsing, entity decoding, quote and dash rewriting, markdown conversion and every other prose-preserving transform in the chain — which is to say, invariant under the entire chain. The attacker pays almost nothing and does not even need anyone to link to the page; they only need it to be reachable by whatever selection step the agent runs. The price is visibility. Prose that survives everything is prose a human reading the page can also read; there is no hiding in this construction, only reliability. And it does have limits that have nothing to do with sanitising: if the fetcher takes only the first N characters, a span below the cut never arrives; if boilerplate removal drops the region the span sits in, it dies with the region; if the agent never ingests page text at all and works from a separately produced summary, the chain is different. Those are the honest boundaries — none of them is the sanitiser. ## How to say it in a loop Name the consumer. `Sanitising is markup removal for a browser; the injection is prose for a model; markup removal preserves prose by definition.` Then say what a clean sanitise result actually licenses you to claim, and what it does not. Candidates who answer at the level of `HTML was stripped so we are fine` have not yet noticed that the two threats travel in different parts of the same document.
- Does the same reasoning apply to an HTML-to-markdown converter rather than a sanitiser?Yes, and more strongly. A converter is defined as a lossy transform in exactly the direction that preserves prose: it exists to throw away presentation and keep the words. Any stage whose contract is `give me the readable text` is a stage that guarantees the planted sentence survives. The chain is not one such stage but several in series, and each one narrows what could have been noticed.
- If sanitising is a no-op here, why do teams keep listing it as an injection control?Because it genuinely is a control — for its own adversary, and for what happens to the agent's own output when a client renders it. The failure is a category slip: one word, `sanitise`, is doing duty for two threat models. In a review, ask which consumer the control protects and which one the finding concerns; if the answers differ, the control is not evidence about the finding.
- A model obeyed a sentence from a fetched page. What has that proved?That the sentence reached the context window carrying no marker that distinguished it from the operator's own instruction, and that the model preferred it in that turn. It has not proved that any stage malfunctioned, that the store or server was compromised, or that the same page will work on the next run. Probabilistic preference means one obeyed instruction is one observation, not a reliability figure.
A metal detector at the door is very good at finding knives and has no opinion whatsoever about what a visitor says once they are inside.
saying these in an interview costs you the question
- Says stripping tags removes the payload
- Treats sanitiser output as trusted text
- Confuses markup removal with meaning removal
- Assumes a payload must be hidden to work
- Claims a clean sanitise proves the text is harmless
- Conflates an HTML sanitiser with a content screening layer