skip to content

Output Encoding & Injection Theory

Every injection is the same bug: data crossing into an interpreter and being read as instructions. Encoding for the specific sink, or better, keeping data out of the instruction stream entirely, is the durable fix — and interviewers want that unified explanation rather than a list of tricks per technology.

part ofApplication security & secure codingoverview, primer and where to startread it →
on this pageshow

questions

4

A single user-supplied value can land in several different positions inside one rendered page. Enumerate those output positions, explain why one generic HTML-escaping helper is insufficient for all of them, and describe what changes when the positions nest inside each other.

level: middleimportance: must knowfreq 62%

answer

  1. text / quoted attr / unquoted attr / handler / URL / script / CSS
  2. unquoted attr: space and slash end it
  3. handler = HTML decode then JS parse → JS-encode first
  4. `</script>` beats JavaScript quoting
  5. javascript: and data: need scheme allow-listing, not escaping

basics

~20 s

A value can land in element text, a quoted attribute, an unquoted attribute, an event-handler attribute, a URL attribute, inside a script block, or in CSS. Each is a different grammar with a different escape set. Where they nest, the value passes through two decoders, so it must be encoded for both, innermost grammar first.

solid answer

~50 s

The distinct positions and what breaks out of each: - **Element text** — `<` and `&` matter; the standard HTML-escape helper is correct here and only here. - **Quoted attribute** — the matching quote and `&`; HTML escaping suffices. - **Unquoted attribute** — whitespace, `/`, `>` and backtick all end the value, so the safe set depends on how the *template author* quoted it. - **Event-handler attribute** (`onclick=`) — the HTML parser decodes entities and hands the result to the JavaScript parser: two grammars, so JavaScript-encode first, then HTML-encode. - **URL attribute** (`href`, `src`) — encoding the characters does not help if the attacker controls the scheme; `javascript:` and `data:` require scheme allow-listing, an enumeration control. - **Inside a script block** — `</script>` terminates the block regardless of JavaScript quoting, so JavaScript string escaping alone is not enough. - **CSS value / `url()`** — its own grammar again. Hence context-aware auto-escaping engines parse the *template* to learn each hole's context, rather than escaping the value in isolation.

code

text · 8 lines
text
<button onclick="greet('{{NAME}}')">

  browser: HTML tokenizer reads attribute value
           -> decodes character references (&#39; becomes ')
           -> hands the decoded string to the JS parser

  therefore encode inner-first:
           NAME -> JS string escape -> HTML attribute escape

go deeper

for a junior

List the main positions — text, attribute, script, URL — and say that each needs its own escaping and that the generic HTML escaper is only right for text.

for a middle

Add the unquoted-attribute break-out, the javascript: scheme problem, and the fact that </script> ends the block regardless of JavaScript quoting.

for a senior

Derive the nesting order from the decoder chain (HTML decodes, then JavaScript parses) and argue for context-aware template compilers or DOM APIs over per-call-site escaping.

for a principal

Set the codebase rule: engines that select the escaper by parsing the template, banned holes in tag and attribute names, data handed to the client as JSON rather than interpolated, and a short reviewed list of raw-markup escape hatches.

## One value, many grammars A rendered page is not one language. The HTML parser produces a tree; some attributes are further parsed as JavaScript; some as URLs; `<style>` and `style=` are parsed as CSS; `<script>` content is JavaScript. A templating hole is a position in *one* of those grammars, and the transformation that keeps the value a single terminal differs for each. ## The positions **Element text.** `<div>HERE</div>`. Only `<` (and `&`, for correctness) can introduce structure. The classic escape of `< > & " '` is correct and slightly over-broad. This is the context people picture when they say "escape HTML", which is why the single-helper habit forms. **Quoted attribute value.** `<a title="HERE">`. The value ends at the matching quote, so that quote plus `&` must be encoded. Encoding `<` is unnecessary but harmless. **Unquoted attribute value.** `<a title=HERE>`. Now whitespace, `>`, `/`, `=` and a backtick all end the value, after which the attacker is writing *attributes* — and adding an event handler is enough; no `<` is required. The critical property is that the required escape set is decided by the template author's quoting, not by the value. This is the strongest argument for template engines that parse the template: an escaper handed only the string cannot know whether it is inside quotes. **Event-handler attribute.** `<button onclick="doIt('HERE')">`. Two grammars in sequence: the HTML parser decodes character references in the attribute value, and the *result* is parsed as JavaScript. So `&#39;` in the source becomes `'` before JavaScript ever sees it. The value must therefore be JavaScript-escaped **first**, then HTML-escaped, and applying only one — in either order — leaves a break-out. The ordering claim is impossible to derive from a single-context view of encoding, and it is the reason nesting is treated as its own topic. **URL-bearing attribute.** `<a href="HERE">`, `<img src="HERE">`. If the attacker controls the *start* of the URL, they control the scheme: `javascript:` executes, `data:` can carry a document with its own script. No amount of character encoding fixes this, because the characters are all legal — the problem is the node's meaning. The control is enumeration: parse the URL, require the scheme to be in a closed allow-list (`https`, `mailto`, or a relative reference), then URL-encode the components. Where a value fills only a query parameter, percent-encode that component; the encoding for a URL component is not HTML escaping. **Inside a script block.** `<script>var x = "HERE";</script>`. Two independent hazards. JavaScript string escaping handles quotes and backslashes and prevents leaving the literal. But the HTML tokenizer is still looking for the end of the script element, so the literal sequence `</script>` inside the string terminates the block no matter how the JavaScript is quoted — as can `<!--` in legacy parsing modes. Safely embedding a value here means escaping for JavaScript *and* neutralising `</` for the HTML tokenizer. The same applies to a JSON blob inlined in a script tag: valid JSON is not automatically safe HTML, and historically the raw line separators U+2028/U+2029 were also a hazard in JavaScript string literals. **CSS.** `<div style="width: HERE">` or a `url()` value. Another grammar with its own escaping and its own URL sub-context. ## Why the generic helper is insufficient — stated as a principle An escaping function encodes a model of one grammar. A page contains at least four, some nested. Applying a helper built for element text to an unquoted attribute, an event handler, a URL or a script literal is not "partially safe" — it is a different lexer, so the result is either a break-out or mangled output. The property that matters is: *did the transformation guarantee this value remains one terminal in the grammar actually running at this position?* ## What good implementations do Context-aware auto-escaping template compilers parse the template itself, track the parser state at every interpolation point, and select the escaper — including composing escapers for nested contexts and refusing constructs they cannot analyse, such as a hole that spans a tag name or an attribute name. That is the **structural separation** rung expressed for markup: the engine builds the structure, and values are attached as data. The next best is to avoid string-built markup entirely and set text and attributes through DOM APIs, where the browser's own parser never sees your value as markup. Escaping by hand at the call site is the rung below, and it is where the ordering and nesting mistakes live. ## Practical rules Quote every attribute so the escape set is predictable. Never place untrusted data into a tag name, an attribute name, or the scheme portion of a URL — those are structure, so they need enumeration, not encoding. Pass data to client-side code through a data attribute or a JSON payload read at runtime rather than by interpolating it into a script literal, which collapses two grammars into one and removes the nesting problem entirely.

  • Why is HTML-escaping a value that goes into an `href` attribute not sufficient?
    HTML escaping keeps the value inside the attribute, but the attribute's contents are then interpreted as a URL. Every character of `javascript:alert(1)` is legal in an attribute, so the escaped value is preserved intact and the link executes when clicked. The control is to parse the URL and require its scheme to be in a closed allow-list, then percent-encode the components — enumeration for the structure, encoding for the data.
  • You need to pass server data to client-side JavaScript. What is the safest shape?
    Do not interpolate it into a script literal, because that nests JavaScript inside HTML and requires two escapers plus neutralising `</`. Emit it as a JSON payload in a data attribute or a `<script type="application/json">` block, HTML-escaped for the attribute or element-text context, and have the client parse it at runtime. That leaves exactly one grammar at the insertion point.

saying these in an interview costs you the question

  • "We escape `<` and `>` everywhere, so we're covered" — useless in unquoted attributes, handlers, URLs and script literals.
  • Assuming JavaScript string escaping protects a value inside `<script>`, when `</script>` closes the block regardless.
  • Believing an encoded value in `href` cannot execute, ignoring the `javascript:` and `data:` schemes.
  • Applying HTML escaping and JavaScript escaping in the wrong order for an event-handler attribute.
  • Leaving attributes unquoted and assuming the value-level escaper can compensate.

context

open as a page

Escaping untrusted data is usually taught as "strip or replace the dangerous characters". Give the definition of output encoding that explains why the same string can be safe in one output position and dangerous in another, and why transforming data when it arrives is the wrong frame.

level: middleimportance: must knowfreq 60%

basics

~20 s

Encoding is a function of the destination grammar, not of the data. At the moment a value is placed into an interpreter's input, it must be transformed into a lexeme that parser cannot exit. Nothing about the string alone is dangerous, so a transformation chosen at arrival — before the destination is known — cannot be correct.

open as a page

Cross-site scripting is commonly split into reflected, stored and DOM-based forms. Explain what actually distinguishes them at the level of where untrusted data becomes code, and why that distinction changes which defences and which detection can possibly work.

level: seniorimportance: must knowfreq 58%

basics

~20 s

The three differ by where the data-to-code transition happens and how the data got there. Reflected and stored transitions occur while the server builds the response; the DOM-based transition occurs in the browser, in client code writing to a script-executing sink. Server-side encoding cannot fix the third, and server-side logs and filters never see it.

open as a page

Some injection sinks are parsers you never invoke: a spreadsheet application opening an exported file, an operator's terminal displaying a log line, a mail server reading headers. What does correct output encoding mean when the interpreter runs outside your process, and how do you reason about it?

level: principalimportance: should knowfreq 26%

basics

~20 s

The rule is unchanged — encode for the grammar that will parse the value — but that grammar belongs to a consumer you do not control and may not know. So you use a real serializer for the format, constrain values to a closed set where the consumer's rules are unknowable, and treat every export or log line as a sink with an owner.

open as a page