Escaping untrusted data is usually taught as "strip or replace the dangerous characters". Give the definition of output encoding that explains why the same string can be safe in one output position and dangerous in another, and why transforming data when it arrives is the wrong frame.
answer
- danger is a property of the sink, not the string
- encode at the sink, once, on the way out
- store raw, render encoded
- escaping = re-implementing someone else's lexer
- separation > escaping > validation > detection
basics
~20 sEncoding is a function of the destination grammar, not of the data. At the moment a value is placed into an interpreter's input, it must be transformed into a lexeme that parser cannot exit. Nothing about the string alone is dangerous, so a transformation chosen at arrival — before the destination is known — cannot be correct.
solid answer
~60 sInjection happens where data crosses into an interpreter and is read as syntax. Output encoding is the transformation that guarantees a value stays a single **terminal** — one string literal, one text node, one attribute value — in that interpreter's grammar. The correct transformation is therefore determined by the *sink*: the same apostrophe is inert in HTML text, ends a quoted attribute, and terminates a JavaScript string literal. Encoding on input fails for three structural reasons: the destination is unknown at arrival, one stored value typically feeds several different sinks over its life, and the transformation is lossy and non-idempotent, so you get mangled names and double-encoding while still being wrong somewhere. The defence ladder is: **structural separation > escaping/transformation > validation > detection**. Escaping is a re-implementation of someone else's lexer and is only as correct as your model of it; structural separation — binding a parameter, creating a text node, a context-aware template compiler — never lexes the value at all, which is why it is a guarantee rather than a good effort.
go deeper
Say encoding depends on where the value is going, that you store the raw value and encode when rendering, and give one example of a character that is safe in one context and not another.
State the invariant — the value must remain a single terminal in the destination grammar — and give the three structural reasons encoding at input fails.
Lead with the ladder and articulate why escaping is weaker in kind: it re-implements a lexer you do not own, whereas separation never lexes the value.
Talk about making the safe path the default across the codebase — template engines and APIs that encode by construction, escape hatches that are few, named and reviewable — rather than relying on per-call-site discipline.
## The definition An interpreter is anything that turns a string into structure: a SQL parser, a shell, an HTML parser, a JavaScript engine, an LDAP filter parser, an XPath evaluator, a template engine, a CSV reader. A vulnerability of this class exists whenever untrusted data is concatenated into text that such an interpreter will later parse, because the parser has no way to know which characters came from the developer and which from the attacker. It sees one string and applies one grammar. **Output encoding** is the transformation applied at the moment of insertion that guarantees the value will be lexed as a single terminal in the destination grammar — one string literal, one text node, one attribute value, one path segment — and cannot terminate that terminal or introduce structure. Two consequences follow immediately. First, **danger is not a property of the string**. `'` is ordinary text in an HTML text node, ends a single-quoted attribute value, and closes a JavaScript string literal. `<` is meaningful only where an HTML parser is running. A character is dangerous exactly when the destination grammar gives it syntactic meaning at that position. So a question of the form "is this input safe?" is malformed; only "is this value correctly encoded for *that* sink?" is answerable. Second, **the transformation must happen at the sink, at the moment of use**. Encoding when data arrives fails structurally: you do not yet know which sinks it will reach; a stored value typically reaches several (an HTML page, a JSON response, a PDF, a log, an email, a CSV export), each with a different grammar; and the encoding is lossy and non-idempotent, so `O'Brien` becomes `O'Brien` in the database, appears literally in the CSV export, gets double-encoded on the next render, and breaks any search or comparison on the column. Store the raw value; encode on the way out; encode exactly once. ## The ladder, and why the top two rungs differ in kind Use one ordering and one vocabulary: 1. **Structural separation** — the data never becomes part of the text the interpreter parses. A bound query parameter, `textContent` or a created DOM text node instead of markup assembly, a builder API that constructs the structure natively. The parser receives the template and the value through *different channels*, so no lexing of the value can change the structure. This is a guarantee. 2. **Escaping / transformation** — the value is embedded in the text but transformed so the lexer cannot exit the terminal. This is strictly weaker for one reason worth stating explicitly: escaping is a **re-implementation of the destination's lexer**, and it is correct only to the extent your model of that lexer matches reality — including its character-set handling, its nested grammars and its edge cases. That is a modelling task you can fail silently, whereas separation has nothing to model. 3. **Validation** — constraining the value's shape (a number, a member of an enumerated set, a known key). Real defence in depth, and the *only* option where structural separation does not exist because the untrusted value must supply structure rather than data: a sort column, a target scheme, a template name. There the top rung is occupied instead by closed-world **enumeration** — mapping the client's token to a code-owned value — and the reason is open-world versus closed-world: a pattern that constrains characters is open-world and still admits every legal identifier, including the sensitive ones, while an enumerated map is finite and owned by the code. 4. **Detection** — response policies, scanners, monitoring, review. Useful containment and evidence; guarantees nothing about a given call site. ## Why blocklisting sits below all of this Blocking "bad characters" or known payload shapes inverts the burden: it requires you to enumerate everything an attacker might send, which is unbounded, over a grammar you do not control, through decoders you may not have modelled. Allow-listing and encoding both invert it back — they enumerate what *you* will emit or accept, which is finite. Any control that fails open on an unforeseen input is a heuristic, and heuristics belong below guarantees on the ladder, never in place of them. ## The unifying claim SQL injection, command injection, XSS, LDAP and XPath injection, template injection and even spreadsheet formula injection are one defect with different parsers behind it: a boundary where data becomes code. The reference case is a bound SQL parameter, where the template is compiled before values are attached. Every other sink is asking the same question — is there a channel that keeps this value out of the parsed text, and if not, do I know this lexer well enough to escape for it? That question, not a list of characters, is what transfers between stacks.
- A colleague proposes HTML-encoding everything on the way into the database so nothing downstream has to think about it. What do you tell them?It picks one sink's grammar for data that will reach many. The stored value is now wrong for JSON, CSV, email, PDF and log sinks, comparisons and search on the column break, and re-encoding on render produces double-encoded output. It also gives false confidence, because a value that reaches a JavaScript or URL context is still unsafe. Store raw, encode at each sink.
- Where does validation belong if encoding is the real defence?Validation constrains the value's shape at the trust boundary and is genuine defence in depth, but it cannot make an unencoded value safe in an arbitrary sink. Its indispensable role is where the untrusted value must supply structure rather than data — a sort column, a URL scheme, a template name — because there is no channel to separate structure from data, so a closed-world enumerated map replaces encoding at the top of the ladder.
A shipping label is not intrinsically dangerous or safe; it depends on which sorting machine reads it. Printing a barcode at the warehouse door, before you know the destination country, means it will be misread somewhere.
saying these in an interview costs you the question
- "Sanitize on input and you're done" — the destination grammar is unknown at input time and one value feeds many sinks.
- Treating specific characters as inherently dangerous rather than dangerous relative to a grammar.
- Claiming escaping and parameter binding are equivalent; one models a lexer, the other never lexes the value.
- Reaching for a blocklist of payload patterns, which requires enumerating the attacker's options instead of your own.
- Encoding more than once "to be safe", producing double-encoded output that hides the real bug.