Some injection sinks are parsers you never invoke: a spreadsheet application opening an exported file, an operator's terminal displaying a log line, a mail server reading headers. What does correct output encoding mean when the interpreter runs outside your process, and how do you reason about it?
answer
- the parser can be outside your process
- CSV quoting ≠ formula prevention
- newline forges log records; escape sequences target the operator's terminal
- log viewers make logs an HTML sink aimed at admins
- top rung shifts to a real format writer, then enumeration
basics
~20 sThe rule is unchanged — encode for the grammar that will parse the value — but that grammar belongs to a consumer you do not control and may not know. So you use a real serializer for the format, constrain values to a closed set where the consumer's rules are unknowable, and treat every export or log line as a sink with an owner.
solid answer
~60 sInjection does not require that *you* run the interpreter; it requires that some interpreter eventually parses text you assembled. - **Spreadsheet formula injection**: a CSV field beginning `=`, `+`, `-` or `@` is evaluated as a formula by the opening application. Correct CSV quoting per the format's rules does **not** prevent evaluation, because quoting is the CSV grammar and formula detection happens after the field is parsed. The control is value-level: prefix or reject leading formula characters, or export a format that has no formula semantics. - **Log injection**: an unescaped newline lets an attacker forge whole log records, defeating later analysis; an ANSI escape sequence can manipulate the terminal of the operator reading it; and if a log viewer renders records as HTML, the log becomes a stored-XSS sink aimed at administrators. - **CRLF injection** into mail or protocol headers splits one header into several. Because you cannot model an unknown consumer's lexer, the top rung shifts from an encoder you write to a **serializer that owns the format** — a real CSV writer, a structured logger emitting one escaped record per event — plus closed-world constraints where the format cannot be trusted at all.
go deeper
Say the same rule applies — encode for whatever will parse the text — and give one example, such as a CSV cell starting with = being run as a formula by a spreadsheet.
Cover formula injection, newline-based log forging and CRLF header splitting, and note that correct CSV quoting does not stop formula evaluation.
Argue the ladder shift: a real format writer replaces the missing separation channel, with value-level neutralisation and closed-world constraints beneath it, and identify logs' multiple consumers.
Make it an inventory problem — every emitted artifact is a sink with a named consumer and an owning serializer — and treat fidelity trade-offs such as prefixing formula characters as product decisions to be recorded.
## The generalisation The injection invariant says: wherever text you assembled will later be parsed, the values inside it must remain single terminals of that grammar. Nothing in that statement requires the parser to live in your process, to run during the request, or to be one you chose. It routinely runs later, elsewhere, and belongs to someone else. That is what makes this family easy to miss — nobody thinks of an export or a log line as calling an interpreter. ## The canonical cases **Spreadsheet formula injection.** An exported tabular file is opened by a spreadsheet application, which treats a cell whose content begins with `=`, `+`, `-` or `@` (and, in some locales and versions, a leading tab or carriage return before those) as a *formula* rather than text. The formula language can reference other cells, and depending on the application and its warning settings, can reach external data or launch other commands, generally behind a confirmation prompt. The instructive part is the defence analysis. Quoting the field correctly per the delimited-text format is what a competent writer already does, and it does not help: quoting governs how the field is delimited, and formula detection is applied to the field's *content* after it has been extracted. So this is a case where the correct encoder for the transport grammar leaves the consumer's second grammar untouched. Controls that do work are value-level: reject or prefix a leading formula character (a leading apostrophe or a space, each with its own fidelity cost), or export a format that has no formula semantics at all and let the consumer import it deliberately. **Log injection.** A log line is a record in a line-oriented format. An untrusted value containing a newline lets an attacker append fabricated records — a forged authentication success, a plausible error attributed to another principal — which corrupts precisely the evidence you would use during an investigation. Carriage returns and terminal escape sequences target a different consumer again: the operator's terminal, where escape sequences can overwrite earlier output, hide lines, or alter the display. And when logs are surfaced in a web console, the log record becomes an HTML sink aimed at administrators, which is a stored-XSS path with unusually high privilege at the far end. The structural answer is a **structured logger**: emit one object per event with fields escaped by the serializer, so newlines inside a value are data within a record rather than a record boundary. That is separation rather than escaping — the format's own writer owns the framing, and no concatenation happens. **Header injection.** Mail headers, and any line-oriented protocol header, are separated by CRLF. An untrusted value placed into a header field that contains CRLF creates additional headers or ends the header block. The control is again a builder that owns the framing plus rejection of control characters in field values — never a hand-assembled header string. **Other members.** Filenames placed into content-disposition metadata; values re-serialised into XML, YAML or JSON by string concatenation; archive entry names that a downstream extractor interprets as paths; terminal output of any kind; and identifiers that a downstream system parses with rules you have never read. ## How the ladder shifts For an interpreter you invoke, the top rung is a channel that keeps the value out of the parsed text — a bound parameter, an argument vector, a created DOM node. For a consumer outside your process there is no such channel, because you hand over a document, not a call. The top rung becomes the closest available equivalent: 1. **A serializer that owns the format.** A real writer for the format — delimited-text writer, structured-log encoder, header builder — so escaping is implemented once by code that models the grammar, and your code never concatenates. 2. **Transformation** — value-level neutralisation for the consumer's *second* grammar, such as prefixing a leading formula character or stripping control characters from a header value. Note this is transformation on the value, and like all such rungs it is only as good as your model of a consumer you do not control. 3. **Validation / closed-world enumeration** — where the consumer's rules are genuinely unknowable, constrain what may be emitted rather than trying to encode everything. This is open-world versus closed-world again: guessing every construct some future spreadsheet or terminal might interpret is unbounded, while an enumerated character or value set is finite and owned by you. 4. **Detection** — sampling exports, alerting on anomalous log records, reviewing new export paths. ## The judgment question The practical principal-level move is to treat **every export, log, notification and file the system emits as a sink with a named consumer**, and to record which format writer owns each one. If nobody can name the consumer's parser, that is the finding: an unowned sink is a place where the invariant cannot be checked, and the appropriate response is to narrow what may be emitted rather than to hope the encoder is right. Fidelity trade-offs are real — prefixing formula characters changes the data the recipient sees — so the decision belongs to whoever owns the product behaviour, and it should be recorded rather than made silently in a utility function.
- Why doesn't correctly quoting the field in a delimited-text export prevent formula injection?Quoting belongs to the delimited-text grammar and decides only where the field starts and ends. Once the reader has extracted the field, the spreadsheet application applies a second grammar to its content and treats a leading `=`, `+`, `-` or `@` as a formula. The transport encoder is correct and irrelevant — the value must be neutralised for the consumer's second grammar, or a format without formula semantics must be used.
- What is the most durable fix for log injection?Stop concatenating log lines. Emit structured records where a serializer owns the framing and escapes field contents, so an embedded newline is data inside a record rather than a record boundary. Then handle the second consumer explicitly: strip or encode terminal escape sequences for operators reading raw output, and encode records for HTML when a viewer renders them, since that consumer is a browser.
saying these in an interview costs you the question
- Assuming injection requires your code to execute the interpreter.
- Believing a compliant CSV writer makes an export safe against formula evaluation.
- Treating logs as inert text rather than as a format with at least three consumers — parsers, terminals and log viewers.
- Building headers or protocol lines by concatenation and relying on validating the value's shape.
- Trying to enumerate everything a downstream application might interpret instead of constraining what you emit.