skip to content

Input validation is often taught as "strip the dangerous characters before you use the value". Give the definition of validation that explains what it can and cannot guarantee, and why treating it as the fix for injection is the wrong frame.

level: middleimportance: must knowfreq 72%

answer

  1. closed-world domain rule, written without the sink's grammar in view
  2. one value, many sinks, contradictory metacharacter sets
  3. ladder: structural separation > escaping > validation > detection
  4. validation exclusively owns domain model + resource bounds
  5. every bypass is a parser differential; stricter rules don't fix it

basics

~20 s

Validation is a closed-world constraint applied at a trust boundary in terms of your own domain, without knowing the grammar of whatever will eventually consume the value. It shrinks the input space and owns resource limits; it cannot guarantee safety at a parser it has never seen.

solid answer

~60 s

Injection exists at a boundary where data is fed into a parser that can read part of it as syntax. Only the consumer knows that grammar. Validation runs far from the sink, usually once, on a value that may later reach a query language, a shell, a filesystem path, HTML, a log pipeline, a directory filter and a template — grammars whose dangerous characters and legal domains contradict each other. A single validator would have to be the intersection of every grammar, which either breaks the product (names with apostrophes, text containing angle brackets) or misses one sink. Rank the controls: **structural separation > escaping/transformation > validation > detection**, weakening from guarantee to heuristic. What validation genuinely owns is the domain model — type, length, range, format, counts, nesting depth — which no sink control addresses and which kills whole classes of resource abuse. And the invariant behind every bypass: **the validator and the consumer must agree on what the input is**; wherever the two parses can diverge, the strictness of the rule is irrelevant.

go deeper

for a junior

Be able to state that validation happens at the boundary against your own domain rules, and that it is not what makes a query, a page or a command safe — the sink has its own control. Give one concrete pair: length and range are validation's job; safe handling at the consumer is the consumer's job.

for a middle

Explain why the grammar belongs to the consumer, so a boundary rule cannot be correct for every sink at once. Name the ladder in order — structural separation, escaping, validation, detection — and say what validation exclusively owns: the domain model and resource bounds.

for a senior

Lead with the definition and the parser-differential invariant, and use it to redirect a review: when a bypass is found, ask where the two parses diverged rather than how to tighten the rule. Be able to point at which controls in a real service are guarantees and which are heuristics.

for a principal

Frame it as ownership across a codebase: the boundary owns the domain and the resource envelope, every sink owns its own safety, and no team may discharge a sink obligation by pointing upstream. Set the rule that decoding happens once at a known point, so validator and consumer see the same bytes by construction.

## A definition that survives contact with reality Validation is: **a constraint applied at a trust boundary that accepts only values belonging to the application's own domain, expressed without knowledge of the interpreter that will eventually consume them.** Every clause is load-bearing. *Trust boundary* says where it happens. *The application's own domain* says the rule is closed-world — a finite description the code owns — rather than a list of the attacker's tricks. *Without knowledge of the interpreter* is the limitation, and it is the reason the whole framing of "validate to prevent injection" is wrong. ## Why validation cannot be the fix for injection Injection bugs exist at a parser boundary where a value is folded into text that a parser will read, so parts of the data become syntax. The set of characters that are syntax, and the escaping rules that neutralise them, are properties **of that parser** — and they differ everywhere. A shell, a query language, an HTML attribute, a JavaScript string literal, a directory filter, a path expression over a document tree and a template engine each have their own metacharacters, their own quoting states and their own encoding assumptions. Some accept structured input where there is no lexing of strings at all, so "dangerous characters" is not even a meaningful category there. A validator sitting at the entry boundary knows none of this. It knows only that a field is called `displayName`. Three consequences follow: 1. **One value, many sinks.** The same string may be stored, then rendered into a page, then written into a log, then used to name a file, then sent to a partner API. A rule that is safe for one is not safe for the others, and each new sink added later is a fresh gamble against a rule nobody re-examined. 2. **The intersection is empty or useless.** To be safe everywhere, the rule must forbid every metacharacter of every grammar: apostrophes, angle brackets, backslashes, semicolons, dollars, backticks, parentheses, asterisks, dots, slashes. That set intersects the legitimate domain of ordinary human data — names, addresses, comments, product descriptions — so you either break the product or you carve out exceptions, and the exceptions are the vulnerability. 3. **Being right about the domain does not make you right about the grammar.** A value can be a perfectly valid twenty-character product code and still be catastrophic in a sink you did not anticipate. ## The ladder Rank controls by the strength of the promise they make: **Structural separation** — the value never enters the parser's syntax channel at all. The consumer receives the structure and the data through different channels, so the data cannot be re-read as syntax. This is a *guarantee*: it holds regardless of the value's content, and it does not require anyone to have enumerated the dangerous characters correctly. **Escaping / transformation** — the value is rewritten so the parser reads it as inert data. Weaker, because correctness depends on modelling the consumer's lexer exactly, and on that lexer's mode and encoding assumptions continuing to hold. Apply it for the wrong context, or after the consumer changes a setting, and it silently stops working while still looking like a control. **Validation** — the value is checked and rejected before use. Weaker again, because it is applied without the sink in view: it can be entirely correct about the domain and still wrong about the grammar. **Detection** — you observe and alert. Weakest: it prevents nothing, it tells you afterwards. One special case matters and must be stated in the same vocabulary: where the untrusted value must select *structure* rather than supply a *value*, structural separation is unavailable and **closed-world enumeration takes the top rung in its place** — see `found-appsec-input-validation-allowlist-versus-denylist` for why. ## What validation genuinely owns Having demoted it, be precise about where it is the *only* control available: - **The domain model.** Types, ranges, lengths, formats, required combinations, referential shape. No sink control tells you that a quantity must be between 1 and 100, and pushing a billion through a perfectly separated query is still a bug. - **Resource bounds.** Body size, string length, array and member counts, nesting depth, decompression ratio, upload size. There is no downstream escaping that makes a four-gigabyte document or a billion-fold entity expansion safe. This is validation's exclusive territory and the strongest single argument for it. - **Sinks with no structural control.** A value used as an array index, a numeric identifier handed to native code, a size used in an allocation. - **Signal.** Rejections are a clean anomaly stream: a rule that suddenly fires a thousand times an hour is either wrong or under attack. So the correct framing is defence in depth with clear ownership: validation makes the input space small and knowable, and the sink still owns its own safety. ## The invariant behind every bypass > A validation bypass is always a **parser differential**: the validator and the consumer disagreed about what the input was. Double-decoding, normalisation applied after the check, duplicate keys resolved differently by two parsers, truncation applied after validation, charset re-interpretation, a value re-decoded downstream — in every case the rule was fine and the two parses were not. The practical corollary is worth saying out loud in an interview: **once a differential exists, making the rule stricter has no effect.** You do not out-pattern two decoders; you remove the second parse, or you validate on exactly the bytes the consumer will see.

  • If validation cannot prevent injection, why not drop it and rely purely on the sink control?
    Because validation owns things no sink control touches. Resource bounds — body size, array counts, nesting depth, decompression ratio — are enforceable only at the boundary, and no amount of correct separation at the query or template layer makes an oversized or hostilely nested document safe. It also enforces the domain model (ranges, formats, required combinations) whose violations are ordinary business bugs, protects sinks that have no structural control at all, and produces a clean rejection signal. The claim is that validation is the wrong *primary* control for injection, not that it is optional.
  • A reviewer says the field is validated to letters and digits only, so the query built from it is safe. What is wrong with that argument?
    It confuses a property of the value with a property of the boundary. The rule was written without the consumer's grammar in view, so it is an accident if the two happen to agree, and the agreement is not maintained: the field gains a new legitimate character, or the value reaches a second sink with different metacharacters, and the safety silently evaporates with no code change at the sink. It also assumes the consumer sees exactly the bytes the validator saw, which fails the moment anything decodes or normalises in between. The safety argument must be made at the sink, where the grammar is known.
  • Where in the request path should validation run, given that a value can be decoded more than once?
    On the representation the consumer will actually use, after all decoding and normalisation the platform will perform, and with any further decode removed rather than tolerated. If a layer downstream decodes again, the validator's view and the consumer's view are different strings and the rule is decorative. In practice that means canonicalising once, at a known point, then validating — and treating a second decode anywhere below it as the defect to fix.

Validation is the customs check at the border: it can confirm the crate contains what the manifest claims, in the quantity and size allowed. It cannot know what the machine at the far end of the supply chain will do when it feeds the contents in — that safety belongs to the machine.

saying these in an interview costs you the question

  • Saying validation prevents injection, or that a strict enough pattern makes a concatenated query safe.
  • Proposing a denylist of dangerous characters as the boundary control, then adding exceptions when real data breaks it.
  • Treating a bypass as evidence the rule was too loose, and responding by tightening the pattern instead of removing the parser differential.
  • Claiming one validator can cover all sinks, without noticing that the metacharacter sets of different grammars contradict each other and intersect legitimate human data.
  • Dropping validation entirely once the sink is separated, leaving no owner for size, count and nesting limits.

context