Why must untrusted input be canonicalised before it is validated, and what class of bug appears when validation runs on the raw form? Give concrete examples of processors that disagree about what the same bytes mean.
answer
- check the bytes, use the decoded value: two different strings
- %252e passes the check, becomes a dot on the second decode
- NFKC folds fullwidth to ASCII; case folding is locale-sensitive and lossy
- JSON duplicate keys: last-wins vs first-wins vs error
- single-pass removal of `AB` turns `AABB` back into `AB`
basics
~20 sBecause a rule applied to one representation says nothing about a different representation of the same value. Every bypass is a parser differential: the validator and the consumer decoded differently. Canonicalise once, validate the canonical form, never decode again downstream.
solid answer
~50 sValidation constrains a *string*; the consumer acts on whatever it decodes. If any decode, normalisation or re-parse happens in between, the check ran on a string that never reaches the sink. Real divergences: percent-decoding depth and ownership, where one component decodes once and another decodes again, so `%252e` becomes the literal text `%2e` at the check and `.` at the sink; Unicode compatibility normalisation, where NFKC folds fullwidth forms (U+FF0F to `/`, U+FF1C to `<`) into ASCII metacharacters, and case folding is locale-sensitive (Turkish dotless i) and non-injective (sharp s folds to `ss`); UTF-8 decoders that differ on ill-formed input between rejecting, replacing with U+FFFD, and historically accepting overlong forms; JSON parsers that disagree on duplicate keys (last-wins, first-wins, error). Rules: establish one canonical form at the boundary, prefer rejecting non-canonical input over transforming it, validate after the last decode, never decode again afterwards. Corollary: once a differential exists, tightening the rule changes nothing.
code
text · 7 linesclient sends: name=%252e%252e
component A (validator, decodes once)
sees: "%2e%2e" -> contains no dot -> PASS
component B (consumer, decodes again during binding)
uses: ".." -> the rule never applied to this string
fix: one decode, owned by one component; validate that result; no decode after itgo deeper
Know the core: the checker and the user of the value must be looking at the same string. Decode once, check the decoded value, and do not decode again afterwards.
Be able to name at least two real divergences with mechanics - double percent-decoding of %252e and NFKC folding a fullwidth character into an ASCII metacharacter - and explain why the check passed while the sink saw something else.
Frame it as a parser differential and state the consequence: strictness does not close it, agreement does. Cover decode depth, Unicode normalisation and case folding, UTF-8 ill-formed handling, and duplicate keys in structured formats, then give the rules: one canonical form, reject rather than transform, validate after the last decode, never decode after validating.
Argue at the system level: two independent parsers over the same bytes is a standing defect, so either the gateway and the service share one parser and one schema or the gateway stops making security decisions. Bound the claim - validation strength is capped by parser agreement, so put the real controls at the sinks and treat every added validating hop as a new differential to justify.
## The bug class This is a time-of-check/time-of-use bug in which what changes between check and use is not the data but its **representation**. The validator inspects one string; the consumer acts on the result of decoding, normalising, re-parsing or truncating that string. If any such stage sits between them, the check was performed on a value that never reaches the sink. That is the **parser differential**, and it is the invariant behind essentially every published validation bypass. State the consequence early, because it is what makes the answer senior: since the defect is a *disagreement between two components*, making the rule stricter does not help. You cannot regex your way out of two decoders. ## Concrete divergences Naming specific, divergent behaviours is the substance here. **Percent-encoding: how many passes, and who does them.** A component that decodes once and an application layer that decodes again during parameter binding disagree about `%252e`: the first decode yields the literal text `%2e`, which passes any check that looks for a dot, and the second decode yields `.`. Stacks also disagree on whether `+` means a space (true for form-encoded bodies and query strings, not for path segments), on whether decoding happens before or after splitting on delimiters (decode-then-split lets an encoded delimiter create an extra field, split-then-decode does not), and on whether encoded control characters are rejected or passed through. **Unicode normalisation.** The four forms (NFC, NFD, NFKC, NFKD) are not interchangeable. The compatibility forms fold characters that merely *look* related: U+FF0F FULLWIDTH SOLIDUS normalises to ASCII `/`, U+FF1C FULLWIDTH LESS-THAN SIGN to `<`, and ligatures such as U+FB01 decompose to `fi`. A validator running before normalisation sees an exotic codepoint that matches no rule; a consumer that normalises afterwards receives the metacharacter. Note that not every lookalike folds - many slash-like characters, U+2044 FRACTION SLASH among them, have no compatibility decomposition at all, which is exactly why you cannot reason about this by eye and must fix the form explicitly. Case folding compounds it: it is locale-sensitive, so uppercase `I` lowercases to a dotless `i` under a Turkish locale and an ASCII comparison silently fails, and it is non-injective, since sharp s full-case-folds to `ss`, collapsing two distinct inputs into one. **UTF-8 decoding.** Decoders differ on ill-formed input between rejecting, replacing with U+FFFD, and historically accepting overlong encodings - sequences that encode an ASCII character in more bytes than necessary, so a byte-level scan for a character misses it while a permissive decoder still produces it. Replacement is itself a transformation: substituting a character can destroy a token the validator relied on, or join two fragments into one. **Structured formats.** JSON parsers disagree on duplicate keys - most take last-wins, some first-wins, a few reject - so a validating proxy and the application can read different values from one document. They also differ on numeric forms and limits (very large integers, exponents, leading zeros, negative zero), which matters when one side checks a bound and the other rounds. XML adds entity expansion and CDATA, so the text seen before expansion is not the text processed after. Multipart parsers differ on filename and boundary handling. When a declared content type contradicts the bytes, some components trust the declaration and others sniff. **Layered decoders downstream.** A value that passes a server-side check may be decoded again by a different layer entirely - an entity decoder, a string-escape decoder, a URL parser inside another service - none of which the original validator simulated. ## The rules that follow **1. One canonical form, established once.** Decide the canonical representation per input (decoded, normalised to one Unicode form, one case convention), produce it at the boundary, validate that. Rules defined on a non-canonical form are rules about the wrong string. **2. Prefer rejecting non-canonical input over transforming it.** Rejecting anything not already well-formed fails closed and is cheap for machine-oriented fields. Normalising is a *transformation*, and transformations sit on the escaping rung of the ladder, not the structural one - they can create the token you are trying to exclude. The general trap is the removal-based sanitiser applied once rather than to a fixed point: given the rule "delete every occurrence of the forbidden two-character sequence `AB`", the input `AABB` becomes `AB` after a single pass, because deleting the inner occurrence leaves the outer characters adjacent and re-forms the token. Every delete-the-bad-thing filter over a multi-character token has this shape, whatever the token is - a comment introducer, a delimiter pair, a separator sequence in a resolved name. Iterating to a fixed point removes that specific trick but leaves you owing a proof that the fixed point is safe; rejecting the input that contained the token needs no such proof. That asymmetry is the whole reason "clean it up and continue" ranks below "reject it". **3. Validate after the last decode, as close to the sink as possible.** A validator several decode stages upstream is guessing about what the consumer will see. **4. Do not decode again after validating.** Every decode after the check reopens the gap. If a downstream component needs another encoding, encode outward, never decode inward. **5. Eliminate the differential rather than defending across it.** If a gateway and a service both parse the request, either they share one parser and one schema, or the gateway stops making security decisions and the service owns them. Two independent parsers over the same bytes is a standing bug waiting for one of them to be upgraded. ## The cross-stack claim This belongs to theory rather than to any one stack because the divergences appear in wholly unrelated processors - a URL decoder, a Unicode library, a UTF-8 decoder, a JSON parser, an XML entity expander, a multipart parser - and all produce the same failure by the same mechanism. That yields a claim no single-stack answer can make: **the strength of validation is bounded by the agreement between the validating parser and the consuming parser, not by the strictness of the rule.** Rule quality is second-order once agreement is lost.
- Your gateway already validates the request and the service validates it again. Is duplicate validation a defence in depth win or a liability?It is only defence in depth if both validators see the same canonical value and share one schema. If they parse independently, the second check does not cover the first's blind spots - it creates a second opinion about what the request means, and the attacker only needs the two opinions to differ. The safe shapes are: one parser and one schema shared by both, or the gateway stops making security decisions and the service alone owns them, with the gateway kept to routing and rate limiting.
- You must accept user text that legitimately contains non-ASCII characters, so rejecting non-canonical input is not an option. What then?Separate the two jobs. Normalise to one declared form (usually NFC) at the single boundary that owns the value, then validate and store that form and never re-normalise later. Rejection still applies to the parts of the input that are machine-oriented - identifiers, enum values, filenames, header values - which is most of what security decisions are made on. For genuinely free text, accept that validation is not the control at all: the safety comes from structural separation at each sink, which is why input validation alone is never sufficient.
- Someone proposes fixing a bypass by adding the specific bypass string to the validator's blocked patterns. What do you say?That fixes one instance of an unbounded class. The defect is that two components disagree about what the bytes mean, and that disagreement generates new encodings faster than patterns can be added - a different decode depth, a different normalisation form, a different overlong sequence. The correct fix is to collapse the differential: decide who decodes, decode once, validate the canonical result, and forbid any decode after the check.
saying these in an interview costs you the question
- Saying stricter regexes or more blocked patterns would have stopped it - the defect is disagreement between parsers, not a weak rule.
- Treating normalisation as free safety: it is a transformation and can manufacture the very metacharacter the rule excludes.
- Assuming one decode pass, or not knowing which component in the chain owns decoding.
- Validating early at the edge and calling it done, while later stages decode, normalise or re-parse the value again.
- Believing a single-pass strip of a forbidden sequence is sound - reconstruction after deletion is the standard bypass.