skip to content

questions

5

Explain the difference between allowlist (positive) and denylist (negative) input validation, why one of them is structurally stronger rather than merely better practice, and where the stronger one stops working.

level: juniorimportance: must knowfreq 80%

answer

  1. denylist = attacker's set = open-world = fail-open
  2. allowlist = your domain = closed-world = fail-closed
  3. format allowlist ≠ value allowlist
  4. free text: length + encoding only, sink owns the rest
  5. denylist survives as detection, the bottom rung

basics

~20 s

A denylist enumerates the attacker's moves — an open-ended, growing set — and fails open on anything unforeseen. An allowlist enumerates the application's own domain — finite and known — and fails closed. The difference is open-world versus closed-world, not optimism versus pessimism.

solid answer

~50 s

A denylist must be complete to be correct, and it can never be: alternate encodings, Unicode confusables, case and whitespace variants, comment syntaxes and the metacharacters of sinks you have not built yet all extend the set. Every omission is a bypass, and the failure mode is fail-open. An allowlist describes what your own domain accepts — a finite, design-time set the code owns — so anything unforeseen is rejected: fail-closed. Two refinements make the answer senior. First, an allowlist of **format** is not an allowlist of **value**: a pattern such as thirty lowercase letters constrains characters but still admits every real identifier in the system, so where the value selects structure the top rung is a finite enumerated map, not a pattern. Second, allowlists degrade where the legitimate domain really is "any text" — comment bodies, names, descriptions. There validation falls back to length and encoding well-formedness, and the security control moves to the sink. Denylists keep one honest role: detection, the bottom rung.

go deeper

for a junior

Define both, say the allowlist is the one to use because anything unexpected is rejected rather than accepted, and give a concrete example field.

for a middle

Explain fail-open versus fail-closed as the mechanism, note that the rule must be applied to the canonical form, and give the free-text limitation.

for a senior

Add that a format allowlist is not a value allowlist and that structure-selecting values need an enumerated map, and place denylists explicitly on the detection rung.

for a principal

Make it a schema-driven default — per-field domains declared once and enforced at every entry point — with enumerated maps mandated wherever input selects structure, and denylist signals routed to detection.

## The two shapes **Denylist (negative) validation** names what is forbidden and accepts the rest: reject strings containing `'`, `<`, `..`, `;`, the word `SELECT`. **Allowlist (positive) validation** names what is permitted and rejects the rest: accept exactly a 24-character hexadecimal identifier; accept an integer from 1 to 100; accept one of `PENDING`, `ACTIVE`, `CLOSED`. ## Why one is structurally stronger The usual answer — "allowlists are safer" — is a preference. The structural answer is about the size and ownership of the set you must enumerate. A denylist enumerates **the attacker's possibilities**. That set is open-world: it is not bounded, not known at design time, and not owned by you. It grows with every encoding the stack accepts (percent-encoding, double percent-encoding, HTML entities, Unicode escapes, overlong UTF-8), every normalisation the runtime performs (compatibility folding turning a fullwidth character into an ASCII one, locale-sensitive case folding), every equivalent syntax in a target grammar (comment forms, alternate quoting, whitespace variants, concatenation), and every new sink someone adds next quarter with metacharacters nobody listed. Correctness requires completeness against an adversary who is actively looking for the omission, and any single miss is a full bypass. The failure mode is **fail-open**: unknown input is accepted. An allowlist enumerates **your own domain**. That set is closed-world: finite, defined at design time, written down in the code, and changed only when you change it. Unforeseen input is rejected by construction — you do not have to have predicted it. The failure mode is **fail-closed**: unknown input is refused, which is at worst a bug report from a legitimate user and at best exactly right. Fail-open versus fail-closed under unforeseen input is the whole argument, and it is why this is a structural property rather than a heuristic preference. ## Refinement one: format allowlist is not value allowlist This is the part that separates a rehearsed answer from a real one. A pattern such as `^[a-z_]{1,30}$` is closed with respect to *characters* and still wide open with respect to *which value*. If the value picks a column to sort by, a template to render, a file to serve, or a host to redirect to, then every legitimate name in the system matches the pattern — including the ones the caller must never reach. Where the untrusted value selects structure rather than supplies data, the top rung is a **closed-world enumeration**: a finite map from the public name the client may send to the internal value the code owns, rejecting anything absent from the map. The reason, again, is open-world versus closed-world: the pattern constrains the alphabet and still admits an unbounded set of meaningful targets, while the map is finite and owned by the code. Say that explicitly and you have said the thing most candidates miss. ## Refinement two: where allowlists stop working Allowlists are excellent for structured fields — identifiers, enumerations, dates, quantities, codes, references. They break down when the legitimate domain genuinely is "any text": a support-ticket body, a product description, a person's name. Real names contain apostrophes, hyphens, accents, non-Latin scripts and, increasingly, emoji; a product description legitimately contains angle brackets and ampersands. For those fields, honest validation reduces to: - **Well-formedness** — valid encoding, no unpaired surrogates, no unexpected control characters. - **Bounds** — maximum length, maximum line count. - **Sometimes a semantic check** — a language or script restriction where the product genuinely requires it. and the security control moves entirely to the sink, which applies context-correct handling at the point of use. The temptation to "sanitize" free text by stripping characters is the worst option available: it corrupts legitimate data (the customer named O'Brien), it creates false confidence, and a remover can synthesise the very token it removes when the input is constructed so the remainder re-forms it after deletion. Reject or store faithfully; do not quietly mutate. ## The honest role of denylists A denylist is a legitimate **detection** control — the bottom rung of the ladder. Logging or alerting when input contains patterns that have no business in your domain gives you an attack signal, and it costs nothing as long as nobody mistakes it for the control. Two rules keep it honest: it must never be the only thing standing between input and a sink, and it must not silently alter the input. ## Practical checklist - Define the allowlist per field, from the domain, not globally. - Define it on the **canonical** form of the input, after decoding and normalisation — a rule applied to a non-canonical representation is a rule about the wrong string. - Reject rather than repair; return a stable error naming the field. - Where the value selects structure, use an enumerated map, not a pattern. - Server-side always; client-side checks are user experience, because the client is attacker-controlled. - Keep the denylist, but wire it to alerting rather than to the request path.

  • A comment field must accept arbitrary user text. What does validation do there?
    It checks well-formedness and bounds — valid encoding, no unpaired surrogates or stray control characters, a maximum length and maybe a line limit — and nothing about content, because every character is legitimate in prose. The security control moves entirely to the point of use, where each sink applies its own context-correct handling to the stored text. Stripping characters to make the text "safe" is the wrong move: it damages legitimate data, gives false assurance, and a remover can be defeated by input crafted so the remaining characters re-form the removed token.
  • Is there any legitimate use for a denylist?
    Yes, as detection — the bottom rung of the defence ladder. Patterns that have no business appearing in your domain make a good alerting signal, and a spike is early warning of probing. The conditions are that it never sits in the request path as the only control, and that it never silently modifies input; it observes and reports while allowlisting and the sink controls do the preventing.

A denylist is a guest list of people who are barred; an allowlist is the list of people invited. Only the second one works when someone shows up whose name nobody had ever heard.

saying these in an interview costs you the question

  • "We block SQL keywords and script tags" — an open-ended set, and every omission is a bypass.
  • Treating a character-class regex as equivalent to choosing from a fixed list of allowed targets.
  • Applying one global allowlist to every field instead of a per-field domain rule.
  • Silently stripping forbidden characters rather than rejecting, which corrupts data and can synthesise the removed token.
  • Relying on a client-side pattern attribute or JavaScript check as the validation.

context

open as a page

Input validation is often taught as "strip the dangerous characters before you use the value". Give the definition of validation that explains what it can and cannot guarantee, and why treating it as the fix for injection is the wrong frame.

level: middleimportance: must knowfreq 72%

basics

~20 s

Validation is a closed-world constraint applied at a trust boundary in terms of your own domain, without knowing the grammar of whatever will eventually consume the value. It shrinks the input space and owns resource limits; it cannot guarantee safety at a parser it has never seen.

open as a page

Why must untrusted input be canonicalised before it is validated, and what class of bug appears when validation runs on the raw form? Give concrete examples of processors that disagree about what the same bytes mean.

level: seniorimportance: must knowfreq 50%

basics

~20 s

Because a rule applied to one representation says nothing about a different representation of the same value. Every bypass is a parser differential: the validator and the consumer decoded differently. Canonicalise once, validate the canonical form, never decode again downstream.

open as a page

In a system built from several services, which inputs actually count as untrusted, and where should validation happen — at the edge gateway, inside each service, or both? Include the input sources teams routinely forget.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Untrusted means "someone outside this component's authority could have influenced these bytes" — not "came from the internet". Validate at every trust boundary, in each service's own domain terms; an edge gateway can only do coarse, schema-and-size filtering and can never be the sole control.

open as a page

You own a public API that must accept free-text fields, arbitrary JSON documents and file uploads. Design the validation strategy: what fails closed, which layers exist, and how do you stop the validation code from becoming a vulnerability in its own right?

level: principalimportance: should knowfreq 40%

basics

~20 s

Declare one schema enforced identically at every entry point; impose size, depth, count and expansion limits during parsing, not after; keep free text unmodified and let sinks handle it; identify files by parsing, not by extension; and treat the validator itself as attack surface — bounded regexes, no external entity or schema fetching, reject rather than coerce.

open as a page