You own a public API that must accept free-text fields, arbitrary JSON documents and file uploads. Design the validation strategy: what fails closed, which layers exist, and how do you stop the validation code from becoming a vulnerability in its own right?
answer
- limits during streaming, not after parse
- one schema compiled into every entry point
- unknown fields: reject on security-relevant endpoints
- free text: bound and store faithfully, never strip
- the validator is attack surface — and reject, never coerce
basics
~20 sDeclare one schema enforced identically at every entry point; impose size, depth, count and expansion limits during parsing, not after; keep free text unmodified and let sinks handle it; identify files by parsing, not by extension; and treat the validator itself as attack surface — bounded regexes, no external entity or schema fetching, reject rather than coerce.
solid answer
~60 sLayer it. **Before parsing**: byte-size caps, content-type restriction, decompression-ratio budget, and streaming limits on nesting depth, member counts and string length — this is the one area with no downstream control, so it must fail closed. **Schema layer**: a single declarative contract compiled into every entry point (HTTP, queue, import) so there is exactly one parse and no differential; reject unknown fields on security-relevant endpoints rather than ignoring them, since silent ignoring enables mass assignment. **Domain layer**: types, ranges, cross-field rules, state-transition legality — inside the service. **Free text**: check encoding well-formedness and length, then store faithfully and let each sink apply its own control; never strip. **Files**: identify by parsing or transcoding with a hardened parser, re-encode, store under a code-generated name, serve from a separate origin with a fixed content type. **The validator itself**: catastrophic regex backtracking, entity-resolving XML validators, remote schema fetches, unbounded normalisation and input echoed back in error messages are all vulnerabilities in the checking layer. Fail with a stable, field-named error; never coerce or truncate.
go deeper
Cover the basics: size limits, a schema for the body, allowed content types, server-generated filenames for uploads, and clear rejection rather than silent fixing.
Add layering with a stated failure mode per layer, the unknown-field decision, and why free text is bounded rather than stripped.
Emphasise limits enforced during streaming, one schema across all entry points to avoid differentials, file identification by parsing, and the validator's own attack surface.
Argue the ownership model — validation strict where it is the only control, permissive where the domain is open, never the barrier in front of a sink — with contract-versioning policy, per-rule rejection metrics, and denylists confined to detection.
## Start from what fails closed A validation strategy is a set of decisions about what happens to input nobody anticipated. Design it as layers, each with a stated failure mode. ### Layer 0 — resource limits, enforced before and during parsing This layer is special because **nothing downstream can compensate for it**. No amount of parameterisation or context-correct output handling makes a four-gigabyte body, a million-element array or a thousand-level-deep document safe. Enforce: - maximum request bytes, and a maximum per part for multipart; - allowed content types, with the server choosing the parser rather than the client's declaration choosing it; - maximum nesting depth, member count, array length and individual string length; - decompression ratio and absolute expansion budget for any compressed or expandable format; - a total time budget for parsing. Crucially these must be enforced **while streaming**, not by parsing the document and then measuring it — measuring afterwards means the attack already succeeded. ### Layer 1 — one schema, one parse Declare the contract once, declaratively, and compile it into every entry point: the HTTP handler, the message consumer, the batch import, the administrative tool. Two hand-written validators over the same payload is how parser differentials are born, and any component that re-parses independently is a standing risk. The live decision here is **unknown fields**. Ignoring them silently hides client bugs and enables mass assignment when a later binding layer does consume them; rejecting them is the closed-world choice and gives clients an immediate, honest error, at the cost of forward compatibility for clients that add fields optimistically. The defensible position is: reject on endpoints where fields drive authorisation, pricing, state or identity, and version the schema so clients have a supported path to add fields. ### Layer 2 — the domain Types, ranges, formats, cross-field consistency, legal state transitions. This lives inside the service because only the service knows it. It is also where enumerated maps belong: any field that selects structure — a sort column, a template, a redirect target, a report type — resolves through a finite, code-owned map rather than a character pattern, because a pattern is open-world with respect to which value it admits. ### Layer 3 — free text Accept it. Check that the encoding is well-formed (valid UTF-8, no unpaired surrogates, control characters excluded unless meaningful) and bound the length, then store exactly what the user typed. Do not strip, do not "clean", do not replace. Stripping corrupts legitimate data, produces false confidence, and a removal step can synthesise the token it removes when the input is crafted so the remainder becomes adjacent. Every sink applies its own context-correct control at the point of use. If the product truly requires rich markup, the control is to parse the input with a real parser and re-serialise from a strict structural model, discarding anything not in the permitted set — a parse-and-rebuild, never a pattern filter. That is a sink-side transformation, and it should be applied once, at a defined layer, with its output treated as the canonical stored form. ### Layer 4 — files The client's declared type and the file extension are claims. Identity comes from **parsing**: decode the file with a hardened, bounded parser and, where the format allows, transcode and re-encode it, discarding metadata. Store under a name the code generated, never one the client supplied. Serve from a separate origin with a fixed content type and no sniffing. Bound decompression for archives, and check entry names and entry sizes before extraction. Magic-byte sniffing alone is not identification — polyglot files satisfy two formats at once. ## The validator as attack surface This is the part most candidates never reach, and it is what the question is really probing. - **Catastrophic backtracking.** A pattern with nested quantifiers over attacker-controlled length can take exponential time — a denial of service delivered *by the security control*. Mitigations: bound length before matching, anchor patterns, avoid nested quantifiers and ambiguous alternation, prefer a linear-time engine, and enforce a matching timeout. - **XML and schema processing.** Validators that resolve external entities or DTDs turn the check into a file-read and outbound-request primitive; validators that fetch remote schemas by URL turn it into a server-side request primitive with the server's network position. - **Unbounded canonicalisation.** Normalisation and decompression during the canonicalisation step can expand input dramatically. Budget it like any other parse. - **Error handling.** Never echo the invalid input back verbatim — that is a reflection sink. Never leak internal identifiers, stack traces or schema internals. Use uniform errors where a distinguishable message would reveal the existence of a resource or an account. - **Logging.** A validator that writes rejected input into logs feeds it to a whole downstream stack — log processors, dashboards, alerting. Encode or bound what you log. ## Failure behaviour **Reject; never coerce.** Silent truncation is the classic self-inflicted differential: the validated value and the stored value differ, and the two systems now disagree — a length limit applied after the check is the standard instance. Silent type coercion (a string accepted where a number was declared, a number rounded) hides client bugs and produces values nobody validated. Return a stable, machine-readable error naming the field and the rule, and apply nothing partially — an entire request either takes effect or does not. ## Operating it Emit a counter per rule. A rule that fires constantly is either wrong, poorly documented, or under attack, and each of those is worth knowing. This is where denylists earn their keep: as detection signals — the bottom rung of the ladder — never as the control in the request path. ## The one-sentence position Validation is strict where it is the *only* control — resource bounds and the domain model — permissive and non-mutating where the domain is genuinely open, and never the thing standing between an attacker and a sink; the sink owns its own safety.
- Should unknown fields in a request body be rejected or ignored?Reject on endpoints where fields influence authorisation, pricing, identity or state, because silently ignoring them conceals client bugs and turns into mass assignment the moment some later binding layer does read them. Ignoring is defensible for telemetry-style or append-only endpoints where forward compatibility matters more than strictness. Either way it must be an explicit, documented contract decision with schema versioning as the client's path to add fields, not an accident of whichever parser configuration was the default.
- How can validation code itself cause an outage?Most commonly through catastrophic regular-expression backtracking: a pattern with nested quantifiers over an attacker-controlled string can consume exponential time, so a handful of requests saturate the CPU. Other paths include XML validators that resolve external entities or fetch remote schemas, unbounded normalisation or decompression during canonicalisation, and validation that parses an entire document before checking its size. Bound the input length before matching, use anchored patterns and linear-time engines with timeouts, disable entity and remote-schema resolution, and enforce limits during streaming rather than after the parse.
- Why is silently truncating an over-long value worse than rejecting it?Truncation creates two different values in two places: the one that was validated and the one that was stored or forwarded. Anything that later compares, joins on, or re-validates the field now disagrees with the component that truncated, which is the same parser-differential failure mode that defeats validation generally. It also destroys user data without telling anyone and hides a genuine client bug. Reject with a clear error naming the field and the limit.
saying these in an interview costs you the question
- Measuring document size or depth after fully parsing it rather than during streaming.
- Hand-writing a second validator at a different layer instead of compiling one schema everywhere.
- Sanitising free text by stripping characters instead of storing it faithfully and handling it at each sink.
- Identifying uploaded files by extension or by magic bytes alone rather than by parsing and re-encoding.
- Unbounded or nested-quantifier regexes over attacker-controlled input, and echoing rejected input back in error messages.