skip to content

In a system built from several services, which inputs actually count as untrusted, and where should validation happen — at the edge gateway, inside each service, or both? Include the input sources teams routinely forget.

level: seniorimportance: must knowfreq 55%

answer

  1. untrusted = who could influence the bytes, not which network
  2. forgotten: Host, forwarded-for, content type, filename, cookies
  3. stored data is re-emitted input — second order
  4. gateway = coarse limits, never authoritative; service = domain model
  5. validation ≠ authorisation — valid id, wrong tenant

basics

~20 s

Untrusted means "someone outside this component's authority could have influenced these bytes" — not "came from the internet". Validate at every trust boundary, in each service's own domain terms; an edge gateway can only do coarse, schema-and-size filtering and can never be the sole control.

solid answer

~60 s

A trust boundary is any point where data crosses from one authority domain into another, so the test is influence, not transport. Routinely forgotten sources: request headers the code makes decisions on (the Host used to build password-reset links, forwarded-for values feeding rate limits and audit logs, content type selecting a parser), cookies, upload metadata such as the filename and declared type as well as the bytes, deserialised object graphs, rows read back from your own database that an older release or another tenant wrote — stored input is re-emitted input — responses from upstream and third-party services, queue and webhook payloads, and anything "validated" in the client, which is attacker-controlled. Placement: both, with distinct jobs. The gateway enforces coarse schema and resource limits and cannot be authoritative, because it does not see service-to-service or asynchronous callers, holds no domain model, and drifts from the contract. Each service validates its own inputs at its own boundary. And validation never answers "may this caller use this value" — that is authorisation, and a perfectly valid identifier can belong to someone else.

go deeper

for a junior

List the untrusted sources beyond the request body — headers, cookies, uploaded filenames, anything a client can set — and say validation must run server-side.

for a middle

Explain the boundary as an authority crossing, add stored data and upstream responses as untrusted, and split gateway versus service responsibilities.

for a senior

Argue the placement from what each layer can know, cover asynchronous entry points, and separate validation from authorisation with a concrete example.

for a principal

Treat it as an architectural rule — no trusted internal network, per-service ownership of its own contract, gateways explicitly non-authoritative, and every new entry point required to declare its boundary and its validation.

## Define the boundary properly A **trust boundary** is a point where data crosses from a principal with one authority into a component with another. "Untrusted" is therefore a statement about *who could have influenced the bytes*, not about which network they arrived on. The moment you phrase it that way, the internal-network exception disappears: a service that receives a message from another service is receiving data influenced by whatever influenced that service, including a compromised deployment, a buggy release, or a tenant whose data flowed through it. ## The inputs teams forget Everyone validates the request body. The recurring incidents come from elsewhere. **Headers used in decisions.** The `Host` header, when used to construct absolute links, produces password-reset and verification URLs pointing wherever the attacker chose. Forwarded-client-address headers feed rate limiting, geo rules and audit trails, and are trivially set by the client unless a trusted-hop count is enforced. `Content-Type` often *selects the parser*, which means the client chooses which of your parsers processes their bytes. `Accept-Language`, `User-Agent` and `Referer` end up in pages, emails, log lines and analytics. **Cookies.** Client-side storage, therefore client-controlled. Structured cookie values are parsed by your code, so they are a parsing surface, not just a value. **Uploads — metadata as well as content.** The filename is a string the client wrote: it carries path components, encoding tricks, control characters and arbitrary extensions. The declared type is a claim, not a fact. Inside archives, entry names and entry sizes are attacker-controlled too, and so is the expansion ratio. **Deserialised payloads.** The dangerous part is the shape and the graph — types, depth, cycles, counts — not only the field values. Content that instantiates a structure before your validation runs has already had an effect. **Your own datastore.** This is the most under-appreciated one. Data was written by an earlier release with weaker rules, by an import job, by an administrator, or by another tenant. When it is read back and re-emitted or re-parsed, it is untrusted input again — the classic second-order pattern. "It came from our database" is not provenance; it is storage. **Upstream and third-party responses.** A partner API can return an oversized list, an HTML error page where JSON was expected, a URL you will fetch or redirect to, or a numeric field outside your assumptions — through a bug, an outage or a compromise. Clients that parse without limits are the reason one dependency's bad day becomes your outage. **Queue, event and webhook payloads.** Same as HTTP, minus the gateway that would have filtered them. Asynchronous entry points are consistently the least-validated surface in a system. **Anything checked in the client.** Browser and mobile validation is user experience. The client is under the user's control, so a client-side check is a hint, never a control. **Configuration and environment**, wherever a tenant or a lower-privileged operator can influence it. ## Where to validate Both, with different jobs — and the reason is what each layer can *know*. **The edge gateway** enforces coarse, universal properties: maximum body size, maximum nesting depth, allowed content types, schema well-formedness, rate limits, obviously malformed requests. It is genuinely useful because it sheds load and enforces resource bounds before your runtime spends memory. It cannot be authoritative, for four reasons: it does not see service-to-service, queue or scheduled-job traffic; it holds no domain model, so it cannot judge that a quantity is out of range for this product or that a state transition is illegal; it drifts from the service's contract as the contract evolves, and drift silently becomes either an outage or a hole; and if it re-parses the request independently of the service, the two parses can diverge — the differential problem. **Each service** validates its own inputs at its own boundary, in terms of its own domain, for every entry point it exposes — HTTP, message consumer, scheduled import, administrative interface. That is the authoritative layer, because it is the only one that knows what a valid value means. The governing principle is that there is no trusted internal network. Once one component is compromised or merely buggy, an internal caller is an attacker, and the only components that stay correct are the ones that checked. ## Two errors worth naming explicitly **"Validated once, trusted forever."** Validity is relative to a use. A string can be a valid display name and an invalid filename; a number can be a valid quantity and an invalid array index. Trust does not survive storage, transformation or a hop. Re-check at each boundary, especially after anything transforms the value. **Confusing validation with authorisation.** Validation answers "is this a well-formed value in the domain?". It never answers "may this caller act on it?". An object identifier can be perfectly well-formed and belong to another tenant; the strictest format rule in the world passes it. Insecure direct object reference is a bug that lives entirely on the far side of a green validation check, and saying so is one of the cleanest signals of seniority on this topic. ## Answering in an interview Define the boundary by influence rather than transport, list the forgotten sources with a concrete consequence for two or three of them, split the gateway and service responsibilities by what each can know, and close with the two errors above.

  • Why is a value read back from your own database still untrusted?
    Because the database records what was written, not what was verified. Rows may come from an earlier release with weaker rules, a bulk import, an administrator, another tenant, or a code path that skipped validation entirely. When such a value is read and re-emitted into a page, a query, a filename or a downstream request, it is untrusted input arriving through storage — the second-order pattern. Storage is not provenance, so the sink at the point of use must still apply its own control.
  • A team argues that because their gateway validates every request against an OpenAPI schema, the services behind it can skip validation. What is wrong with that?
    The gateway sees only the traffic that passes through it, so any message consumer, scheduled job, administrative path or service-to-service call is unvalidated. It also has no domain model, so it cannot judge range, state-transition legality or cross-field consistency, and it cannot make authorisation-dependent decisions at all. It drifts from the service's real contract over time, and if it re-parses the request independently it may read different values than the service does. Keep it for coarse limits and schema shape, and keep the authoritative check inside the service.

saying these in an interview costs you the question

  • Treating requests from inside the network or from another service as trusted.
  • Validating the request body while using headers, cookies and upload filenames unchecked.
  • "It came from our database, so it's clean" — storage is not validation.
  • Assuming a client-side check constrains what the server receives.
  • Conflating validation with authorisation — a valid identifier says nothing about who may use it.

context