skip to content

Should request validation run before or after authorization in a web framework's pipeline, and what does each order give away?

level: middleimportance: should knowfreq 50%

answer

  1. rights before shape
  2. a validation error is a schema hint
  3. parse before load, load before ownership
  4. referenced ids can leak across tenants

basics

~20 s

Authorize first for anything a stranger can reach. A validation error returned before the rights check hands an unauthorized caller your field names, accepted values and sometimes which identifiers exist; validating first only buys a friendlier error.

solid answer

~50 s

Two orders are possible and they trade different things. Validating first gives the caller the most complete error, but a `400` or `422` describes your schema, so an unauthorized caller learns field names, types, enum members and length limits, and the parsing work is spent on a request that will be refused. Authorizing first means such a caller only ever sees `401` or `403` and learns nothing about the body contract, at the cost of a second round trip for a legitimate client that got both wrong. The sequence most pipelines produce is: authenticate, apply the coarse route-level rights check, validate the syntactic shape, load the record, then apply the rules that need the loaded data. The trap is a validation message naming another resource, because it can confirm a record in someone else's tenant before any ownership check has run.

go deeper

for a junior

Know that a validation error describes your request format, so an API may refuse on credentials first and tell you nothing about the body until you are allowed in.

for a middle

Explain the ordering and its consequence: field names, enums and limits are disclosed by a validation error, so it belongs behind the rights check on any protected endpoint.

for a senior

Show the nuance: a validation rule that resolves a referenced identifier is data access in disguise and must sit behind the ownership check, or it confirms records across a tenant boundary.

for a principal

Set the policy across the surface: which endpoints are generous with error detail, which are terse, and how a team gets the safe order by default rather than by remembering it each time.

"Validate before or after authorization" sounds like an aesthetic preference about error quality. It is really a question about what an unauthorized caller is allowed to learn, and about where the work of a request is spent. ## What a validation error hands over A body-validation failure is a partial publication of your schema. Depending on how detailed the message is, a caller can learn: - which fields exist and what their types are; - which values an enumerated field accepts; - length, range and format constraints; - which combinations of fields are mutually required or exclusive; - that the endpoint exists at all and accepts this method and media type; - occasionally, that some referenced identifier does or does not exist. None of that is catastrophic on its own, and much of it may be in a published contract anyway. But on an internal or partner-only surface it is exactly the map an attacker wants, and it is free if validation runs before the rights check. ## What the two orders actually cost | Order | Unauthorized caller learns | Cost to a legitimate client | Work spent on a doomed request | |---|---|---|---| | Validate, then authorize | schema shape, sometimes referenced ids | none | parsing, validation, occasionally lookups | | Authorize, then validate | only that they are refused | one extra round trip if both were wrong | credential check only | For a public, unauthenticated write endpoint the first row is fine. For anything behind a credential, the second row is the default worth defending. ## The sequence a typical pipeline produces 1. Transport-level limits: body size and media type, before anything is read fully. 2. Authentication: the credential is verified and a caller established, or `401`. 3. Coarse authorization: the caller's roles or scopes against the route, or `403`. 4. Binding and syntactic validation: the body is parsed into the declared shape, or `400` / `422`. 5. The lookup: the identifier from the path becomes a record, or the not-found path. 6. Fine-grained authorization: rules that need the loaded record, such as ownership or state. 7. Business validation: rules that need both the body and the loaded data. Steps 1 and 2 cannot be reordered usefully, because a hook must be able to read the request before it can inspect it, and a rule needs a subject. Step 4 sits after step 3 by design, and before step 5 by necessity - you cannot load a record until the identifier that addresses it has been parsed. ## The cross-tenant validation trap The subtle failure is not the schema leak; it is a validation rule that *references other data*: 1. The request body carries an identifier belonging to another aggregate. 2. A validation rule resolves it to confirm it points at something real. 3. That resolution runs at step 4, before any ownership rule at step 6. 4. The caller learns whether that identifier exists anywhere in the system, including in another tenant's data. The fix is to move the existence-shaped part of the validation behind the ownership check, or to answer identically whether the referenced record is absent or invisible. A rule that checks only shape - format, length, character set - is safe at step 4 because the shape is public contract. ## What ordering rights first costs, honestly - A partner integrating against your API fixes the credential, retries, and only then discovers the body was wrong too. - Tooling that wants to report every problem at once cannot, because the chain short-circuits. - Error-driven client development becomes a little slower. Those are real costs, and they are the reason the opposite order keeps being proposed. The usual resolution is asymmetric: be generous with validation detail once the caller is authenticated and authorized, and be terse before that. ## A working rule - Anything a stranger can reach: refuse on credentials and coarse rights before describing the payload. - Anything that needs the loaded record: it runs after the lookup by necessity, so put the ownership rule immediately after the load and before any rule that could reveal what was loaded. - Anything that resolves a referenced identifier: treat it as data access, not as validation, and order it accordingly. - Keep the detailed reason in the log line for every terse response, so the terseness costs you nothing internally.

  • Why must the path identifier's shape still be checked before the ownership rule runs?
    Because the identifier has to be parsed before anything can be loaded, and the ownership rule needs the loaded record. Refusing a malformed identifier early is safe as long as the answer does not depend on whether a well-formed one exists - the format is public contract, the existence is not.
  • Does authorizing first mean an unauthorized caller never sees a validation error?
    Not entirely. Checks that must run before any hook can inspect the request - body size limits and unsupported media types - still answer first, because the server has to bound what it reads. What the caller stops seeing is the endpoint's field-level contract, which is the part worth withholding.

saying these in an interview costs you the question

  • Validates the body first so every caller gets the most useful error
  • Thinks a 400 leaks nothing because it carries no stored data
  • Assumes validation is too cheap to be worth refusing early
  • Puts ownership rules in the same stage as syntactic validation
  • Reports 'unknown identifier' for another tenant's record during validation