skip to content

In an OPA Envoy ext_authz policy, where in input do the request method, path and body live?

level: juniorimportance: must knowfreq 72%

answer

  1. one request, one policy query
  2. Envoy's check request becomes input.attributes
  3. method, path, headers, body under request.http
  4. parsed_ fields sit at input's top level
  5. header names arrive lowercased

basics

~10 s

Under input.attributes.request.http: method, path, headers with lowercased keys, and body as a string. The OPA Envoy plugin also adds three top-level conveniences: input.parsed_path, input.parsed_query and input.parsed_body.

solid answer

~40 s

OPA runs as Envoy's external authorization service, so each request becomes one policy query. Envoy's `CheckRequest` attributes land verbatim under `input.attributes`: `input.attributes.request.http.method`, `.path`, `.host`, `.headers` (a flat map with lowercased header names) and `.body` as a raw string, plus peer information under `input.attributes.source` and `input.attributes.destination`. On top of that the OPA plugin adds three pre-parsed fields at the *top* level of input, not under attributes: `input.parsed_path`, an array of URL-decoded path segments with the query removed; `input.parsed_query`, a map from parameter name to an array of values; and `input.parsed_body`, the decoded JSON body. They exist because `path` is the raw request target and `body` is an undecoded string, and re-parsing those correctly in Rego on every request is exactly where policy bugs come from.

code

json · 16 lines
json
{
  "attributes": {
    "source": { "principal": "spiffe://prod/ns/web/sa/frontend" },
    "request": {
      "http": {
        "method": "POST",
        "path": "/api/v1/orders?dry_run=true",
        "headers": { "content-type": "application/json", "authorization": "..." },
        "body": "{\"sku\":\"A-1\",\"qty\":3}"
      }
    }
  },
  "parsed_path": ["api", "v1", "orders"],
  "parsed_query": { "dry_run": ["true"] },
  "parsed_body": { "sku": "A-1", "qty": 3 }
}

go deeper

for a junior

Be ready to name the path to the basics without hesitating: input.attributes.request.http for method, path, headers and body, and the three top-level parsed_ fields the plugin adds.

for a middle

Explain why the parsed_ fields exist at all — raw path carries the query string and percent-encoding, raw body is an undecoded string — and what makes each of them undefined.

for a senior

Show that you treat input as the complete and only evidence: nothing about other requests, no session, no upstream response, so anything else has to be data you loaded into the engine deliberately.

for a principal

Own the contract question: which request attributes your platform guarantees will always be present for policy, and what it costs the proxy to supply them, so teams do not write rules against fields that are only sometimes there.

## What is being evaluated With OPA's Envoy external authorization plugin, OPA implements the gRPC external authorization service that Envoy calls once per HTTP request, before the request is routed upstream. Envoy sends a `CheckRequest` describing the request; OPA evaluates a Rego decision document against it and answers. Everything the policy can see about the request is in the `input` document built from that `CheckRequest` — the policy is not looking at a live connection, and it cannot go and fetch more. ## The attributes tree The `CheckRequest`'s attribute context is copied under `input.attributes`: - `input.attributes.request.http.method` — the HTTP method, uppercase (`GET`, `POST`, ...). - `input.attributes.request.http.path` — the **raw request target**. This is the value as it appeared on the wire: it still carries the query string, and it is still percent-encoded. - `input.attributes.request.http.host` — the authority the client asked for. - `input.attributes.request.http.headers` — a flat object mapping header name to value, with **names lowercased**. A rule that looks up `Authorization` finds nothing; the key is `authorization`. - `input.attributes.request.http.body` — the request body as a string, present only when the proxy has been told to buffer the body for the authorization check. If it was not, there is no body to read, and a rule that depends on one is undefined. - `input.attributes.source` and `input.attributes.destination` — the two peers of the connection, each with an address and a `principal`. `source.principal` is the authenticated identity of the caller as the proxy established it, for example a SPIFFE URI taken from the peer certificate on an mTLS connection. It is empty when the connection was not mutually authenticated. ## The three parsed_ conveniences The plugin adds three fields at the top level of `input`, **not** inside `attributes`: - `input.parsed_path` — the path with the query string removed, split on `/`, each segment percent-decoded. For `GET /api/v1/orders?page=2` it is `["api", "v1", "orders"]`. - `input.parsed_query` — a map from query parameter name to an **array** of values, because a name may legitimately appear more than once. - `input.parsed_body` — the body decoded as JSON, so a rule can write `input.parsed_body.account.id` instead of running a string parse itself. It is populated when the body is present and the content type says JSON. They exist because the two things a policy most wants to match on — a path and a JSON field — are the two things that arrive in a form you cannot compare against safely. Matching `input.attributes.request.http.path` against a literal string breaks the moment a caller appends a query parameter, and it can be sidestepped by percent-encoding a character in the path. Matching `input.parsed_path` against a list of segments has neither problem. ## What input does not contain Just as important for a policy author: the input describes **this request and nothing else**. There is no other in-flight request, no session, no history of what this caller did a minute ago, and no upstream response — the decision is made before the request is forwarded, so the policy cannot see what the service would have replied. Any fact that is not on the request has to come from data loaded into OPA (a bundle of roles, an allowlist of identities) rather than from `input`. ## Writing against it Two habits follow directly. First, always start the decision document with an explicit default deny, because an undefined rule produces *no result*, not `false` — leaving the answer to the enforcer rather than to your policy. Second, prefer the parsed fields for matching and reserve the raw ones for logging or for cases where you deliberately want the bytes as sent. If you need to debug what a rule actually saw, OPA's decision logs record the `input` it evaluated alongside the result, which is usually faster than guessing at the shape from the rule. ``` input.attributes.request.http.method "GET" input.attributes.request.http.path "/api/v1/orders?page=2" input.parsed_path ["api", "v1", "orders"] input.parsed_query.page ["2"] ``` Knowing this shape cold is the price of entry: nearly every mistake in a request-path policy is a rule reading the right idea out of the wrong field.

  • Why does the plugin bother adding parsed_path when path is already there?
    Because `path` is the raw request target: it carries the query string and is still percent-encoded. Comparing it to a literal breaks as soon as a caller appends `?page=2`, and an encoded character can make an intended match miss. `parsed_path` gives you the query-free, decoded segments as an array, so a segment comparison means what you think it means.
  • A rule reads input.parsed_body and never fires. What are the two likeliest causes?
    Either the proxy was never told to buffer the request body for the authorization check, so there is no body in the check at all, or the body is present but not JSON, so nothing was decoded. In both cases the reference is undefined, the rule body fails, and — with a default deny — the request is refused for a reason that has nothing to do with its content.
  • Can the policy see the response the upstream would have sent?
    No. External authorization runs before the request is forwarded, so the input describes only the inbound request: method, path, headers, an optionally buffered body, and the connection peers. Anything the decision needs beyond that — roles, allowlists, ownership — must be data loaded into OPA, not something fetched from the service being protected.

The attributes tree is the envelope exactly as it was posted; the parsed_ fields are the same envelope already opened, with the address split into lines for you.

saying these in an interview costs you the question

  • Thinking input.path or input.method exist at the top level
  • Assuming the query string is stripped from request.http.path
  • Looking up a header by its capitalised name
  • Expecting a body to be present when the proxy never buffered one
  • Believing the policy can see the upstream's response

context