When should you call OPA's /v1/compile instead of /v1/data?
answer
- one endpoint answers with conditions
- you name what you cannot supply
- no boolean comes back
- empty result means never true
basics
~20 sWhen you cannot supply every fact at query time. /v1/data needs a complete input and answers with a value; /v1/compile takes a query plus the references you name as unknown and answers with the conditions that are still left to check.
solid answer
~50 s`/v1/data` is the ordinary path: give OPA everything, get a verdict. `/v1/compile` is for the case where one fact is not available yet - you send a `query` string, whatever `input` you do have, and an `unknowns` list naming the references you cannot resolve. Instead of true or false you get back `{"result": {"queries": [...], "support": [...]}}` - the residual conditions under which the query would hold. Two results are worth memorising: an empty result object means the query can never be true, so you can answer no immediately, and a result whose single condition set is empty means it is unconditionally true. The catch is entirely on the client side. A residual is not a decision, so something you own has to consume it - evaluate it once the missing fact arrives, or translate it into a check of your own. Without that consumer the endpoint buys you nothing over `/v1/data`.
go deeper
Just know that OPA can be asked in two ways: /v1/data answers with a value when you can supply every fact, and /v1/compile answers with leftover conditions when you cannot. You will not be expected to use the second one yet.
Explain the request fields - the query string, the optional input, and the unknowns you name - and be clear that the reply carries residual conditions rather than a true or false verdict.
Show judgment about when the extra machinery pays for itself. Name what your client would do with a residual, and be willing to say that with no consumer for it the endpoint is a moving part you should not add.
Weigh whether the platform should expose this at all. It hands callers a piece of policy evaluation to run themselves, which spreads the decision boundary across services you do not operate and quietly multiplies the places a rule has to be right.
## Two ways to ask, and what each one costs the caller The Data API assumes you can state the whole world before you ask. You POST `{"input": ...}` to `/v1/data/<path>`, OPA evaluates the document, and you get a value. That assumption holds for most gates: a provisioning service knows the database spec it is about to submit, so it can ask whether a thirty-day backup retention rule permits it and get back `{"result": false}`. Sometimes it does not hold. One fact the rule depends on is not knowable at the moment you want to ask - it is decided by a downstream system, it arrives later in the workflow, or it lives somewhere your caller cannot cheaply enumerate. The naive workarounds are both bad: guess the missing value and get a decision about a world that may not happen, or wait until you know and discover the answer too late to act on it. `/v1/compile` is the endpoint for that situation. ### What the caller sends A POST to `/v1/compile` with a JSON body carrying: - **`query`** - the Rego query you want an answer to, as a string, for example `data.db.policy.allow`. Note this is a query, not a URL path; the document you are asking about is named inside the body rather than in the URL. - **`input`** - optional, and exactly the same shape as a Data API request's input: everything you *do* know. - **`unknowns`** - an array naming the references OPA must not resolve, written as reference strings such as `input.database.retention_days`. If you leave this out, OPA treats the entire `input` document as unknown, which is almost never what you want: everything you carefully supplied is then ignored and the residual you get back is far larger and far less useful than it should be. Name the specific references you cannot supply and pass the rest as ordinary input. ### What comes back The reply is `200` with `{"result": {"queries": [...], "support": [...]}}`. `queries` holds the residual conditions - the parts of the policy that could not be settled without the unknown values, expressed as a disjunction of condition sets. `support` holds any rule definitions the residual still refers to. Two results carry more meaning than the rest: - **No queries at all** (an empty result object) means the query can never be true given what you fixed. There is nothing left to check; you can answer *no* right now. - **A single, empty condition set** means the query is unconditionally true. Also nothing left to check; you can answer *yes* right now. Anything in between is real work handed back to you. ### The client-side obligation, which is the whole point This is where interviewers separate people who have used the endpoint from people who have read its description. `/v1/compile` does not return a decision. It returns conditions. That means **you must own a consumer for them**, and there are only a few honest shapes for one: - Hold the residual and evaluate it later, once the missing fact is known - which means you have taken on a piece of policy evaluation inside your own service. - Translate the residual into a check or a filter your own layer already knows how to run. - Recognise only the two degenerate answers above and treat everything else as 'ask again later with complete input'. This is a legitimate, cheap and underrated option: you get an early definite no without committing to interpreting anything. If your integration has none of these, the endpoint is a moving part that adds latency and a second response shape to handle without changing any outcome. A client written against `/v1/data` that starts calling `/v1/compile` will break immediately, because it is looking for `result` to be a boolean and instead finds an object of conditions. ### How to talk about the tradeoff The honest framing is that `/v1/compile` moves part of the evaluation across a service boundary. The engine answers as far as it can and hands the rest to the caller. That is powerful when the caller is genuinely better placed to finish the job, and it is a liability when it means several teams each grow their own half-implementation of the policy's tail end. On a platform, the question is not only 'can we use it' but 'do we want the decision boundary to be in one place, and are we willing to spread it'. For everyday gates the answer stays the same: gather the facts, POST to `/v1/data`, read `result`, and handle a missing `result` as an outage. `/v1/compile` earns its place only when you can genuinely name what you do not know and you have somewhere to put the answer.
- What does an empty result object from /v1/compile tell the caller?That the query can never be true given the facts you fixed, so nothing is left to check and you can answer no immediately. The mirror case is a single empty condition set, which means unconditionally true - answer yes without checking anything. Recognising just those two is often enough to be useful.
- What must your client be able to do before this endpoint is worth calling?Consume a residual. You are handed conditions rather than a verdict, so something you own has to evaluate them once the missing fact arrives or translate them into a check your own layer runs. Without that consumer you have added a second response shape and some latency for no change in outcome.
- If you omit the unknowns field, what does OPA assume?That the whole `input` document is unknown. Everything you did supply is then left unresolved, and the residual you get back is much larger and much harder to use than it should be. Name the specific references you cannot resolve and pass the rest as ordinary input.
saying these in an interview costs you the question
- Expects /v1/compile to answer true or false
- Calls compile for ordinary complete-input decisions
- Leaves unknowns unset and wonders why nothing resolved
- Assumes OPA will fetch the missing facts itself