skip to content

What does opa build --optimize do to a policy, and what must you give it?

level: middleimportance: should knowfreq 32%

answer

  1. partial evaluation pointed at build time
  2. input unknown, bundled data known
  3. the optimiser needs a target rule
  4. entrypoint is mandatory with -O
  5. inlined values are frozen at build

basics

~20 s

It runs partial evaluation at build time, specialising the policy for named entrypoints against the data packaged in the bundle. You must supply at least one entrypoint; without one there is nothing for the optimiser to specialise towards.

solid answer

~50 s

`opa build -O=1` or `-O=2` applies the same partial-evaluation machinery you would use for pushdown, but in the other direction: `input` is the unknown, the `data` shipped in the bundle is known, and OPA precomputes everything that does not depend on the request. The output is a generated policy specialised for the entrypoints you named — which is why `--entrypoint` is mandatory with optimisation; the optimiser needs to know which rules matter. Higher levels inline more aggressively, which trades build time and bundle size for a cheaper decision at request time. The catch worth stating in an interview: anything inlined is frozen into the generated policy, so a value you intend to replace at runtime must not be visible to the optimiser. And the Rego that runs is no longer the Rego you wrote, so run your tests against the built bundle, not just the source.

go deeper

for a junior

Know that a policy bundle can be built in an optimised form that is faster to evaluate, and that the flag is a build-time choice rather than something you turn on in the running engine.

for a middle

Be ready to explain the mechanism: input is the unknown, bundled data is known, an entrypoint is required, and higher levels trade build time and bundle size for cheaper decisions.

for a senior

An interviewer expects the inlining trap and its symptom — a data push that stops taking effect — plus the discipline of testing the built artifact rather than only the source.

for a principal

Own the judgment that optimisation is a tuning step against a stated latency number, and that a debuggable unoptimised bundle is the right default for gates that run once per change.

## The same machinery, pointed the other way Partial evaluation is usually explained as "defer what you cannot decide". Build-time optimisation is the mirror image: **decide everything you can, now, so the request path does less**. `opa build --optimize` (short form `-O`) treats `input` as the unknown and the `data` packaged into the bundle as known, evaluates the policy as far as it can go, and writes out a generated policy that is equivalent but pre-specialised. ## What you must supply An **entrypoint**. Optimisation requires at least one `--entrypoint` (`-e`) naming the rule the bundle exists to answer, for example `-e backup/retention/compliant`. Without it, OPA has no target to specialise towards — every rule in the package would have to be preserved in full generality, which is the unoptimised case. This trips people up because a plain `opa build` needs no entrypoint at all; adding `-O` changes the requirement. ## What the levels mean `-O=1` is conservative and `-O=2` is aggressive: the higher level does more inlining of `data` documents and more specialisation of rule bodies. Practically, more optimisation means: - **longer builds** — the optimiser is doing evaluation work up front; - **potentially larger bundles** — inlined data is copied into the generated policy, sometimes several times over; - **cheaper evaluation at request time** — the work you paid for at build time is not repeated per decision. That is the whole trade, and which side you want depends on how latency-sensitive the decision path is and how often the bundle is rebuilt. ## The trap: inlined data is frozen This is the failure that catches teams. If a `data` document is visible to the optimiser, its values can be inlined into the generated policy. Pushing a new version of that data at runtime does not retroactively change the copies that were baked in — the decision keeps using what was compiled. So any value designed to change independently of the policy build, a per-tenant list or a threshold you tune without a release, must not be part of what the optimiser sees; it belongs in `input`, or in a data path deliberately kept out of the optimised build. ## The second trap: what runs is not what you wrote The generated policy is machine-produced. It is valid Rego, but it is not written for a human: rule names are synthesised, bodies are expanded, and reading it to debug a wrong decision is unpleasant. Two consequences. First, **test the built bundle**, not only the source modules, so your tests exercise the artifact that will actually answer requests. Second, when you are diagnosing a wrong answer, reproduce it against the unoptimised policy first; if the two disagree, that is a much more interesting bug than the one you started with, and it usually points at data that should not have been inlined. ## When to reach for it Reach for optimisation when the decision path is genuinely latency-sensitive — a sidecar or library evaluating on every request — and the policy does meaningful work over data that is stable between builds. Do not reach for it as a default. An unoptimised bundle is readable, debuggable, and fast enough for the overwhelming majority of gates that run once per change rather than once per request. Optimisation is a tuning step you apply when you have a number you are trying to hit, not a box you tick because the flag exists.

  • Why does optimisation require an entrypoint when a plain build does not?
    Because specialisation is relative to a question. The optimiser precomputes the parts of a specific rule's evaluation that do not depend on `input`; with no named rule it would have to preserve every rule in full generality, which is just the unoptimised bundle. Naming the entrypoint tells OPA what the bundle exists to answer.
  • A threshold you push as data stopped taking effect after you enabled -O=2 — why?
    Because the optimiser inlined the value it saw at build time into the generated policy. Pushing new data does not rewrite the copies that were baked in. Either keep that value out of the optimised build so it stays a live lookup, or accept that changing it now requires rebuilding and redeploying the bundle.
  • How should testing change once you ship optimised bundles?
    Run the policy tests against the built artifact as well as the source, because the Rego that answers requests is generated code, not the modules you wrote. If an optimised and unoptimised build ever disagree on the same input, treat that as a defect in what you allowed the optimiser to inline rather than a curiosity.

saying these in an interview costs you the question

  • Enables optimisation by default without a latency target
  • Expects runtime data pushes to override inlined values
  • Does not know an entrypoint is required with -O
  • Debugs a wrong decision by reading the generated policy
  • Tests only the source modules, never the built bundle

context