skip to content

What lets OPA index the rules of a Rego document, and which body shapes defeat it?

level: middleimportance: nice to knowfreq 31%

answer

  1. many rules share one name
  2. equality against a literal constant
  3. ground reference into input
  4. helpers and negations hide the key

basics

~20 s

OPA indexes rules that share a name using equality tests between a plain reference into input or data and a constant, so a query only evaluates the rules whose key matches. Negations, comprehensions, walks and helper calls cannot be indexed.

solid answer

~50 s

When a package defines many rules under one name — say forty `deny` rules — OPA does not have to evaluate all of them. It builds an index from the expressions in each body that it can interpret, chiefly an equality between a ground reference into `input` or `data` and a constant value, plus `glob.match` against a constant pattern. A query then selects only the rules whose keys match the input; a rule with no indexable expression has no key and is evaluated every time. What defeats it is hiding the discriminator: moving `input.request.kind.kind == "Namespace"` into a helper function, expressing it as a negation, deriving the compared value with a comprehension or arithmetic, or comparing against a variable. Indexing prunes rules — it never speeds up a single rule that iterates a large document.

go deeper

for a junior

Know that one Rego package can define the same rule name many times and that OPA does not necessarily evaluate all of those bodies for every input.

for a middle

Explain what the indexer can read — an equality between a ground reference into input and a literal constant — and name shapes it cannot, such as negations, comprehensions and helper calls.

for a senior

Show how you would restructure an existing rule set so its discriminators are indexable, and prove the change with a profile rather than asserting that the style is faster.

for a principal

Own the house style: every rule in a large same-named set opens with a plain equality on the object kind, so the set stays selectable as teams keep adding to it.

## The problem indexing solves A Rego package typically defines one virtual document many times: dozens of rules all named `deny`, each matching a different object kind and each contributing a message. Evaluating the document means taking the union of every rule that produces a value. Done naively, every rule body runs for every input, and most of them go undefined on their very first expression because the object is the wrong kind. Rule indexing is OPA's answer. Before evaluation, the compiler inspects the bodies of the rules that define one virtual document and builds a lookup structure from the expressions it can interpret. At query time OPA reads the corresponding values out of the input and uses them to select the candidate rules; the rest are never evaluated. ## What the indexer can read The indexer understands a deliberately small set of expression forms. The workhorse is an equality between a **reference whose path is ground** — no variables in it, such as `input.request.kind.kind` or `input.request.object.metadata.namespace` — and a **constant** value written literally in the policy. Pattern matching with `glob.match` against a constant pattern is also indexable. Everything else is opaque to it. A rule whose body contains no expression the indexer can interpret simply has no key: it joins the set of rules that are always evaluated, which is the correct conservative behaviour, because skipping it could change the answer. ## What defeats it, in order of how often it happens - **Hiding the discriminator in a helper.** `is_namespace(input)` reads better than `input.request.kind.kind == "Namespace"`, but the indexer inspects expressions in the rule body, and a function call is opaque. Every rule that adopts the helper loses its key. - **Expressing the discriminator as a negation.** `not exempt_kind(input)` cannot serve as an index key; a negation tells the indexer nothing about what value would match. - **Comparing against a computed value.** `kind := lower(input.request.kind.kind); kind == "namespace"` puts a builtin call between the input and the constant. The comparison is now against a variable whose value the compiler does not know. - **Deriving the fact with a comprehension or `walk`.** Anything that has to run to produce the compared value is by definition not something the compiler can look up ahead of evaluation. - **Non-ground paths.** A reference that iterates, such as `input.spec.containers[_].image`, is not the ground reference the index is built on. ## How to write for it Open each rule of a large same-named set with a plain, literal equality on whatever discriminates it — object kind, apiVersion, operation — as a top-level conjunct of the body, before any helper does anything clever. Keep the constant a literal. Then let the helpers handle the parts that genuinely need logic. It costs nothing in readability, and it is the difference between a document whose rules are selected and one whose rules are all run. ## What indexing is not Indexing chooses **which rule bodies to evaluate**. It does nothing at all for a rule that has been selected and then iterates a large reference document inside its body — that is a data-shape problem, solved by how the document is keyed, not by how the rule opens. Nor is it caching: nothing is remembered between decisions; the index is a compile-time structure over the policy, not over the answers. ## Convincing yourself it is working Do not assert it — measure it. Profile the same query against the same input before and after restructuring: when the index is doing its job, the bodies of the rules whose key does not match the input stop appearing with evaluation counts at all, rather than appearing with a count of one and an immediately undefined first expression. That difference is visible directly in the profile output, which makes it a claim you can put in a review rather than a piece of folklore.

  • Why does wrapping the kind check in a helper function hurt indexing?
    The indexer inspects expressions in the rule body and understands equality against a constant. A function call is opaque to it, so the rule no longer advertises a key and joins the set that is evaluated for every input. The readability win costs you rule selection.
  • Does rule indexing help a rule that scans a large data document?
    No. Indexing decides which rule bodies run. Once a rule is selected, its body iterates exactly as it would have. Cost inside a body is a question of how the reference data is shaped and keyed, and it is fixed there, not in the rule head.
  • How would you tell whether indexing is actually doing anything?
    Profile the same query and input before and after the change. When the index works, the rules whose key does not match the input stop showing up with evaluation counts at all. That is an observation you can paste into a review, unlike an assertion that the style is faster.

saying these in an interview costs you the question

  • Thinks indexing speeds up iteration inside a rule body
  • Puts the kind check behind a helper and expects an index
  • Believes any expression in a body can serve as a key
  • Confuses rule indexing with remembering past decisions
  • Assumes a negation can act as the index key

context