A Rego rule checking namespaces against thousands of quota records is slow — how do you find and fix it?
answer
- profile before guessing
- high redo count means backtracking
- the scan hides in a helper
- arrays scan, objects look up
- reshape the document before loading
basics
~20 sProfile with opa eval --profile: the top row is the comparison inside the scan, its redo count near the record count. Fix the data shape — key the inventory by namespace name so the lookup is one reference.
solid answer
~40 sMeasure before guessing. `opa eval --profile` with the real policy, the real inventory and a representative input prints a table of expressions with total time, evaluation count and redo count; sorting by redos puts the culprit on top. A redo count in the thousands is the signature of backtracking through a collection — here, a helper comparing `data.quotas.namespaces[_].name` against the namespace, once per record, on every decision. The fix is not the rule text but the document: build the inventory as an object keyed by namespace name before it is loaded, so the helper becomes a single key lookup with the name already bound. Building that map with a comprehension inside the rule just moves the scan. And profile against realistically sized data — a five-record fixture makes the rule look fine.
code
rego · 12 linespackage quota
deny contains msg if {
input.request.kind.kind == "Namespace"
name := input.request.object.metadata.name
not quota_exists(name)
msg := sprintf("namespace %v has no ResourceQuota entry", [name])
}
quota_exists(name) if {
data.quotas.namespaces[_].name == name # walks every record
}go deeper
Know that opa eval has a --profile flag and that a policy's cost usually comes from the data it walks, not from how many lines the rule has.
Explain the profile columns — total time, evaluations, redos — and why a redo count close to the number of records means the body is scanning a collection.
Diagnose from the profile, then fix the document shape rather than the rule text, and show that you profiled the query the gate runs with a realistically sized inventory.
Own the standard that reference data enters the engine in the shape the policies query it in, making that a property of how the document is built rather than a fix rediscovered per rule.
## The setup A rule requires every namespace to appear in a quota inventory: the inventory is a document loaded into `data`, holding a record per namespace, and the rule denies a namespace object that has no matching record. It passed review, it passed its tests, and once the real inventory grew to thousands of records the evaluation got noticeably expensive. ## Step one: profile, do not guess Run the query the gate actually asks for, with the real data loaded and a representative input: ``` opa eval --profile --format=pretty \ -d policy.rego -d quotas.json -i review.json 'data.quota.deny' ``` The profiler prints one row per expression with the columns that matter: total time, number of evaluations, number of redos, and the source location. `--profile-sort` lets you order by something other than time; sorting by redo count is usually the fastest route to the answer. Two details decide whether the exercise is worth anything: - **Query what the gate queries.** Asking for `data` or a whole package evaluates everything in it, and you end up profiling rules the gate never asks about. - **Load what production loads.** Cost here is a function of the document, not of the policy text. A fixture with five records reproduces nothing. ## Step two: read the columns Evaluations count how many times the evaluator entered an expression. **Redos** count re-entry on backtracking, which is what iteration looks like from the inside: for each element of a collection, the body is retried. So an expression whose redo count sits near the number of records in your inventory is not merely slow, it is scanning, and you now know exactly which expression and how many times. In this rule, the top row is the comparison inside the membership helper. The outer `not` around that helper does not help either: negation still has to establish that the helper is undefined, which means running the scan to exhaustion for every namespace that *is* compliant — the common case. ## Step three: fix the shape, not the rule The rule text is fine. What is wrong is that the inventory arrives as an **array of records** while the policy asks a **membership question by name**. An array can only be searched by walking it. An object can be looked up by key: when the key is a variable that is already bound — the namespace name pulled from the object under review — a reference like `data.quotas.by_namespace[name]` is a direct lookup rather than an iteration. So transform the document where it is built, before it is loaded: ``` {"by_namespace": {"team-a": {...}, "team-b": {...}}} ``` and the helper collapses to a single reference into it. The rule reads better as a bonus; the point is that the lookup is no longer proportional to the inventory. ## Two traps on the way **Rebuilding the map inside the policy.** Writing a comprehension that turns the array into an object at the top of the rule feels like a fix, but it constructs thousands of entries as part of evaluating the decision. The cost returns wearing a different hat. Reshape once, outside, in whatever produces the document. **Shipping more than the question needs.** If the rule only asks *whether* a namespace has an entry, the loaded document does not need the full record per namespace — a set of names gives the same key lookup and a much smaller document to build, ship and hold. Decide what the policies actually query before deciding what the inventory contains. ## Step four: prove it, and keep proving it Re-profile the same query with the same data. The row that dominated should be gone, not merely smaller, and the redo count should no longer track the record count. Keep the realistic inventory around as a fixture for the next person: the cost of this rule is invisible to tests and to coverage, both of which pass happily on five records, so a profile against realistic data is the only artefact that will catch the regression when someone reintroduces a scan. ## Why this generalises Almost every expensive Rego policy is one of two things: a rule set where every rule is evaluated because nothing is selectable, or a rule whose body walks a document that could have been keyed. The profiler's redo column separates the two in about a minute, and the second one is nearly always fixed in the data pipeline rather than in the policy.
- Which column of the opa eval --profile table points at iteration?The redo count. Total time tells you where the milliseconds went; redos tell you why. Re-entry on backtracking is how iteration appears from the evaluator's side, so a redo count near the size of the collection identifies the scan and the expression doing it.
- Why not build the keyed object with a comprehension inside the rule?Because it constructs thousands of entries as part of evaluating the decision, so the scan cost comes back in another form. The transformation should happen once, in whatever produces the document, so the engine loads a shape the policies can query directly.
- You profiled with a five-record fixture and it looked fine — what went wrong?Cost here is a function of the loaded document, not of the policy text, so a tiny fixture cannot reproduce it. Profile with an inventory the size of the real one and an input representative of what the gate sees, and keep that fixture around for the next change.
Scanning an array of records for one name is reading every page of a phone book; an object keyed by name is the index at the back.
saying these in an interview costs you the question
- Guesses at the slow rule instead of profiling it
- Reads total time only and ignores the redo count
- Rebuilds the lookup object inside the rule each evaluation
- Blames the engine rather than the document shape
- Profiles with a tiny fixture and declares the rule fast