What does opa test --coverage mark as covered in a Rego policy, and what does it not prove?
answer
- which expressions the tests evaluated
- covered is not asserted
- text coverage, not input coverage
- one deny test can reach 100%
basics
~20 sopa test --coverage marks every expression the tests actually evaluated, reported per file with a percentage. That proves the policy text ran; it does not prove the rule decided correctly, and it says nothing about which inputs were tried.
solid answer
~40 sAdding `--coverage` to `opa test` makes OPA record every expression it evaluates while the tests run, and report per file the covered and not-covered source ranges plus a percentage. Covered means reached at least once — not that the expression was true, and not that any test asserted on the result. Because a Rego body stops at its first undefined expression and then yields no value at all, an uncovered tail on a `deny` rule means no test ever drove that rule to its conclusion; that is the useful half of the report. Conversely 100% only tells you some fixture reached the end of each body once. Coverage is over the policy text, never over the input space, so it cannot tell you the rule denies the other nine shapes it should.
go deeper
Be ready to say what --coverage adds to opa test and to read the output out loud: covered and not-covered source ranges per file, plus a percentage.
Explain that coverage records an expression being evaluated rather than being true, and that a body stops at its first undefined expression, so an uncovered tail means no test reached that rule's conclusion.
Show what you ask for beyond the number: a case per deny message, a case proving the rule stays quiet on a legitimate object, and fixtures loaded with realistically shaped reference data.
Own the argument that a coverage percentage is weak evidence for a guardrail, and be able to state what artefact you would require in its place before a rule is rolled out.
## What the flag actually records Running `opa test --coverage` turns on an evaluation tracer: while the test run executes, OPA records the source location of every expression the evaluator reaches. The report (readable as JSON with `--format=json`) is organised per file — a list of covered source ranges, a list of not-covered ranges, counts of each, and a coverage percentage for the file — plus an overall percentage across the files that were loaded. The unit is a source location, not a "line" in the sense a Go or Java coverage tool means. A Rego rule body is a conjunction of expressions, and each expression is tracked on its own. That granularity matters, because in Rego the interesting failures live inside a body rather than between statements. ## What "covered" means, precisely Covered means one thing only: during this run, the evaluator reached that expression at least once. It does **not** mean the expression evaluated to true. It does **not** mean a test asserted anything about the result of the rule containing it — the coverage tracer has no idea which assertions exist or which of them passed. And it does not mean the enclosing rule produced a decision. ## The undefined interaction, where Rego differs from most languages A rule body is a conjunction, and evaluation stops at the first expression that is **undefined**. Undefined is not `false`; it is the absence of a value, and a body that goes undefined produces no result at all rather than a negative one. So coverage within a body is prefix-shaped: if a fixture makes the first comparison undefined, nothing below it is covered. Read positively, that is the report's most actionable signal. An uncovered tail on a `deny` rule means no test in the suite ever drove that rule all the way to its message — nothing you have written has ever seen the rule deny. A percentage averages that away; the `not_covered` ranges name it. ## What 100% does not buy you Take a rule requiring that every namespace has an entry in a quota inventory loaded as `data`. Suppose every expression in it is covered. All that establishes is that at least one fixture reached the end of the body: one namespace object, one inventory, one message. Coverage is over the policy **text**. What decides whether the guardrail works is the **input space**: - the namespace whose name matches a record with different capitalisation; - the object where the metadata your rule reads is simply absent; - the inventory document that was never loaded at all — in which case the reference into `data` is undefined, the rule quietly produces nothing, and the outcome is indistinguishable from "allowed". None of those appear as coverage gaps, because the same expressions are the ones being evaluated. A rule that denies everything can report exactly the same coverage as a correct one. So can a rule whose message is wrong, or a rule that is never reached in the environment where it is deployed. ## Reading it as a lead, before a rollout The honest framing at sign-off: coverage is a floor detector. It finds untested policy cheaply and reliably, and that is worth having. It is not evidence that the rule is right. So do not sign off on the number. Ask for: - at least one test per `deny` message that makes the message appear, and one that proves the rule stays quiet on a legitimate object — a rule that only ever fires is as broken as one that never does; - fixtures loaded with a document shaped like the real reference data, not a two-record stub, so the `data` lookups in the body are exercised against a realistic shape; - a walk of the `not_covered` list looking for whole rules and helper functions that no test touches — those are either dead policy or the branch that will surprise you; - an explicit case for the "reference data missing" scenario, because that is the failure that turns a deny rule into a no-op. ## Using the number honestly Treat rising coverage as "we have stopped shipping untested rules" and nothing more. The claim "the guardrail works" is carried by which decisions the tests pin down, not by how much of the file was executed while pinning them.
- A deny rule shows 100% covered but fired in only one test — would you sign the rollout off on that?No. All 100% establishes is that one fixture reached the end of the body once. Coverage is over the policy text, not the input space. I would ask for a case per message, a case proving the rule stays quiet on a legitimate object, and a case for the reference data being absent, which makes the rule silently produce nothing.
- If a rule body goes undefined halfway through, what does the coverage report show?The expressions up to and including the one that went undefined are covered; everything after it is not. A Rego body stops at its first undefined expression and yields no result, so an uncovered tail is a reliable sign that no test drove the rule to its conclusion.
- How would you use the not_covered section rather than the percentage?It names the rules and helper functions no test ever reached — dead policy, unreachable branches, or a rule someone forgot to test. That is a concrete work list. The percentage blends all of it into one number that goes up when you test the easy rules.
Coverage is a receipt showing which lines of the recipe you read out loud, not proof the cake tastes right.
saying these in an interview costs you the question
- Claims 100% coverage means the policy is correct
- Treats a covered expression as one that returned true
- Thinks coverage measures which inputs were tested
- Reads only the percentage, never the not-covered ranges
- Assumes undefined counts as false, so the deny fired