In opa eval, how do --explain=notes and --explain=fails differ?
answer
- Both filter one underlying event stream
- One shows your messages, one shows dead ends
- Full is correct and unreadable
- A Fail event is usually normal iteration
- Narrow the input before reading either
basics
~20 sBoth filter the same evaluation trace. notes shows only what the policy itself said — print and trace output. fails shows only the expressions that did not hold, which is how you find where an undefined rule stopped.
solid answer
~40 s`opa eval` can emit a trace of evaluation, and `--explain` chooses which events you get. The unfiltered stream (`full`) records entering and leaving every rule, evaluating every expression and re-trying every iteration, which on a real document is thousands of lines. `--explain=notes` narrows it to note events — the messages the policy emitted itself via `print` and `trace` — so you read your own commentary in evaluation order. `--explain=fails` narrows it to failure events: each expression that evaluated to undefined or false, with its location. That second one answers the question you usually have, namely "my rule returned nothing, which line stopped it". Pair it with `--format=pretty` so the trace is readable, and narrow the input first, because normal iteration produces failures constantly and the interesting one drowns.
go deeper
Know that opa eval can show why a rule produced nothing, and that --explain=fails is the flag for it while --explain=notes shows messages the policy printed itself. Pair either with --format=pretty.
Explain that both flags filter one underlying event stream of Enter, Eval, Exit, Fail, Redo and Note events, and say why the unfiltered full mode is unusable on anything but a tiny fixture.
Demonstrate the workflow rather than the flag list: narrow the input first, use fails to find the expression that stopped the body, add a print and read it with notes, then keep the minimised document as a test fixture.
Frame explanations as author-time diagnostics with no operational reach, and be clear about what the organisation relies on instead for after-the-fact answers about production decisions.
## The trace behind both flags When OPA evaluates a query it can record events as it goes. The event kinds are the vocabulary of every explanation you will read: | Event | Meaning | | --- | --- | | Enter | evaluation entered a rule or a comprehension body | | Eval | an expression is being evaluated | | Exit | a body succeeded and produced a value | | Fail | an expression did not hold — it was undefined or false | | Redo | evaluation backtracked to try the next binding | | Note | the policy emitted a message itself | `--explain` does not change evaluation. It selects which of those events are printed alongside the result. ## `--explain=full` Everything. This is correct and almost never usable: every expression in every rule for every binding of every iteration variable. Evaluate a partial set rule over a fifty-resource template and you get thousands of lines in which the two you needed are indistinguishable. Full has one honest use — a tiny hand-made fixture where you genuinely want to watch the whole resolution unfold. ## `--explain=notes` Only note events: the messages the policy produced on purpose. This is the explain mode that pairs with `print` and `trace` in the body. Instead of a firehose you get your own annotations, in the order the evaluator reached them, interleaved with nothing else. It is the right mode when your question is *what values did the rule actually see*, and it scales to a real document because you control how many notes exist — one print in the body of interest produces one line per iteration of that body. ## `--explain=fails` Only failure events: expressions that evaluated to undefined or to false, each with a source location. This is the mode for the single most common Rego problem, which is silence. A rule that returns nothing gives you no error, no message and no clue; `fails` names the first expression that did not hold, which is nearly always either a field path that does not exist in the input or a comparison against a value whose shape you guessed wrong. ## The trap in reading `fails` A Fail event is not an error. It is the ordinary outcome of the ordinary Rego pattern: a body iterates over a collection and most members do not match, so most iterations fail — that is how filtering works. On a real input document `--explain=fails` therefore prints a wall of perfectly healthy failures. The discipline that makes it usable is narrowing. Cut the input down until it contains one instance of the thing you care about — one listener, one container, one plan change — and re-run. Now the trace is short enough that the failure you are hunting is either obviously present or obviously absent, and "absent" is itself the answer: the body never got as far as the expression you suspected, so the problem is earlier. Narrowing also leaves you with something worth keeping — the minimised document is exactly the fixture your regression test should use. ## Choosing between them A workable rule of thumb: - *The rule produced nothing and I do not know why* → `fails`, on a narrowed input. - *The rule ran but the values it matched were not the ones I expected* → `print` in the body plus `notes`. - *I am teaching someone how resolution works on a five-line example* → `full`. They compose in sequence rather than competing. A typical session is: `fails` on the narrowed document to find the expression that stopped the body, then a `print` immediately above that expression with `notes` to see the value it was actually looking at, then the fix, then the narrowed document saved as a test fixture. ## Practicalities Use `--format=pretty` so the trace renders as indented text rather than being embedded in JSON; the JSON form is for tooling that wants to post-process an explanation. Remember that an explanation covers the query you asked for — if you evaluate the whole package you trace everything in it, so query the specific rule you are debugging. And note that an explanation is a local diagnostic: it is produced by the tool you ran, on the input you supplied, and is not something a deployed server hands back with a decision.
- You run --explain=fails on a real template and get hundreds of lines. What do you do?Narrow the input, not the flag. Most Fail events are ordinary iteration: a body walks a collection and every non-matching member fails, which is how filtering works in Rego. Cut the document down to a single resource of the kind you care about and re-run; the trace becomes short enough to read, and the minimised document doubles as the fixture for a regression test.
- When is --explain=full actually the right choice?On a tiny hand-written fixture where you want to watch the whole resolution — entering the rule, each expression, each backtrack — usually to understand or teach how a body resolves rather than to fix a specific bug. On anything resembling a real input document it produces thousands of events and buries the answer.
- Does an explanation tell you anything a decision log would not?Yes, a different thing entirely. An explanation is the inside of one evaluation — which expressions held and which did not — produced locally by the tool you ran. It is a debugging artefact for the policy author, not a record of what the system decided in production, and you would not build operational reporting on it.
saying these in an interview costs you the question
- Treats every Fail event as an error
- Reaches for --explain=full on a real document
- Thinks notes shows failed expressions too
- Expects an explanation from a deployed server's response
- Reads a trace before narrowing the input