skip to content

OPA allowed an unencrypted volume though a Rego rule forbids it — how do you check the rule was loaded?

level: seniorimportance: should knowfreq 42%

answer

  1. an empty result and a missing rule look alike
  2. ask the engine what it holds
  3. bundles activate all or nothing
  4. the queried path must match the package
  5. revision age is the whole incident

basics

~20 s

Separate evaluated-and-found-nothing from never-loaded. Ask the running engine which modules it holds and whether its bundle activated, then replay the exact input against that same bundle. A bundle that fails to compile is rejected whole, leaving older policy serving.

solid answer

~50 s

A gate that passes and a rule that is absent look identical downstream: both produce an empty set of violations. So the first question is not why the rule did not match, it is whether it was there. Ask the running engine which modules it holds — the policy API returns each one with its source — and check bundle activation: the health endpoint with the bundles parameter stays non-200 until every configured bundle activates, and the status plugin reports the active revision and the last activation error. Activation is atomic: if any module fails to compile the whole bundle is rejected and the previously activated policy keeps serving, so you can be a month stale behind a green gate. Quieter still is a file that was never in the bundle — a build glob or manifest that missed the path — where nothing errors anywhere. Then replay the offending input offline against that same bundle.

go deeper

for a junior

Know that a rule producing no violations and a rule that is not loaded look identical to whatever consumes the decision, so the absence of a block is not evidence the rule ran.

for a middle

Be able to name where to look: the policy API for loaded modules, the bundle-aware health endpoint for activation, decision logs for what was asked and answered.

for a senior

Demonstrate the diagnosis order — presence, then addressing, then semantics — and explain atomic bundle activation and why it leaves a green gate serving stale policy.

for a principal

Own the systemic answer: what the policy repository's CI must block on, what the engine must page on, and how you bound and communicate an enforcement gap without overstating it as a breach.

## Why this is hard to see A policy gate reads a decision. When the decision is an empty set of violation messages, the gate passes. It cannot tell the difference between: 1. the rule ran and this input genuinely satisfies it, 2. the rule ran but its body was not satisfied for a reason nobody intended, 3. the rule was never in the loaded policy set at all. All three produce the same silence. On call, the instinct is to debug the rule's logic — case 2 — and that instinct wastes the first hour if the truth is case 3. Establish presence before you debate semantics. ## Step one: what does the engine actually hold? Ask the running engine, not the repository. OPA's policy API (`GET /v1/policies`) returns every module the server has loaded, keyed by module id, with its raw source. Look for the package that should contain the rule and read the rule out of the response. If it is not there, you are done diagnosing the class of failure and can start on the cause. If policy is delivered as a bundle, also check activation. `GET /health?bundles=true` returns non-200 until every configured bundle has activated at least once — a useful readiness signal that many deployments wire up and then never look at again. The status plugin reports the revision of the bundle currently active and the error from the last failed activation, either to your status service or as metrics you can graph. A revision that has not moved in weeks against a repository with weekly merges is the whole incident in one number. ## The mechanic people get wrong Bundle activation is **atomic**. If any module in a bundle fails to compile, the bundle is rejected as a unit and the engine keeps serving the last bundle that activated successfully. There is no partial load, no skipping the broken file and taking the rest. That has an uncomfortable consequence. The loud version of this failure — a rule rewritten in a syntax the engine no longer parses, say a v0-era encryption-at-rest rule after a move to a v1 engine — does not take the gate down. It leaves the engine happily enforcing month-old policy while every new rule merged since sits unloaded. Downstream, everything is green. ## The genuinely silent version Worse is the file that was never handed to the compiler. The rule is in the repository, reviewed and merged, but the bundle build globbed a directory that did not include it, or the manifest's roots excluded its path, or the publishing job ran from a branch that does not have it. Nothing fails to compile because nothing tried. No error exists anywhere in the system. The only evidence is behavioural: a control that used to block things blocks nothing. A close cousin is a package rename. The gate queries a fixed path, and if the module now declares a different package the rule is loaded, compiles, and lives somewhere the query never looks. The policy API will show you the module and its package, which is why reading the source out of the running engine is worth the extra minute over trusting the repository. ## Step two: prove it evaluated Presence is not evaluation. Decision logs record, per decision, the input and the result returned, so a log entry for the deploy in question tells you what the engine was actually asked and what it answered — including whether it was asked at all, which is its own class of incident when the gate is misconfigured. Then reproduce offline. Take the exact input from the decision log and evaluate it against the same bundle artifact with `opa eval -b <bundle> -i input.json 'data.volumes.deny'`. If the current bundle denies it, the bundle in production is not the one you are holding. If the current bundle also allows it, the rule's logic is the problem after all and you have moved cleanly from case 3 to case 2. ## Step three: close it so it cannot recur silently The durable fixes live in the policy repository's own CI and in the engine's alerting, and they map one to one onto the causes: - A file that will not compile: `opa check --strict` over the whole tree as a blocking step, plus `opa fmt --diff --fail` and a Rego linter so dead imports and shadowed names are caught at the same moment. - A file that was never included: after building the bundle, assert that the packages and rule names your gates query are present in the artifact. This is the check nobody writes and the one that catches the silent case. - A bundle that stopped activating: alert on activation failures and on bundle revision age. Both are already exported; they just need someone to page on them. ## What you say on the call Be precise about scope and honest about the direction of the claim. The control was not enforced between the last successful activation and now; that is a window, not a breach. Replay the decision inputs recorded over that window to find which changes would have been denied, and treat those as the remediation list. Distinguish clearly between what you know the gate allowed and what you merely suspect: the decision log is evidence, and the absence of a decision is also evidence.

  • The engine reports the module loaded and the bundle current — where do you go next?
    Then it is semantics or addressing, not presence. Compare the path the gate queries with the package the module actually declares, since a rename leaves the rule loaded but unreachable. If those match, pull the exact input from the decision log and evaluate the rule against it offline; a body that is not satisfied yields no result at all, so an undefined condition on a field the input does not carry looks the same as a compliant resource.
  • Why is a module that fails to compile quieter than you would expect?
    Because bundle activation is atomic. The bundle is rejected as a unit and the engine keeps serving the last bundle that activated, so the gate stays up and stays green against stale policy. Nothing downstream sees an error. The signal exists — activation failure and a frozen bundle revision — but it lives in the engine's status and metrics, not at the gate, so it is only useful if somebody alerts on it.
  • Which single CI check would have caught this before it shipped?
    It depends which cause you found. For a file that no longer compiles, `opa check --strict` over the tree as a blocking step. For a file that was never included in the bundle, no compile check helps: you need a post-build assertion that the packages and rule names the gates query are present in the artifact. The second one is rarer to see and catches the failure mode that produces no error at all.

A checkpoint that waves everyone through looks exactly like one where nobody is standing. You check whether the guard is on the roster before you ask why they did not stop anyone.

saying these in an interview costs you the question

  • Assumes a broken module is skipped and the rest still loads
  • Reads a green gate as proof the rule ran
  • Debugs the rule's logic before confirming it was loaded
  • Trusts the repository over the running engine's loaded modules
  • Calls an enforcement gap a breach without evidence

context