Your Rego secrets rule has denied nothing in six weeks of evaluation — how do you find out why?
answer
- silence has two explanations
- evaluate a known-bad fixture first
- then replay real captured objects
- check the shapes production actually sends
- make an unreadable document a denial
basics
~20 sSilence is not evidence: a clean estate and a broken rule produce identical output. Replay a violating fixture and a sample of real objects offline — if the bad fixture also passes, the rule's body is undefined.
solid answer
~50 sRefuse both easy conclusions first — “it works” and “it is broken” produce the same empty result. Evaluate the rule against a manifest you know violates it, with a credential pasted into `env[].value`. If that fixture is not denied, the rule cannot deny anything and its body is undefined for the shape you gave it. If the fixture *is* denied, the rule works for that shape but maybe not for what production sends, so sample the real workload kinds and replay them: bare Pods and CronJobs put containers where the rule never looks. Check whether a built-in in the body could be erroring, since OPA treats a failing built-in as undefined by default. Only when a violating real object is denied offline can you call the estate clean — and then say so with the replay as evidence, not with the absence of alerts.
go deeper
Know the core fact you would need here: a Rego rule that cannot read the field it needs produces no denial and no error, so zero denials never proves the objects were compliant.
Be able to run the diagnosis: evaluate a known-bad manifest, then narrow the body expression by expression to find the one with no value, and name the workload shapes that put containers at different paths.
Show a method, not a hunch — falsify with a counterexample, replay real captured objects, rule out silently degraded built-ins, and then change the design so an unreadable document denies instead of passing.
Own the standard behind it: a control whose failure mode is silence is not evidence of anything. Decide what your platform requires before a rule counts as enforced, and what you tell the people who trusted six weeks of green.
## The trap in the question Six weeks of zero denials has two indistinguishable explanations: nobody violated the rule, or the rule was never able to fire. A Rego rule whose body is undefined contributes nothing and reports nothing — no error, no log line, no metric — so the observable output of a perfectly healthy rule over a clean estate and the output of a rule reading a path that does not exist are byte-for-byte identical. Any answer that reads silence as success has failed the question. ## Diagnose in order of cheapness **1. Can the rule deny at all?** Take a manifest that unambiguously violates it — a Deployment with `env: [{name: DB_PASSWORD, value: "..."}]` and no `valueFrom` — and evaluate the rule against it offline. This costs a minute and splits the problem in half. If the known-bad fixture is not denied, stop looking at the platform; the rule is broken. **2. Which expression died?** If the fixture passes, narrow the body: evaluate the first reference on its own, then each subsequent expression, until you find the one with no value. In practice it is nearly always the first one — the path into the document. **3. Does the rule match what production actually sends?** If the fixture *is* denied, the rule works for the shape you fixtured and possibly nothing else. Pull a real sample of the objects that have been going through — the actual workload kinds, from the manifests in the repositories that deploy or from the objects live in the cluster — and replay the rule over them. A pod template lives at `spec.template.spec` for a Deployment, at `spec.containers` for a bare Pod, and at `spec.jobTemplate.spec.template.spec` for a CronJob; a rule written for one is undefined for the others. Also check `initContainers`, which a rule that only walks `containers` never sees. **4. Could a built-in be failing?** By default OPA treats a built-in function error as undefined rather than as a hard error, so a `sprintf`, `split` or `to_number` that chokes on unexpected data silently drops the rest of the body. Re-running the evaluation with strict built-in errors turns that class of failure into a visible message. **5. Only now consider that the estate is clean.** If a violating real object is denied on replay and the sample contains no violations, the zero is honest. Say that with the replay as the evidence. ## What you change afterwards The incident is not "one rule had a bad path". It is "our gate's failure mode is silence", and that is a design property you can fix. **Every rule ships with a counterexample.** A fixture that must be denied, asserted in CI on every policy change, with one fixture per document shape the rule claims to cover. A test suite made only of compliant fixtures is passed by a rule that denies nothing. **Make an unreadable shape a denial.** Add a companion rule that denies any document whose container path cannot be resolved from a known set. This is the only mechanism that reports a shape nobody anticipated, and it converts the silent class of failure into the loud one. **Give complete rules a `default`.** A decision that always has a value is easier to reason about than one that may be absent. It does not fix an undefined body — nothing about a default makes a rule match — but it removes one source of ambiguity in what the caller receives. **Baseline the estate.** Run the rule in a reporting pass over everything currently deployed. That gives you a number to compare against, and a number that drops to zero overnight is a signal; a number that was always zero never was. **Watch for a shape change, not just a rule change.** The rule that worked when it shipped stops working when a team adopts a workload kind it does not cover — and nothing in the change that broke it touched the policy repository at all. ## How to talk about it The answer an interviewer is listening for has three beats: name the ambiguity (empty and broken look the same), give a concrete falsification step (the must-deny fixture, then a replay of real objects), and then make the failure mode structurally visible rather than promising to be more careful. Blaming the author of the path, or proposing more review, misses that the whole class of bug is invisible by construction.
- The known-bad fixture is denied, but production still denies nothing. What next?The rule works for the shape you fixtured, so go and get the shapes production sends. Sample the real manifests or live objects, note their kinds, and replay the rule over them. Deployments, bare Pods and CronJobs place the pod template at three different paths, and `initContainers` is a fourth blind spot. If real violating objects are denied on replay, the zero may be honest — report it with the replay as evidence.
- Could a failing built-in function explain six weeks of silence?Yes. OPA's default is to treat a built-in error as undefined rather than as a hard error, so one expression choking on unexpected data drops the remainder of the body exactly as a missing key does. Re-running with strict built-in errors surfaces it. It is worth enabling in CI evaluation even if you are more cautious about it in the enforcement path.
- What single change would have caught this on day one?A fixture the rule must deny, asserted in CI, with one per document shape the rule claims to cover. It is the only cheap test that fails when the rule loses the ability to fire — compliant fixtures all pass against a rule that does nothing, which is why a suite made only of those is worthless for this class of bug.
saying these in an interview costs you the question
- Reports the gate as healthy because no denials appeared
- Assumes zero denials means the estate is clean
- Verifies only that the policy compiles and is loaded
- Never evaluates the rule against a known-bad object
- Expects an error log for the missing field path
- Proposes more code review instead of a must-deny fixture