A Rego advisory rule warns 'TLS policy is not approved' with no resource name — how do you find which listeners matched?
answer
- Reproduce on the exact document first
- Make the rule report what it read
- Notes for values, fails for dead ends
- Shrink the template until it is readable
- The message itself is the real defect
basics
~20 sReproduce it locally with opa eval on the same template, add a print of the resource id and the value the rule reads, and read it back with --explain=notes. Then fix the message to carry that context.
solid answer
~50 sGet the exact document the pipeline evaluated and run the rule locally: `opa eval -i template.json -d policy.rego --format=pretty 'data.cfn.tls.warn'`. Add a `print` just above the suspect expression, dumping the logical id and the field the rule reads, and read it with `--explain=notes` so you see one line per resource instead of the whole trace. That usually settles it immediately — either the field path is not where you assumed, or the value has a shape the comparison never matches. For resources you expected to warn on but did not, switch to `--explain=fails` and cut the template down to a single listener so the trace is short enough to read. The permanent fix is not the diagnosis, it is the message: build the resource id and the observed value into it with `sprintf`, and keep the narrowed template as a test fixture so the next person is never sent on this errand.
code
rego · 12 linespackage cfn.tls
approved_policies := {"ELBSecurityPolicy-TLS13-1-2-2021-06"}
warn contains msg if {
some name, res in input.Resources
res.Type == "AWS::ElasticLoadBalancingV2::Listener"
res.Properties.Protocol == "HTTPS"
print("listener:", name, "policy:", res.Properties.SslPolicy)
not approved_policies[res.Properties.SslPolicy]
msg := "TLS policy is not approved"
}go deeper
Know that you reproduce a policy result by running the rule against the same input file locally, and that adding a print inside the body is how you see the values it is reading.
Explain the two tracing modes you would use and why: notes to read the values the body saw, fails to find the expression that stopped it, with the input narrowed first so either output is short enough to read.
Show the whole loop, including the part after the diagnosis: the message gets the resource id and the observed value built into it, and the narrowed document becomes a regression fixture. Be able to say why a rule warning on compliant resources is a serious problem, not a cosmetic one.
Own the standard rather than the incident. Decide what every advisory finding must carry before it may be turned on for other teams, and treat a stream of context-free warnings as a defect in the policy programme, because it teaches an entire organisation to ignore the channel.
## The situation An advisory rule is running in warn mode over a CloudFormation template. It emits a single unadorned string — `TLS policy is not approved` — once per matching resource. The template has a dozen listeners. Nobody can act on the output: it does not say which listener, and it does not say what the rule saw. Somebody has to reproduce the result and explain which expressions did and did not evaluate. ## Step 1: get the same input A reproduction that uses a different document proves nothing. Take the template the job evaluated, as it was evaluated — the same file, the same rendering step ahead of it if there was one. Then run the rule directly: ``` opa eval -i template.json -d policy.rego --format=pretty 'data.cfn.tls.warn' ``` Query the specific rule, not the whole package: a package-wide query drags every other rule's evaluation into anything you trace afterwards. If the local run does not reproduce the warning, stop and fix that first. The likely causes are a different input shape (a template rendered by another step, or JSON versus YAML) or different data loaded alongside it. ## Step 2: make the rule say what it sees The rule's message tells you nothing, so make the body tell you instead. Put a `print` immediately above the expression you suspect, dumping both the identifier and the value being read: ``` print("listener:", name, "policy:", res.Properties.SslPolicy) ``` Read it with `--explain=notes`, which filters the trace down to exactly those messages. You now have one line per resource the body reached, in evaluation order. This is where most of these end: the printed value is `<undefined>` because the field lives somewhere other than where the author assumed, or it is a string whose form the comparison never matches — a named security policy such as `ELBSecurityPolicy-TLS13-1-2-2021-06` rather than a bare version number. Note what the `<undefined>` case means for an advisory rule written with a negation. If the body says *not approved* and the value it checks is missing, the inner check is undefined, the negation succeeds, and every listener warns — including the ones that are perfectly configured. An advisory rule that fires on everything is worse than one that fires on nothing, because it trains the team to skim past it. ## Step 3: narrow until the behaviour is isolated For the resources you expected to match but that produced nothing, switch to `--explain=fails`, which shows the expressions that did not hold. On the whole template that output is unreadable, and not because anything is wrong: a body that iterates over resources fails on every resource it does not match, which is how filtering works. So cut the template down — one listener, the properties the rule touches, nothing else — and re-run. The trace collapses to a handful of lines and the answer is usually visible at a glance: either the expression you suspected failed, or evaluation never reached it, which relocates the problem earlier in the body. Work by bisection if the minimal case still misbehaves: remove properties until the behaviour flips, and the last thing you removed is implicated. ## Step 4: fix the thing that sent you here The missing field is the bug of the hour. The message is the bug of the year. An advisory finding that names no resource and quotes no value costs somebody this entire exercise every time it fires. Rewrite it to carry its own evidence: ``` msg := sprintf("listener %v uses TLS policy %v, which is not approved", [name, res.Properties.SslPolicy]) ``` Now the warning is self-contained: the developer who receives it can see which resource and what value, and can decide without opening a policy repository they have never read. ## Step 5: keep the fixture The narrowed template is the most valuable artefact of the session. Save it as a test case — the input, plus the expectation that the rule now produces the message with the right resource named — so the fix is pinned and the next change to the rule cannot silently reintroduce the same blind spot. ## A note on the hosted playground The browser playground is excellent for sharing a reduced reproduction with a colleague or with someone on another team who does not have OPA installed: a link, a rule, a fixture, evaluated in front of both of you. It is not the place to paste a real infrastructure template. By the time you have narrowed the input to a handful of lines you can usually anonymise it trivially — replace identifiers, keep the shape — and share that instead.
- The print shows <undefined> for the field the rule checks, and every listener warns. Why does that follow?Because the check sits under a negation. The inner membership test on a missing value is undefined rather than false, so it never holds, so the negation always succeeds and the body reaches the message for every listener — the compliant ones included. The printed `<undefined>` is the evidence; the fix is to read the field where it actually lives and to make the rule refuse to run against a shape it does not recognise.
- Your local run does not reproduce the warning at all. Where do you look?At the input, before the rule. The job may evaluate a rendered or converted artefact rather than the file in the repo, it may load data alongside the input that you did not, or it may be running a different revision of the policy bundle. Reproducing the exact document and the exact bundle is a precondition for any of the tracing work; without it you are debugging a different question.
- What do you hand back to the team that received the unactionable warning?A message they can act on, not a debugging story. The rewritten finding names the resource and quotes the value observed, so the next occurrence needs no interpreter. Alongside it goes the narrowed fixture as a test, and — if the rule was warning on compliant resources — a note that the previous run's findings were not trustworthy, since people will otherwise assume the old output was real.
saying these in an interview costs you the question
- Debugs against a different template than the job used
- Reads a full trace before narrowing the input
- Fixes the field path and leaves the message empty
- Treats an advisory rule firing on everything as harmless
- Pastes a real infrastructure template into a hosted playground
- Escalates warn to block instead of adding context