skip to content

Why does a rule denying deprecated TLS policies quietly pass a listener whose value is unknown until apply?

level: middleimportance: must knowfreq 52%

answer

  1. nothing to match against
  2. zero findings looks like a clean run
  3. denylist under-reports, allowlist mislabels
  4. the unknown needs its own branch
  5. count what the gate could not evaluate

basics

~20 s

A value that does not exist matches nothing. The rule hunts for a forbidden TLS policy, finds no value at all under that attribute, and reports no violation — a pass that looks identical in CI to a genuine one.

solid answer

~50 s

The rule is phrased as a denylist: violate when the listener's TLS policy is one of the deprecated ones. When the plan marks that attribute unknown until apply, there is no value under it, so it matches none of the forbidden entries and the rule produces nothing. Zero violations is exactly what a compliant listener produces, so the pipeline shows green and nobody learns that the attribute was never examined. Flipping to an allowlist — violate unless the value is one of the approved ones — changes the direction of the mistake, not the mistake: now every unknown is reported as a bad TLS policy, which fails closed but with a false and confusing message. The real fix is to branch on unknown explicitly: read the plan's unknown marker for that attribute and emit a deliberate outcome, so the gate's output separates "checked and clean" from "could not check".

code

json · 19 lines
json
{
  "resource_changes": [
    {
      "address": "aws_lb_listener.public",
      "type": "aws_lb_listener",
      "change": {
        "after": {
          "port": 443,
          "protocol": "HTTPS",
          "ssl_policy": null
        },
        "after_unknown": {
          "arn": true,
          "ssl_policy": true
        }
      }
    }
  ]
}

go deeper

for a junior

Know that a rule searching for a forbidden value finds nothing when the value does not exist yet, so the change passes and looks clean.

for a middle

Explain both rule shapes — deny-on-forbidden and deny-unless-approved — and predict which way each one fails on an unknown attribute, including why the second gives a misleading message.

for a senior

Demonstrate that you keep plan fixtures containing unknowns and that your gate reports how many resources it could not evaluate rather than folding them into the clean count.

for a principal

Set the expectation across the estate that a gate result which cannot separate clean from unevaluated is a number nobody should act on, and require rules to report both.

## The failure in one sentence Absence of a match is being read as evidence of compliance. ## The setup You own a rule that says: a public load-balancer listener must not use a deprecated TLS policy. The natural way to write it is a denylist — collect the listener resources from the plan, look at the attribute naming the TLS policy, and raise a violation if it is one of the entries in a set of retired policies. That rule works perfectly on a plan where somebody wrote the policy name literally in the configuration. It also works on a plan where they wrote a bad one, and it catches them. Now the team refactors: the policy name comes from a module variable fed by a lookup that only resolves during apply. The plan now carries no value under that attribute and flags the path as unknown. Your rule collects the listener, looks up the attribute, finds nothing to compare, matches none of the forbidden entries, and emits zero violations. ## Why this is worse than a crash A rule that errored would be noticed within an hour. This one produces the exact output a compliant plan produces: a green check, no findings, merge allowed. From the outside — the pull request, the pipeline summary, the monthly count of blocked changes — a listener that was never examined and a listener that was examined and found clean are indistinguishable. That is why this is called a false green rather than a false negative: the signal is not merely wrong, it is *identical to the right answer*. And it is stable. It does not fire once and recover; every plan that keeps computing the value at apply time keeps passing, silently, for as long as the pattern lives in the module. ## The two rule shapes and how each breaks | Rule shape | On a good value | On a bad value | On an unknown | | --- | --- | --- | --- | | Deny when the value is in a forbidden set | passes | violates | **passes silently** | | Deny unless the value is in an approved set | passes | violates | violates, with a wrong reason | Neither is correct. The denylist under-reports; the allowlist over-reports and lies about why. Candidates who answer "just use an allowlist" have understood half of it: yes, the change is blocked and nothing unsafe ships, but the message says the team chose a forbidden TLS policy when in fact nobody knows what they chose. That message sends an engineer hunting for a misconfiguration that is not there, and it is the fastest way to teach a team that the gate cries wolf. ## Writing the rule so unknown is a real case The fix is structural, not a change of operator. Make the rule ask three questions rather than one: 1. Is the attribute unknown at plan time? Read the plan's unknown marker for that specific path, not just the value. 2. If it is known, is it acceptable? 3. If it is neither known nor set, is "absent" itself a violation for this control? Each branch gets its own outcome and its own message. What you do with the unknown branch — block, warn, or record it as unevaluated — is a policy decision, and it can differ per control. What must not happen is that the branch does not exist, because then the decision has been made by accident. ## Reporting, not just deciding A gate that handles unknowns properly should be able to answer the question "how many resources could you not evaluate on this run?" If that number is not in the output, nobody can tell a clean run from a blind one, and the rule's pass rate is not measuring what people think it measures. Even in a report-only rollout this number is the most useful thing the gate produces: it tells you, before you make anything blocking, how much of the estate the rule can actually see. ## Test the rule against the shape that breaks it The reason this defect ships is that policy rules are usually tested against hand-written fixtures where every value is a literal. Keep at least one fixture per rule in which the attribute under test is unknown, and assert on the outcome you intended. A rule that has only ever been run against complete plans has not been tested against the one document shape that is deliberately incomplete.

  • How do you rewrite the rule so an unknown value cannot produce a silent pass?
    Stop relying on a comparison failing to match. Read the plan's unknown marker for that attribute path first and branch on it, emitting a deliberate outcome — violation, warning, or a recorded "not evaluable" finding carrying the resource address. The requirement is that the gate's output distinguishes checked-and-clean from could-not-check; which of the three outcomes you pick is a separate decision per control.
  • Does phrasing the rule as an allowlist fix the problem?
    It changes the direction of the error, not the error. Deny-unless-approved does fire on an unknown, so nothing unsafe slips through, but the finding claims the listener uses a forbidden TLS policy when the truth is that the value is not computable yet. That is a fail-closed outcome with a false message, and it burns engineer time chasing a misconfiguration that does not exist.
  • How would you test a rule against unknowns?
    Keep plan fixtures that genuinely contain unknown values for the attribute under test, not only hand-written literals, and assert the intended outcome on each. Generating one is easy — plan a configuration where the value comes from a resource being created in the same run. Without that fixture you have tested only the case the rule already handles.

saying these in an interview costs you the question

  • Claims an unknown value is treated as a violation by default
  • Says a green gate run proves the attribute was checked
  • Thinks flipping the comparison operator solves it
  • Assumes the engine errors out when the value is missing
  • Reports pass counts without counting unevaluated resources

context