skip to content

A Rego complete rule ran green for months and now errors on one plan. Why?

level: seniorimportance: should knowfreq 44%

answer

  1. how many resources matched this time?
  2. one value per path, per evaluation
  3. iteration bound the variable twice
  4. the head promised a cardinality the body broke
  5. default fills undefined, not over-defined

basics

~20 s

Most likely the input finally contained two matching resources. A complete rule may produce only one value, so a body that iterates and succeeds twice with different values raises an evaluation conflict and the whole query fails.

solid answer

~50 s

A complete rule assigns one value to its path, and that limit applies across every body and every variable binding, not just across bodies you wrote. If the body iterates — `some rc in input.resource_changes` — it is evaluated once per element, and each success produces a candidate value. One offending resource yields one value and everything works; two offenders yield two different messages and the engine raises a conflict error rather than picking one, so the query fails outright instead of returning a decision. Tests stayed green because no fixture contained two violating resources. The fix is to match the shape to the cardinality: make it a partial set of messages, or a partial object keyed by the resource address so each failure gets its own entry. `default` does not help — it fills an undefined rule, and this one is over-defined.

go deeper

for a junior

Know that a complete rule holds one value and that a body containing iteration can try to produce several — enough to recognise the error message rather than be baffled by it.

for a middle

Explain that the one-value limit spans all bodies and all bindings, that identical values are fine while differing ones conflict, and why a partial rule cannot hit this.

for a senior

Diagnose it from the input rather than the rule: identify the binding that multiplied, choose the shape that matches the real cardinality, and add the two-offender fixture that would have caught it.

for a principal

Decide what a gate does when the engine itself errors rather than denies, since a failed query is neither allow nor deny, and set the review habit that catches cardinality mismatches before they reach a pipeline.

### A complete rule may produce one value — across every binding, not just every body The failure looks like a mystery because the policy did not change and the engine did not change. What changed is the *input*: for the first time, two resources matched the same rule. Consider a rule written to explain a machine-family violation: ```rego package estate reason := msg if { some rc in input.resource_changes not rc.change.after.machine_type in data.catalog.families msg := sprintf("%v uses an unapproved machine type", [rc.address]) } ``` `reason` is a **complete rule**: its head assigns one value to `data.estate.reason`. The body, however, iterates. `some rc in input.resource_changes` binds `rc` to each element in turn, so the body is evaluated once per resource, and every evaluation that succeeds produces a value for the head. With one offending resource that is one value and everything works. With two offending resources the body succeeds twice with two different strings, and the engine cannot pick one: it raises an evaluation conflict — a complete rule must not produce multiple outputs. The whole query fails; the gate does not return "deny", it returns an error. Two details make this sharper: - Multiple successes with the **same** value are fine. The constraint is one *value*, not one success. That is why a rule like `deny := true if { ... }` never trips over this, and why the bug hides in rules whose value is derived from the matched item — a message, an address, a severity. - The same trap exists in **partial object** form. Keying by something that is not unique per match — `reason[rc.change.after.machine_type] := msg` when two resources share a type — reproduces the conflict one level down. ### Why the tests were green Because the fixtures had at most one violating resource. A conflict of this kind is invisible to any test whose input cannot produce two distinct bindings. The regression test to add is not "assert the message" but a fixture with **two** offending resources — and, for the object form, two offenders that collide on the chosen key. ### `default` does not fix it The most common wrong answer in the room. `default reason := "ok"` supplies a value when the rule is **undefined** — when no body succeeded. This rule is the opposite failure: it is *over*-defined. `default` is never consulted, and it is not permitted on a partial rule at all, so it cannot be smuggled in that way either. ### The fixes, in order of preference 1. **Change the shape to match the cardinality.** If the rule is inherently per-resource, it should not have been a complete rule. Write it as a partial set of messages (`reason contains msg if { ... }`) or, better for a gate that must attribute failures, a partial object keyed by the resource address, which is unique within a plan: `reason[rc.address] := msg if { ... }`. The caller now receives every failure in one query instead of an error. 2. **Aggregate deliberately.** If the caller genuinely needs one value, compute it from the set: `summary := sprintf("%d resources violate the catalogue", [count(reason)])`. The complete rule then depends on the partial one and can only ever produce a single value. 3. **Narrow the body so it can match once.** Legitimate only when uniqueness is guaranteed by the data — a lookup by primary key, for instance. It is fragile as a *fix* for this bug, because the guarantee usually lives in someone's head rather than in the input. ### What an interviewer is really testing Whether you understand that a rule's head is a cardinality declaration, and that iteration inside the body multiplies results regardless of what the head promised. A candidate who reaches straight for `default`, or who says the engine "returns the first match", has the mental model of a function rather than of a document — and that model will keep producing this bug.

  • Why did the unit tests never catch this?
    Because every fixture had at most one violating resource, and the conflict only appears when the body succeeds with two distinct values. The regression test is not another message assertion but a fixture containing two offending resources — and, if the rule is a keyed object, two offenders that collide on the chosen key.
  • Would keying a partial object by resource type instead of address fix it?
    No, it reproduces the same conflict one level down. Two instances of the same type both insert under one key with different reasons, and a partial object rejects a key assigned two different values. Key by something unique per match — the resource address within a plan — so every failure gets its own entry.
  • When is a complete rule still the right shape for something derived from many resources?
    When the caller genuinely needs one value and you compute it deliberately from a partial rule: a boolean, a count, or a summary string built with `count()` over the set. The complete rule then depends on the partial one and can only ever produce a single value, whatever the input contains.

saying these in an interview costs you the question

  • Blames a syntax error the parser should have caught
  • Says the engine returns the first matching value
  • Adds default to fix an over-defined rule
  • Assumes the two values are collected into a list
  • Keys a partial object by resource type, reproducing the conflict

context