You own a shared Rego policy library for 40 teams — standardize on deny sets or one allow decision?
answer
- who can contribute, and how loudly it breaks
- additive files versus one shared conjunction
- blast radius of a bad predicate
- buy fail-closed at the harness
- a short, written list of exceptions
basics
~20 sStandardize on deny sets for almost everything, because they are additive, tool-native and produce usable messages, then buy back the fail-closed property in the harness. Reserve an explicit allow decision for the few gates where a silently absent rule is unacceptable.
solid answer
~50 sDeny sets win on contribution economics: a team adds a rule as a new file without touching anyone else's, the runner finds it by name with no configuration, and each violation carries its own message. One central `allow` conjunction is fail-closed by construction, but every new rule edits a shared expression — merge contention, one bad predicate blocking all 40 teams, and a single boolean where developers need reasons, so you end up rebuilding the message set alongside it. My call is deny sets as the default with the safety bought back around them: pin the package and namespace in the gate invocation, keep non-compliant fixtures that the library's own CI asserts are rejected, and fail the gate when it reports zero rules evaluated. Then pick the handful of controls where a silent no-op is genuinely unacceptable and give those an explicit allow decision with a harness that treats undefined as denial. Two shapes with a written rule for which is which beats one shape applied dogmatically.
go deeper
Know that the two shapes exist and that a deny set is what most published policy libraries use because the runner resolves it by name.
Be able to explain the concrete differences a contributor feels: adding a file versus editing a shared expression, and one message per violation versus a single boolean.
Show how you would make a deny-set library detect its own silence — pinned resolution, rejected fixtures, and failing when no rules evaluated.
Own the decision and its cost, keep the deviation list short and written, and make sure the property you actually depend on is that someone notices when the gate stops working.
## What is actually being decided This is not a Rego syntax preference. It is a decision about who can contribute to the guardrails, how loudly the system fails when it breaks, and what a developer sees when it stops them. **Deny sets.** Each rule is a partial set of messages, resolved by name by the runner. Rules are additive and independent. **One allow decision.** A complete rule with `default allow := false`, true only when every condition holds. Approval must be earned explicitly. ## The case for deny sets across an estate **Contribution is cheap and local.** A platform team adds a file. Nothing existing is edited, so review is scoped to the new rule and merge contention is near zero. Across 40 teams that difference compounds: the library grows because contributing is a pull request, not a negotiation over a shared expression. **The runner needs no configuration.** Rule names are the interface — `deny`, `violation`, `warn`, optionally suffixed, inside a namespace. Any published policy drops into the standard invocation. An allow-shaped library has no such convention; you write the query and the harness, and every consumer must agree on them. **Messages come for free.** A set element *is* the sentence a developer reads. A boolean is one bit; to match the usefulness of a deny set you must collect reasons next to it, and now two things must stay in sync. **Blast radius is bounded.** A badly written deny rule blocks the changes it matches. A bad predicate inside one central conjunction makes `allow` false for everybody, and you have taken 40 teams offline with a typo. ## The case against them, which is the whole problem A deny set cannot report its own absence. Empty is what compliance looks like, so a renamed package, an unselected namespace, an uncopied policy directory or a rule that no longer matches the document all render as green. You do not find out from the gate; you find out from the incident it was supposed to prevent. An allow decision that stops working blocks everything and is escalated within the hour. That is a real asymmetry and it is why the question gets asked. But notice what it is really an argument for: not a different rule shape, a different **harness**. ## Where I land, and the reasoning Standardize on deny sets, and treat the fail-open property as an operational problem you solve once, centrally, rather than a language problem you solve 400 times in every rule body: 1. **Pin resolution.** The gate invocation names the package and namespace explicitly, so a rename becomes a configuration error instead of a quiet no-op. 2. **Assert the gate fires.** The library ships non-compliant fixtures — a pipeline definition whose build-log retention sits below the floor — and its own CI asserts each is rejected. Any change that silences a rule turns the library's build red before a consumer ever sees it. 3. **Treat zero as suspicious.** The gate reports how many rules it evaluated, and zero fails the job. This alone converts the entire class of "nothing ran" failures from invisible to loud. 4. **Watch the aggregate.** A control that has produced no findings anywhere for a long stretch is a hypothesis about the estate, not a fact about it — worth a look either way. Then carve out the exceptions deliberately. For the small number of gates where a silent no-op is unacceptable, write an explicit allow decision with a harness that admits only a literal `true` and treats undefined, missing or errored results as denial. Keep that list short and written down, with the reason each entry is on it. The failure mode of this design is a library where nobody remembers which shape a given control uses, so the rule for choosing has to be documented, not folklore. ## The organisational half of the answer Uniformity has value beyond either shape's merits: 40 teams reading one library should not have to learn two conventions to read a rule. So the deviation must be rare and justified per control, not per author's taste. And whichever shape you pick, the property that actually protects you is not in the policy language at all — it is that somebody would notice if the gate stopped working. Build that first; it is worth more than the shape debate. ## What an interviewer is listening for A position, held for stated reasons, with the cost of the position named. Candidates who say "allow, because fail-closed" without pricing the merge contention, the blast radius and the missing messages have optimised one property in isolation. Candidates who say "deny sets, everyone does" without a plan for the silent no-op have not thought about how it fails. The strong answer picks the default, buys back the missing property at the harness, and reserves the other shape for a short, deliberate list.
- What is the blast radius argument against one central allow decision?Every rule is a term in a single conjunction, so one over-broad predicate makes the decision false for every team at once. Deny sets fail in proportion to what they match: a bad rule blocks the changes it matches and the rest of the estate keeps shipping.
- How do you decide which controls get the allow shape?The ones where a silent no-op is unacceptable and the condition is narrow enough to state positively — a small set, written down with the reason each entry qualifies. If the list starts growing, that is a signal the harness safeguards are not trusted and should be fixed instead.
- A team argues the library should be fail-closed everywhere. What is your counter?That fail-closed is worth having, but the shape is an expensive way to buy it: it costs merge contention, blast radius and per-violation messages. The same property comes from pinning resolution, asserting fixtures are rejected, and failing when zero rules evaluated — bought once, centrally.
saying these in an interview costs you the question
- Picks a shape without naming what it costs
- Assumes an allow decision cannot break silently
- Ignores merge contention on one shared expression
- Forgets developers need reasons, not one boolean
- Leaves the choice to each author with no written rule