A legacy namespace cannot meet your encrypted-storageClass rule — why express the exception as a cluster object rather than editing the rule?
answer
- where does the carve-out physically live
- one rule text, or one per cluster
- how do you list live carve-outs
- revoking it: release or delete
- exceptions are themselves policy
basics
~20 sEditing the rule weakens it for everyone and buries the carve-out in policy source. An exception object keeps one rule text everywhere, states exactly what is excluded and where, and is listable, access-controlled and revocable by deleting it.
solid answer
~50 sIf I add `namespace != "legacy-billing"` to the rule, three things break at once. The rule text stops being identical across clusters and in the CI check, so the offline tests now have to encode a carve-out that has nothing to do with encryption. The exception becomes invisible — you find it by reading policy source, not by listing anything. And withdrawing it is a policy release rather than deleting an object. Expressing it as an object the engine reads fixes all three: the rule stays byte-identical wherever it runs, the exception names the policy, the namespace and the resource kinds it covers, RBAC decides who may create one, `kubectl get` answers "what exceptions are live right now", and removing it restores enforcement with no change to the rule. Kyverno models this as a dedicated exception resource; Gatekeeper constraints exclude namespaces in their match block. Native support for this is an adoption criterion, not a nice-to-have.
go deeper
Know that an exception should be recorded as its own object in the cluster rather than by editing the rule, so the same rule keeps running everywhere and the carve-out can be found and removed.
Explain what the object narrows — which policy and rule, which namespaces, which kinds — and why a scoped object beats an in-rule condition whose reach is inferred rather than stated.
Show the operational consequences you would insist on: one rule text across cluster and CI, a listable inventory of live carve-outs, deletion as the revocation path, and the excepted objects still visible in reporting.
Own the position that exceptions are policy too. Decide who may grant one, constrain their breadth with a rule of their own, and treat native exception support as a criterion an engine must meet before it is adoptable.
## The criterion Every guardrail eventually meets a case it cannot accommodate. A legacy namespace runs on a storage class that predates the encrypted one and the migration is a quarter away. The question is not whether to allow the exception — you will — but **where the exception lives**. The criterion for an in-cluster engine is that it can express one as a first-class object the engine reads, without touching the rule. ## What breaks when the exception lives inside the rule The tempting fix is a clause in the rule body excluding that namespace. It works, and it costs you four separate things. **One rule text, everywhere.** The reason a rule can be unit-tested offline and re-run in CI over rendered chart output is that there is exactly one artifact. Fork it to carry a per-cluster carve-out and you now have variants: the production copy with the exclusion, the CI copy without, and a test suite that must encode both. They drift, quietly, in the direction of the copy nobody runs. **Discoverability.** An exception in the rule body is found by reading policy source. Ask "what carve-outs are live in this cluster right now?" and the honest answer is a `grep` across a repository plus a hope that what is deployed matches. An exception object answers the same question with a list. **Blast radius.** A clause in the rule is evaluated for every object the rule matches. A mistyped or over-broad condition silently widens the hole for the whole cluster, and it looks like a normal policy edit in review. An exception object is scoped by construction — it names the policy and rule it relaxes, the namespaces it applies to, and typically the resource kinds and names — so its reach is stated rather than inferred. **Reversibility.** Withdrawing an in-rule exception is a policy change: edit, review, test, release. Withdrawing an exception object is deleting the object, which any operator can do during an incident and which leaves no trace in the rule's history. ## What the object should be able to say When judging an engine's exception support, look for the ability to narrow along the axes you will actually need: which policy and which rule inside it, which namespaces, which resource kinds, and ideally which specific object names. A mechanism that can only say "turn this policy off in this namespace" is coarse — it exempts the namespace from every rule in the policy, including the ones it could easily satisfy. Also check that an exception is scoped to a namespace when it should be: an exception object that lives in the excepted namespace, and can only affect that namespace, is far safer than a cluster-scoped one whose selector could accidentally match everything. ## Guarding the guard Exceptions are policy. Once they are objects, they inherit the cluster's own controls, and you should use them deliberately: - **RBAC.** Decide who may create an exception object. If any namespace owner can create one for their own namespace, that is a legitimate design — self-service with an audit trail — but it must be a decision, not an accident of default permissions. - **Validate the exceptions themselves.** The dangerous shape is an exception with an empty or wildcard selector, which silently disables the rule cluster-wide. A rule that constrains exception objects is a normal and worthwhile thing to write. - **Keep them visible in reporting.** An excepted object should still show up somewhere as excepted rather than disappearing from the report entirely, otherwise the carve-out becomes invisible again by a different route. ## Interaction with the rest of the criteria This criterion is not independent of the others, which is why it belongs in the same evaluation. Because the exception is an object, the rule keeps passing the same offline tests — the test suite never learns about `legacy-billing`. Because the rule text is unchanged, the CI check over rendered chart output keeps working for every other team. And because the exception is an object, the background scan can still report what is inside the excepted namespace, so the migration has a number attached to it. An engine that only lets you weaken the rule takes all three of those away at once. ## The answer that lands A strong answer names the mechanism and then the operational consequences: one rule text, a listable inventory of live carve-outs, RBAC over who grants one, deletion as the revocation path, and exception objects that are themselves subject to policy. A weak answer treats "add an if-clause" and "create an exception object" as two spellings of the same thing.
- What specifically breaks if the carve-out is an `if namespace == "legacy-billing"` clause inside the rule?The rule text stops being identical across clusters and in the CI check, so its offline tests must encode a carve-out unrelated to encryption. Nobody can list live exceptions without reading policy source. An over-broad condition widens the hole for every matched object and looks like an ordinary policy edit in review. And revoking it is a policy release rather than deleting an object.
- Who should be allowed to create an exception object, and what would you constrain?Grant it through RBAC deliberately — often the namespace owners for exceptions scoped to their own namespace, so it is self-service with an audit trail. The shape to constrain is breadth: an exception with an empty or wildcard selector disables the rule cluster-wide. Writing a policy that validates exception objects, and preferring namespace-scoped ones, keeps the carve-out as narrow as it claims to be.
- How does the exception interact with your background report?An excepted object should still appear in reporting, marked as excepted rather than omitted. Otherwise the carve-out becomes invisible again by a different route and the legacy namespace drops out of the migration count. The point of an object is that the exception stays countable.
saying these in an interview costs you the question
- Adds a namespace check inside the rule and calls that an exception
- Grants a cluster-wide exception when one namespace asked
- Disables the whole policy while the team migrates
- Assumes exception objects need no access control of their own
- Maintains a separate copy of the rule per cluster to hold carve-outs