skip to content

A Kyverno rule stopped blocking after its ConfigMap was renamed - how would you have caught it?

level: seniorimportance: should knowfreq 43%

answer

  1. zero violations has two causes
  2. the report result column, not the count
  3. error and skip are different signals
  4. submit something known-bad on purpose
  5. the lookup is an undeclared dependency

basics

~20 s

Alert on rule results, not on violations. A failed context lookup produces an error result and a skipped precondition produces a skip - both show up as zero violations, so watch error and skip counts and run a deliberately non-compliant canary against the cluster.

solid answer

~50 s

The trap is that "no violations" and "the rule never ran" look identical on a compliance dashboard. Kyverno records per-rule outcomes in policy reports with distinct results - `pass`, `fail`, `warn`, `error` and `skip`. A context entry that cannot resolve, because the ConfigMap was renamed or Kyverno's ServiceAccount lost the read, yields `error` for that rule, plus events and controller logs; an unmet precondition yields `skip`. So the monitoring has to be built on those: alert when a policy's error or skip rate rises, or when a policy that normally emits thousands of `pass` results emits none. Second, prove enforcement positively - submit a known-bad manifest as a server-side dry run on a schedule and page if it is admitted. Third, treat the ConfigMap as a deployment dependency of the policy: same GitOps application, same review, and a name nobody can rename in isolation.

go deeper

for a junior

Know that a rule which cannot read its data does not block - it reports an error - and that a dashboard counting violations will show nothing wrong.

for a middle

Explain the distinct report results and which one a failed lookup versus an unmet precondition produces, and where the controller logs name the failing entry.

for a senior

Demonstrate the operational answer: alert on error and on the disappearance of pass results, and run a deliberately non-compliant canary as a dry run to prove enforcement.

for a principal

Argue for the structural fix - policy and its data shipped and reviewed together - and for treating a control that fails silently as a higher-priority defect than one that fails loudly.

## Why this failure is quiet A Kyverno rule that reads an allow list from a ConfigMap has a runtime dependency nothing validated up front. The policy was accepted when it was created - its structure is fine. The ConfigMap is resolved per admission request. Rename it, move it to another namespace, or narrow the RBAC that lets Kyverno read it, and from that moment the rule cannot evaluate. It does not start blocking everything and it does not start allowing loudly. It produces an error for that rule and the request continues down the pipeline. Meanwhile the security dashboard everyone actually looks at counts violations. Violations go to zero. Zero violations is what success looks like. ## The signal that distinguishes the two Kyverno writes per-resource, per-rule outcomes into policy reports, and the result values are distinct on purpose: | result | what it means | | --- | --- | | `pass` | the rule evaluated and the resource satisfied it | | `fail` | the rule evaluated and the resource violated it | | `warn` | reported without being treated as a violation | | `error` | the rule could not be evaluated at all | | `skip` | the rule did not apply - typically an unmet precondition | Those last two are the ones that matter here, and they are the ones nobody alerts on. Practical monitoring: - **Alert on `error`.** Any sustained error rate for a policy is an outage of that control, whatever the admission outcome was. - **Alert on the absence of `pass`.** A policy that normally evaluates every Pod creation and suddenly evaluates none has stopped working, even if nothing errored - for example because a `match` block or a precondition was narrowed by a well-meaning edit. - **Watch `skip` as a ratio.** Skips are legitimate; a step change in the skip ratio is not. - **Read the events and controller logs.** Kyverno emits events for policy execution problems, and the controller logs name the entry that failed, which is what turns "a control is down" into "the allowed-placement ConfigMap is gone". The admission controller also exposes Prometheus metrics for policy results labelled by policy, rule and outcome, which is usually the cheapest place to hang the alerts. ## Prove it positively Monitoring tells you when something changed. A canary tells you the control still works. Keep a manifest that is deliberately non-compliant - a Pod targeting a zone that is definitely not on the list - and submit it on a schedule as a server-side dry run from a CI job or a small controller. If it is admitted, page. This is the single highest-value check for any guardrail whose enforcement depends on data it fetches, because it exercises the whole path: match, context lookup, precondition, condition, decision. ## Why CLI tests did not save you The Kyverno CLI can test a policy against fixture resources and lets you supply context values from a values file. That is exactly what makes it fast, and exactly why it is blind to this failure: the ConfigMap the test used came from a file, so the test stays green while the cluster's real ConfigMap does not exist. CLI tests verify the *logic* of the rule. They cannot verify its *dependencies*. Both need covering, in different places. ## Make the dependency explicit The durable fix is structural rather than observational: - Ship the ConfigMap and the policy in the same GitOps application so they are created, updated and pruned together, and a rename is one commit that touches both. - Keep the ConfigMap in a namespace owned by the platform team, not in the namespaces the policy governs, so an ordinary namespace admin cannot rename or delete it in the course of tidying up. - Guard it with a Kyverno policy of its own if it matters enough: a rule that refuses deletion or an update from an unexpected subject. - Record, in the policy's own repository, that the rule has an external dependency. The next person to read the YAML sees a `context` block; they do not see that a ConfigMap in another namespace is load-bearing. ## What to say in the interview The strong answer names the specific asymmetry: a control that fails by not firing produces the same telemetry as a control that is passing. Everything else - error alerts, canaries, GitOps coupling - follows from taking that seriously. The weak answer is "we would see it in the policy reports", which is true only if someone is looking at the column nobody looks at.

  • How do you tell an errored rule from a skipped one, and why does the distinction matter?
    The policy report result says which: `error` means the rule could not be evaluated - typically the context lookup failed - while `skip` means it did not apply, usually an unmet precondition. Error is an outage of the control and should page. A rising skip ratio is quieter but just as interesting: it usually means a match block or precondition was narrowed and the rule now covers less than you think.
  • Your policy tests pass in CI. Why did they not catch this?
    CLI tests supply context values from a fixture file, so they exercise the rule's logic against data the test provides. The failure here is the rule's dependency on a real ConfigMap in a real cluster, which no fixture can represent. You need a check that runs in the cluster - a scheduled server-side dry run of a known-bad manifest is the cheapest one that proves the whole path end to end.
  • What stops this recurring after you fix the ConfigMap?
    Couple the data to the policy: same GitOps application so they deploy and prune together, a platform-owned namespace so team admins cannot rename it, and an alert on the policy's error and pass counts. Optionally a policy guarding the ConfigMap itself. The goal is that the rename becomes impossible in isolation rather than merely detectable afterwards.

saying these in an interview costs you the question

  • Treats zero violations as proof the control works
  • Alerts only on fail results, never on error or skip
  • Believes policy creation validates the ConfigMap exists
  • Relies on CLI tests to cover a live cluster dependency
  • Cannot name any positive test that the gate still blocks

context