Your storageClass admission rule has enforced for a month — why can it still not tell you how many existing volumes violate it?
answer
- which requests does the engine ever see
- write path only, one object at a time
- nothing re-reads what is already stored
- violations as queryable objects, not log lines
basics
~20 sAdmission runs only on the write path. It sees objects as they are created or updated, and it was never shown anything already stored in the cluster. Counting today's violations needs a background evaluation that reports on existing objects as data.
solid answer
~50 sAn admission decision is made about one object on its way into the API server. The review the engine receives carries that object, the previous version on an update, the requesting user and groups, and whether the request is a dry run — and nothing else: no other objects, no cluster state, no history. So a rule turned on today has an opinion about every future write and no opinion at all about the thousand claims already in etcd; those are only re-evaluated if something writes them again. That is why the report surface is a separate adoption criterion. The engine needs a background scan that lists matching objects, evaluates them, and records violations as queryable data — Kyverno writes report resources into the API, Gatekeeper's audit controller records violations on the constraint's status. That is what answers "how many, in which namespaces, owned by whom" and lets you watch the number fall.
go deeper
Know that an admission rule only ever sees objects being created or updated, so switching one on today changes nothing about workloads already running in the cluster.
Explain the mechanics: the review the engine receives carries one object plus the requester, and no other cluster state, so answering "how many violate this now" requires a separate background scan that writes results somewhere you can query.
Demonstrate that you would size the problem before enforcing, use the report as the migration backlog, and know the operational costs — cluster-wide watches, interval-based staleness, and per-policy caps that truncate a first scan.
Own the position that an engine without a usable report surface cannot support the conversation the organisation will actually have about a guardrail, and make that a stated criterion at adoption rather than a discovery six months in.
## Admission is a write-path control A validating admission rule is invoked because someone is writing an object. The API server sends the engine a review containing the object being written, the old object when it is an update, the identity of the requester and their groups, and a flag saying whether this is a dry run. What it does **not** contain is anything else in the cluster: no other objects, no aggregate state, no history. The rule renders a verdict on one object, in isolation, at the moment of the write. Everything else about the criterion follows from that single fact. ## What the gap looks like in practice You add the rule that dynamically provisioned volumes must use an approved encrypted storage class. From that moment, every new or updated claim naming an unapproved class is rejected. Meanwhile: - Claims created last year with an unencrypted class are untouched. Nothing writes them, so nothing evaluates them. - A workload that never gets redeployed never gets re-checked. - Your admission logs, no matter how clean, describe only the requests that arrived. A month of zero denials could mean the estate is clean or it could mean nobody created a volume. So the enforcing rule cannot answer the first question anyone asks about a new guardrail: how big is the problem we already have? Treating a green admission history as a compliance statement is the canonical mistake here, and it is the version of this mistake a platform interviewer probes. ## The report surface The capability that closes the gap is a background or continuous evaluation: the engine lists (and watches) the objects a policy matches, evaluates them outside the write path, and records the results. The design point to insist on is **where the results land**. Log lines are nearly useless for this — they are unstructured, they scroll away, they cannot be grouped by namespace and they cannot be joined to ownership. What you want is results written back into the API as objects: Kyverno emits report resources you can query with `kubectl`, and Gatekeeper's audit controller records violations in the status of the constraint. Either shape gives you a queryable inventory with the cluster's own RBAC on it, which you can list per namespace, aggregate onto a dashboard, and diff week over week. That inventory does several jobs at once: - **Sizing.** Forty-one claims across nine namespaces is a migration plan; "some volumes are probably unencrypted" is not. - **Progress.** The count going down is the only evidence the migration is happening. - **Drift after an exception.** When one namespace is excepted, the report still shows what is inside it, so the carve-out does not become invisible. - **Reconciliation.** If the engine was ever disabled during an incident, the background scan is how you find out what slipped in while it was off. ## What background scanning does not do It reports; it does not retroactively fix. Finding a violating claim does not delete the volume or evict the workload — that would be a catastrophic default for a control whose whole purpose is preventing new mistakes. Some engines can be configured to act on existing resources; treat that as a separate and much more dangerous capability with its own review, not as part of turning reporting on. ## The costs you should ask about Background evaluation is not free, and this is where the platform chair matters: - The engine has to list and watch every object kind its policies match, cluster-wide. On a large cluster that is real memory in the engine and real load on the API server. - Scans run on an interval, so the report is eventually consistent — it is a picture of a few minutes ago, not of now. - Engines cap how many violations they will record per policy so a report object does not grow without bound. A brand-new rule on a big estate can therefore return a **truncated** list, and reading that truncated number as the total is a real and easy mistake. Check the limit before you quote a figure. ## The interview shape The question is usually asked as a trap with an obvious wrong answer available: the rule is enforcing, so the cluster must be compliant. The strong answer separates the two controls cleanly — admission decides about writes, background evaluation describes state — and then treats "does this engine have a usable report surface" as something you check before adopting it, not something you discover you need after the first audit-style question arrives.
- What does a background scan do when it finds an existing violating volume?It records it. Reporting is the safe and normal outcome — deleting or evicting an existing workload because a new rule landed would be an outage caused by a policy change. Enforcement stays on the write path, and the report becomes the migration backlog. Some engines can be configured to act on existing resources; that is a separate, far riskier capability to review on its own.
- Why insist violations land as API objects rather than in the engine's logs?Objects are queryable with the cluster's own tooling and RBAC, have a stable shape you can aggregate by namespace or owner, survive pod restarts, and can be diffed week over week to show progress. Logs are unstructured, scroll away, and answer no question more complex than "did something happen".
- A fresh rule's report shows exactly twenty violations across a huge cluster. What would you check first?Whether the number is truncated. Engines cap how many violations they record per policy so report objects stay bounded, and a round number on a first scan is a strong hint you are reading the cap rather than the total. Raise or check the limit, rescan, and only then quote a figure.
A turnstile controls who walks in from now on. It tells you nothing about who is already inside the building; for that you have to walk the floors and count.
saying these in an interview costs you the question
- Treats an enforcing rule as proof the cluster already complies
- Believes admission re-evaluates objects already stored in the cluster
- Offers webhook logs as the compliance report
- Thinks a background scan deletes or evicts offending workloads
- Quotes a first-scan violation count without checking it is truncated