What does Gatekeeper's audit controller do, and where does it write what it finds?
answer
- a second loop, not the webhook
- runs on a timer, not on a write
- looks at objects already there
- writes findings back onto the rule object
- status.violations plus totalViolations
basics
~20 sGatekeeper's audit controller periodically re-evaluates objects that already exist in the cluster against every enforced Constraint and records the violations in each Constraint object's own status field. It only reports: it never blocks, deletes or changes anything.
solid answer
~50 sGatekeeper runs two loops over the same rules. The admission webhook decides on writes as they arrive; the audit controller is a separate deployment that wakes on `--audit-interval` (60 seconds by default), evaluates the objects already living in the cluster against every enforced Constraint, and records what it finds. The findings go back onto each Constraint object's own `status`: a `violations` list naming the kind, name, namespace and the rule's message, plus `totalViolations` and an `auditTimestamp`. Audit is a detective control with no teeth of its own — nothing is denied, evicted or mutated. That is why it is how you answer "how bad is it already?": turn on a rule requiring PersistentVolumeClaims to use an approved encrypted StorageClass, wait one interval, and the Constraint's status tells you how many of the four hundred claims already in the cluster fail it.
go deeper
Be ready to say plainly that Gatekeeper has two loops: the webhook that decides on writes, and audit, which runs on a timer over objects that are already there and only records what it finds.
Expect to name the mechanics: the audit interval, that findings land in each Constraint's own status with a violations list and a total, and that a dryrun Constraint is audited just like an enforcing one.
Show that you treat audit as a periodic snapshot with a cost. Know that a stale timestamp, a dead audit pod or a disabled interval all look like clean results, and that the sweep's read load is a real budget line on a large cluster.
Own the question of what the audit signal is for. Findings sitting in a status field reach nobody by default, so decide deliberately whether audit is an investigation tool, a feed into something that pages a team, or an expense you turn down.
## Two loops over the same rules Gatekeeper enforces policy in two places, and they are separate processes with separate failure modes. The **admission webhook** is synchronous and per-request: the API server calls it while a create or update is in flight, and the answer decides whether that write lands. It sees exactly one thing — the object being written. The **audit controller** is asynchronous and periodic. It ships as its own deployment (a single replica by default, separate from the webhook pods) and its job is the opposite one: take the objects that are *already* in the cluster and check them against the rules as they stand today. Nothing is being written when audit runs; nothing is waiting on its answer. ## What a sweep actually does On each cycle — the period comes from `--audit-interval`, 60 seconds by default, and setting it to `0` disables audit entirely — the controller walks the Constraints currently installed, works out which kinds they match, reads those objects, and evaluates each one against the matching Constraint's rule. The result is a set of violations per Constraint. A Constraint whose `enforcementAction` is `dryrun` is audited exactly like an enforcing one. Dryrun only tells the *webhook* not to deny; it does not tell audit not to look. That combination — a rule that blocks nothing while its status fills up with findings — is the normal first phase of introducing a new rule. ## Where the findings land There is no central report object. Each Constraint carries its own results in its own `status`: - `violations` — a list of entries, each with the offending object's `kind`, `name` and `namespace`, the `enforcementAction` in force, and the `message` the rule produced. This list is **capped**; it is a sample, not an inventory. - `totalViolations` — the count the sweep found, reported separately from that capped list. - `auditTimestamp` — when the results were last refreshed. Everything in the status is as of that moment, not as of now. - `byPod` — per-pod bookkeeping, since more than one Gatekeeper pod may report status. If two Constraints match the same object, each records its own violation on its own status. To see everything audit found you read every Constraint. ## What audit does not do It does not enforce. It does not evict a running Pod, it does not delete a PersistentVolumeClaim, and it does not mutate anything into compliance. It also does not notify: a `status` field is not an alert, and nothing watches it unless you build that. A Constraint can sit in dryrun for a quarter with hundreds of findings recorded and no human ever look. It is also not continuous. Between sweeps the picture is stale by up to one interval, and if the audit pod is down or a sweep errored partway, the status simply keeps the last numbers it managed to write — old results look identical to fresh ones apart from the timestamp. ## What it costs A sweep is not free. Reading every object of every constrained kind, every interval, is real load on the API server and real memory in the audit pod on a large cluster. Gatekeeper gives you dials for this: lengthen `--audit-interval`, restrict the sweep to kinds some Constraint actually matches with `--audit-match-kind-only`, page the reads with `--audit-chunk-size`, or evaluate a replicated cache instead of the live API with `--audit-from-cache`. On small clusters the defaults are invisible; on a cluster with hundreds of thousands of objects, audit is often the most expensive thing Gatekeeper does. ## The example to carry You write a rule that PersistentVolumeClaims must name a StorageClass from an approved, encrypted-at-rest set. At admission, that rule affects exactly one population: claims created from now on. Meanwhile four hundred claims already exist, and several teams' data sits on them. Audit is the half of Gatekeeper that can tell you anything about those four hundred: after one interval, the Constraint's status carries a total and a sample of names. Nothing has moved, nothing has broken, and you now have a number to plan against — which is precisely the value, and precisely the limit, of a detective control.
- Does audit still report on a Constraint whose enforcementAction is dryrun?Yes. `dryrun` tells the admission webhook not to deny; it says nothing to the audit controller. The sweep evaluates that Constraint like any other and its violations accumulate in status. That is exactly why teams ship a new rule in dryrun first — audit gives them the size of the existing problem while nothing is blocked.
- How do you turn Gatekeeper's audit off, and why would anyone want to?Set `--audit-interval=0`. On a very large cluster the periodic read of every constrained kind is the heaviest thing Gatekeeper does, both for the API server and for the audit pod's memory. Teams under pressure lengthen the interval, narrow it with `--audit-match-kind-only`, or disable it and get their inventory picture from somewhere else.
- If two Constraints match the same PersistentVolumeClaim, where do the violations show up?On each Constraint's own status, separately. Gatekeeper produces no cluster-wide findings object; the Constraint is the unit of reporting. To see everything a sweep found you enumerate the Constraints and read each status, which is one reason teams scrape the results rather than reading YAML by hand.
The admission webhook is the door check on the way in. Audit is the night-shift stocktake that walks the shelves and writes down what is already on them — it counts, it does not confiscate.
saying these in an interview costs you the question
- Says audit deletes or evicts the offending objects
- Thinks audit runs per write like the webhook
- Expects a single cluster-wide findings object
- Assumes a dryrun Constraint is skipped by audit
- Treats the status list as a live view of the cluster