skip to content

Rules as Cluster Objects

Gatekeeper packages a rule as Kubernetes objects, so one team ships a reusable template and others configure it - at the price of a second control loop. Interviewers probe what admission never sees.

on this pageshow

explore

questions

23

What does Gatekeeper's audit controller do, and where does it write what it finds?

level: juniorimportance: must knowfreq 66%

answer

  1. a second loop, not the webhook
  2. runs on a timer, not on a write
  3. looks at objects already there
  4. writes findings back onto the rule object
  5. status.violations plus totalViolations

basics

~20 s

Gatekeeper's audit controller periodically re-evaluates objects that already exist in the cluster against every enforced Constraint and records the violations in each Constraint object's own status field. It only reports: it never blocks, deletes or changes anything.

solid answer

~50 s

Gatekeeper runs two loops over the same rules. The admission webhook decides on writes as they arrive; the audit controller is a separate deployment that wakes on `--audit-interval` (60 seconds by default), evaluates the objects already living in the cluster against every enforced Constraint, and records what it finds. The findings go back onto each Constraint object's own `status`: a `violations` list naming the kind, name, namespace and the rule's message, plus `totalViolations` and an `auditTimestamp`. Audit is a detective control with no teeth of its own — nothing is denied, evicted or mutated. That is why it is how you answer "how bad is it already?": turn on a rule requiring PersistentVolumeClaims to use an approved encrypted StorageClass, wait one interval, and the Constraint's status tells you how many of the four hundred claims already in the cluster fail it.

go deeper

for a junior

Be ready to say plainly that Gatekeeper has two loops: the webhook that decides on writes, and audit, which runs on a timer over objects that are already there and only records what it finds.

for a middle

Expect to name the mechanics: the audit interval, that findings land in each Constraint's own status with a violations list and a total, and that a dryrun Constraint is audited just like an enforcing one.

for a senior

Show that you treat audit as a periodic snapshot with a cost. Know that a stale timestamp, a dead audit pod or a disabled interval all look like clean results, and that the sweep's read load is a real budget line on a large cluster.

for a principal

Own the question of what the audit signal is for. Findings sitting in a status field reach nobody by default, so decide deliberately whether audit is an investigation tool, a feed into something that pages a team, or an expense you turn down.

## Two loops over the same rules Gatekeeper enforces policy in two places, and they are separate processes with separate failure modes. The **admission webhook** is synchronous and per-request: the API server calls it while a create or update is in flight, and the answer decides whether that write lands. It sees exactly one thing — the object being written. The **audit controller** is asynchronous and periodic. It ships as its own deployment (a single replica by default, separate from the webhook pods) and its job is the opposite one: take the objects that are *already* in the cluster and check them against the rules as they stand today. Nothing is being written when audit runs; nothing is waiting on its answer. ## What a sweep actually does On each cycle — the period comes from `--audit-interval`, 60 seconds by default, and setting it to `0` disables audit entirely — the controller walks the Constraints currently installed, works out which kinds they match, reads those objects, and evaluates each one against the matching Constraint's rule. The result is a set of violations per Constraint. A Constraint whose `enforcementAction` is `dryrun` is audited exactly like an enforcing one. Dryrun only tells the *webhook* not to deny; it does not tell audit not to look. That combination — a rule that blocks nothing while its status fills up with findings — is the normal first phase of introducing a new rule. ## Where the findings land There is no central report object. Each Constraint carries its own results in its own `status`: - `violations` — a list of entries, each with the offending object's `kind`, `name` and `namespace`, the `enforcementAction` in force, and the `message` the rule produced. This list is **capped**; it is a sample, not an inventory. - `totalViolations` — the count the sweep found, reported separately from that capped list. - `auditTimestamp` — when the results were last refreshed. Everything in the status is as of that moment, not as of now. - `byPod` — per-pod bookkeeping, since more than one Gatekeeper pod may report status. If two Constraints match the same object, each records its own violation on its own status. To see everything audit found you read every Constraint. ## What audit does not do It does not enforce. It does not evict a running Pod, it does not delete a PersistentVolumeClaim, and it does not mutate anything into compliance. It also does not notify: a `status` field is not an alert, and nothing watches it unless you build that. A Constraint can sit in dryrun for a quarter with hundreds of findings recorded and no human ever look. It is also not continuous. Between sweeps the picture is stale by up to one interval, and if the audit pod is down or a sweep errored partway, the status simply keeps the last numbers it managed to write — old results look identical to fresh ones apart from the timestamp. ## What it costs A sweep is not free. Reading every object of every constrained kind, every interval, is real load on the API server and real memory in the audit pod on a large cluster. Gatekeeper gives you dials for this: lengthen `--audit-interval`, restrict the sweep to kinds some Constraint actually matches with `--audit-match-kind-only`, page the reads with `--audit-chunk-size`, or evaluate a replicated cache instead of the live API with `--audit-from-cache`. On small clusters the defaults are invisible; on a cluster with hundreds of thousands of objects, audit is often the most expensive thing Gatekeeper does. ## The example to carry You write a rule that PersistentVolumeClaims must name a StorageClass from an approved, encrypted-at-rest set. At admission, that rule affects exactly one population: claims created from now on. Meanwhile four hundred claims already exist, and several teams' data sits on them. Audit is the half of Gatekeeper that can tell you anything about those four hundred: after one interval, the Constraint's status carries a total and a sample of names. Nothing has moved, nothing has broken, and you now have a number to plan against — which is precisely the value, and precisely the limit, of a detective control.

  • Does audit still report on a Constraint whose enforcementAction is dryrun?
    Yes. `dryrun` tells the admission webhook not to deny; it says nothing to the audit controller. The sweep evaluates that Constraint like any other and its violations accumulate in status. That is exactly why teams ship a new rule in dryrun first — audit gives them the size of the existing problem while nothing is blocked.
  • How do you turn Gatekeeper's audit off, and why would anyone want to?
    Set `--audit-interval=0`. On a very large cluster the periodic read of every constrained kind is the heaviest thing Gatekeeper does, both for the API server and for the audit pod's memory. Teams under pressure lengthen the interval, narrow it with `--audit-match-kind-only`, or disable it and get their inventory picture from somewhere else.
  • If two Constraints match the same PersistentVolumeClaim, where do the violations show up?
    On each Constraint's own status, separately. Gatekeeper produces no cluster-wide findings object; the Constraint is the unit of reporting. To see everything a sweep found you enumerate the Constraints and read each status, which is one reason teams scrape the results rather than reading YAML by hand.

The admission webhook is the door check on the way in. Audit is the night-shift stocktake that walks the shelves and writes down what is already on them — it counts, it does not confiscate.

saying these in an interview costs you the question

  • Says audit deletes or evicts the offending objects
  • Thinks audit runs per write like the webhook
  • Expects a single cluster-wide findings object
  • Assumes a dryrun Constraint is skipped by audit
  • Treats the status list as a live view of the cluster

context

open as a page

What does Gatekeeper's gator test command evaluate, and what must you feed it?

level: juniorimportance: must knowfreq 60%

basics

~20 s

gator test evaluates Gatekeeper policy locally: you hand it ConstraintTemplates, their Constraints, and the Kubernetes objects to review, and it reports which object violated which constraint. It needs no cluster, no kubeconfig and no admission webhook.

open as a page

What does a Gatekeeper Assign resource do, and how does it differ from a Constraint?

level: juniorimportance: must knowfreq 64%

basics

~10 s

A Gatekeeper Assign is a mutator: at admission it writes a value into a field of the incoming object, so the object is stored changed instead of rejected. A Constraint only inspects and denies.

open as a page

In a Gatekeeper ConstraintTemplate, what must the Rego produce to deny a request?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Gatekeeper denies a request when the template's violation rule produces a non-empty set of objects, each carrying a msg string. An empty set means the constraint is satisfied. Each object may also carry an optional details object.

open as a page

In Gatekeeper, what do the Config and SyncSet objects do, and where does the synced data appear to a rule?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Gatekeeper's Config (spec.sync.syncOnly) and SyncSet objects choose which resource kinds are replicated into its in-memory cache. Those objects then appear to Rego under data.inventory, so a constraint can read cluster objects other than the one being admitted.

open as a page

In OPA Gatekeeper, what is the difference between a ConstraintTemplate and a Constraint?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A ConstraintTemplate defines a reusable rule plus a schema for its parameters, and makes Gatekeeper generate a new custom resource kind. A Constraint is an instance of that kind: it switches the rule on, scoped and parameterised.

open as a page

When Gatekeeper evaluates a violation rule, what is in input and what is not?

level: middleimportance: must knowfreq 67%

basics

~20 s

Two things only: input.review, holding the resource under decision as input.review.object (and its previous version as oldObject on an update), and input.parameters, holding the matched Constraint's parameters. No other object, no cluster state, no history.

open as a page

In Gatekeeper's gator verify, what can a Suite case assert beyond a plain denial?

level: middleimportance: should knowfreq 46%

basics

~20 s

A gator verify case can assert how many violations the object produced and what the violation message said, not merely that something was denied. That pins which rule fired and what a developer will be told when it does.

open as a page

Your Gatekeeper mutator must add ALL to every container's dropped capabilities — Assign or ModifySet?

level: middleimportance: should knowfreq 50%

basics

~20 s

ModifySet. Assign writes the whole value at the location, so it would replace whatever drop list the author already had. ModifySet treats the list as a set and merges the new member in, leaving the existing entries alone.

open as a page

How does a Gatekeeper rule read a sibling object from data.inventory, and what happens if that kind was never synced?

level: middleimportance: should knowfreq 48%

basics

~10 s

A rule indexes data.inventory by scope, groupVersion, kind and name. An unsynced kind makes that lookup undefined, not false — so the surrounding logic decides whether the rule blocks everything or waves everything through.

open as a page

What do Gatekeeper's enforcementAction values deny, dryrun and warn each do to a request?

level: middleimportance: should knowfreq 61%

basics

~20 s

deny rejects the request and returns the violation message. dryrun admits it and only records the violation on the Constraint's status. warn admits it too, but the API server returns a warning that the person applying sees immediately.

open as a page

A Gatekeeper Constraint's status lists 20 violating PVCs; why is that not the offender count?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Because the violations list in a Constraint's status is truncated, not complete: Gatekeeper stores at most --constraint-violations-limit entries per Constraint, 20 by default. The real count sits beside it in status.totalViolations, as of the last audit timestamp.

open as a page

A Gatekeeper constraint passes its gator suite but never fires on Deployments — how do you catch that offline?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Use gator expand with an ExpansionTemplate: it renders the Pod a Deployment, Job or CronJob would produce, and gator test and verify then review that generated Pod. The false green is a Pod rule tested only on bare Pods.

open as a page

Your Gatekeeper Assign is installed but the Pods come out unchanged — how do you debug it?

level: seniorimportance: should knowfreq 43%

basics

~20 s

A mutator that matches nothing fails silently, so work outward: confirm which object is actually admitted, then applyTo's group/version/kind, then the match block, then the location path and pathTests — and verify on a server-side dry-run of a real object.

open as a page

Why can't a Gatekeeper ConstraintTemplate call http.send, and what replaces it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Gatekeeper evaluates template Rego without http.send. A call during admission would put an outside service on every matching write's path and make results non-reproducible. Facts are pushed in instead: parameters, replicated objects, or an external data provider.

open as a page

A Gatekeeper constraint blocked a Deployment for a missing NetworkPolicy applied seconds earlier — why, and what do you do?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Replication lag. A Gatekeeper rule reads data.inventory, a watch-fed replica rather than a live query, so the NetworkPolicy existed in the API server but had not reached the cache when the decision was made.

open as a page

Your Gatekeeper Constraint is blocking system-namespace workloads — how do you scope or exempt it?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Narrow the Constraint's match block: list kinds, restrict namespaces or add excludedNamespaces, or select by label. For system namespaces that must never be evaluated at all, exempt them engine-wide in Gatekeeper's Config rather than per rule.

open as a page

In Gatekeeper, what changes when the audit controller runs with audit-from-cache enabled?

level: middleimportance: nice to knowfreq 34%

basics

~10 s

Gatekeeper's audit normally reads each constrained kind from the API server every cycle. With --audit-from-cache it evaluates Gatekeeper's replicated object cache instead: far cheaper, but unreplicated kinds are invisible and silently report zero violations.

open as a page

gator test in your rule repo's CI exits 0 on a manifest you know violates the constraint — why?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Usual causes: the constraint is set to dryrun or warn, so violations print without failing the exit code; the rule files were never in the input set; the match block selects nothing; or the CI step discards the status.

open as a page

How does the location path in a Gatekeeper Assign select a field inside a Pod?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

The location is a dotted path from the object's root, with list entries picked by a key filter such as spec.containers[name:*]. Missing intermediate fields are created on the way; a glob only matches list entries that already exist.

open as a page

How do you share Rego helpers across Gatekeeper ConstraintTemplates?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Through the ConstraintTemplate's libs field: each helper is Rego source shipped inside the same object, under the lib package namespace and imported as data.lib.<name>. A template is self-contained, so every one carries its own copy.

open as a page

In Gatekeeper, what does an external data Provider add to a rule, and what does a slow one do to admission?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

A Provider registers an HTTPS service Gatekeeper may call during evaluation, letting a rule resolve a fact no cluster object holds — an image tag to its digest. The call runs inside the admission request, so provider latency becomes request latency.

open as a page

In a Gatekeeper ConstraintTemplate, what belongs in the rule body versus in its parameters?

level: seniorimportance: nice to knowfreq 40%

basics

~20 s

Put the invariant in the template body and the values that vary between teams in parameters, declared with types in the openAPIV3Schema. Fork into a second template only when the logic differs, not when a number does.

open as a page