In Gatekeeper, what changes when the audit controller runs with audit-from-cache enabled?
answer
- changes where the objects come from
- API server versus Gatekeeper's own copy
- the copy is opt-in per kind
- a missing kind is silence, not an error
- no violations can mean not looked at
basics
~10 sGatekeeper's audit normally reads each constrained kind from the API server every cycle. With --audit-from-cache it evaluates Gatekeeper's replicated object cache instead: far cheaper, but unreplicated kinds are invisible and silently report zero violations.
solid answer
~50 sBy default each audit cycle asks the API server for the objects of every kind some Constraint matches. That is a live, authoritative read, and on a large cluster it is also the most expensive thing Gatekeeper does. `--audit-from-cache=true` switches the source: audit evaluates Gatekeeper's own replicated copy of cluster objects instead, so a sweep costs almost nothing at the API server. The catch is that the cache only holds kinds you have explicitly configured Gatekeeper to replicate. A Constraint over a kind that is not replicated does not error — it finds nothing, and its status reads like a clean estate. That silent false green is the failure mode interviewers are probing for. The results are also only as fresh as the informer feeding the cache, so you trade authority and completeness for cost.
go deeper
Know that Gatekeeper's audit normally reads objects straight from the Kubernetes API on each cycle, and that there is a setting to read a cached copy instead, mainly to save load.
Be able to state both sides of the trade: cheaper sweeps versus a cache that only holds kinds someone configured, so a Constraint over an unreplicated kind reports nothing at all.
Demonstrate that you would verify coverage after flipping the flag — compare constrained kinds against replicated kinds — and that you would tune interval, match-kind and chunk size before changing the data source.
Frame it as a question of which failure you would rather have: an audit that is expensive and honest, or one that is cheap and can go quietly blind. Decide who owns keeping the replicated kind list in step with the rule set.
## The default: a live read, every cycle With no special configuration, Gatekeeper's audit controller does the obvious thing on each interval: for every kind that some installed Constraint matches, it lists those objects from the Kubernetes API server and evaluates them. Two properties follow. It is **authoritative and complete**. Whatever exists in the cluster at that moment and matches a Constraint gets looked at. Nothing has to be configured in advance for a new kind to be covered — install a Constraint over a kind and the next sweep reads it. It is **expensive**. Listing hundreds of thousands of objects on a schedule is load on the API server and on etcd behind it, and the objects have to be held in the audit pod while they are evaluated. On a big cluster this is a real capacity conversation, not a rounding error. ## What the flag changes `--audit-from-cache=true` redirects the sweep's input. Instead of listing from the API server, audit evaluates Gatekeeper's replicated copy of cluster objects — the same cached data a rule can consult when it needs to see objects other than the one under review. The sweep now costs one pass over an in-memory dataset that is being kept warm anyway by watches, rather than a fresh set of list calls. The crucial consequence: **the cache is opt-in per kind**. Gatekeeper only replicates the kinds an operator has configured it to replicate. That configuration is a deliberate, separate act — you choose kinds and pay memory for them. ## The false green Put those two facts together. You enable `--audit-from-cache` for cost reasons. You have a Constraint requiring PersistentVolumeClaims to use an approved encrypted StorageClass, and four hundred non-compliant claims exist. If `PersistentVolumeClaim` is not among the replicated kinds, the sweep sees an empty population for that kind. It does not fail. It does not warn. It reports **no violations**, and the Constraint's status — the artefact people actually read — becomes indistinguishable from a clean cluster. This is the reason the flag deserves an interview question at all. Every other way of getting audit wrong is loud: a crashed pod, a stale timestamp, an error in the logs. This one produces a confident, well-formed, completely wrong answer. Whenever you enable cache-based audit you have to pair it with a check that every kind referenced by a Constraint is actually replicated, and re-check it whenever someone adds a Constraint over a new kind. ## Staleness, and what it is not Cache-backed audit results are as fresh as the watch feeding the cache, which in practice is close behind the cluster but not identical to it. That is a second-order concern next to the completeness problem: audit is already a snapshot bounded by its interval, so a little extra lag rarely changes a decision. The kind-coverage gap changes the answer entirely. A related trap: the cache also costs memory in Gatekeeper's own pods, since it holds the replicated objects. Cache-based audit does not make the objects free — it moves the cost from repeated API reads into steady-state memory. On a cluster where you replicate a handful of small kinds that is an excellent trade; on one where someone replicates every ConfigMap and Secret to make audit cheap, it is not. ## The other dials `--audit-from-cache` is not the only way to make sweeps affordable, and it is usually not the first one to reach for: | Dial | Effect | | --- | --- | | `--audit-interval` | Longer gap between sweeps; the cheapest change, at the price of a staler picture. Set to 0 to disable audit. | | `--audit-match-kind-only` | Only read kinds some Constraint actually matches, instead of the broader set. | | `--audit-chunk-size` | Page the list calls so the audit pod holds less at once and the API server sees smaller responses. | | `--audit-from-cache` | Change the source from the API server to the replicated cache. | ## How to answer in an interview Say what changes (the source of the objects), say what you buy (API-server load), and say what you risk (kinds outside the cache are silently absent, which reads as compliance). Then say what you would do: prefer interval and match-kind tuning first, and if you do turn on cache-based audit, treat "is every constrained kind replicated?" as a standing check rather than a one-time setup step.
- You enabled audit-from-cache and a Constraint's violations dropped from 412 to zero overnight. What is your first hypothesis?That the kind is not replicated, so the sweep evaluated an empty population. Nothing was remediated overnight and nothing errored — audit simply had nothing to look at. Confirm by checking which kinds Gatekeeper replicates against the kinds your Constraints match, and by comparing the count with a direct list of the objects.
- Which knob would you reach for first if audit is straining your API server?The interval, then `--audit-match-kind-only`, then chunking the reads. All three keep the authoritative live read and just do less of it, so they cannot produce a false clean result. Cache-based audit is the bigger change and carries the coverage risk, so it is a deliberate decision, not a default reflex.
The default sweep is a stock count done by walking the warehouse. Cache-based audit counts from the inventory system instead — instant, but it can only count the shelves someone remembered to register.
saying these in an interview costs you the question
- Thinks the flag makes audit results fresher
- Assumes an unreplicated kind produces an error
- Believes cache-based audit removes the memory cost
- Treats zero violations as proof of compliance
- Enables it cluster-wide without checking constrained kinds