Kyverno background scanning is churning PolicyReport objects and loading etcd. How do you cut the cost?
answer
- cost is resources times rules times rescans
- every result is an API object
- rescans rewrite what already existed
- narrow the match before the interval
- admission-only rules need no background evaluation
basics
~20 sA background scan re-evaluates every stored object against every matching rule and writes results into PolicyReport objects that live in etcd. Cut the cost by narrowing what matches, scanning less often, and disabling background evaluation for admission-only rules.
solid answer
~50 sMeasure before tuning: count the report objects, look at API-server write rate by resource, and check etcd size. The load is roughly matching resources times rules times rescans, so the biggest wins are in the match block — narrow the kinds, use a namespace selector, exclude noisy namespaces, and match the controller object rather than every replica Pod, which collapses N results into one. Then set `background: false` on rules that only make sense against an admission request, which is required anyway for rules reading the requesting user. Lengthening the scan interval helps but costs freshness. Also filter noisy kinds and namespaces out before the controller sees them, and give the reports controller its own resources. If you need long retention or cross-cluster views, export findings out of the cluster rather than treating etcd as the results database.
code
yaml · 24 linesapiVersion: wgpolicyk8s.io/v1alpha2
kind: PolicyReport
metadata:
name: pol-disallow-deprecated-field-7f3c
namespace: payments
summary:
pass: 214
fail: 9
warn: 0
error: 0
skip: 3
results:
- source: kyverno
policy: disallow-deprecated-field
rule: check-deprecated-field
result: fail
severity: medium
message: "deprecated field in use; removed in a later API version"
resources:
- apiVersion: apps/v1
kind: Deployment
namespace: payments
name: ledger-api
# ... one entry per rule per matched resourcego deeper
Know that a background scan re-checks objects already stored in the cluster and records the outcome in report objects, and that those reports are ordinary Kubernetes resources kept in etcd like anything else.
Explain the cost model — matching resources times rules times rescans — and why matching Pods rather than their controllers multiplies both the write volume and the number of findings a human has to read.
Diagnose before tuning: report counts by namespace, API-server write rate, etcd size. Then order the levers correctly, match block first, and name what visibility each reduction costs you.
Own the standing decision about how much of the cluster is continuously evaluated and how long findings are kept in-cluster. That is a budget question about API-server capacity as much as a security one, and it belongs in a written platform standard.
## What background scanning actually does Admission only ever sees what someone submits. Objects that were created before a policy existed are never re-examined by it, and nobody resubmits a two-year-old Deployment. Background scanning closes that hole: a controller periodically lists the resources that policies match and evaluates them, writing the outcomes into `PolicyReport` objects (namespaced) and `ClusterPolicyReport` objects (for cluster-scoped resources). Nobody is blocked; the outcome is a record. That is exactly the population you care about for a rule like "this field is deprecated and disappears at the next API version". Nothing new is being created with it. The interesting instances are already stored, and only a scan will ever find them. A report is an ordinary API object with a `summary` of counts and a `results` list, one entry per rule per matched resource. Which is where the cost comes from. ## Where the load comes from Think of it as **matching resources x rules x rescans**, with four multipliers people miss: 1. **Match breadth.** A policy matching Pods in all namespaces matches every Pod in the cluster. Add ten such rules and every scan produces ten results per Pod. 2. **Replica fan-out.** One misconfigured Deployment with fifty replicas is fifty Pod-level results. Matching the controller instead gives you one — and one that a human can actually fix, since editing a Pod is pointless. 3. **Churn.** Namespaces full of short-lived objects — CI runs, Jobs, autoscaled workloads — generate a continuous stream of report creations and deletions on top of the scheduled passes. This is often the real source of a paged operator's write storm, not the interval. 4. **Rewrites.** Every rescan rewrites report content. Every write goes through the API server into etcd and fans out to every watcher of that resource type. Etcd cares about write volume and object size, not about whether the writes were interesting. A cluster whose reports are mostly unchanged from pass to pass still pays for the pass. ## Diagnosing it Before turning knobs, establish which of the multipliers is dominant: - Count report objects and their sizes, broken down by namespace. A handful of namespaces usually dominates. - Look at API-server request rate by resource and verb. If report writes dwarf everything else, you have confirmation. - Check etcd database size and how compaction and defragmentation are keeping up. - Look at the controller's own CPU and memory — a controller being throttled produces stale reports as well as load. The answer to "is it the interval or the match?" is almost always the match, and tuning the interval first is the classic wasted change. ## The levers, roughly in order of payoff **Narrow what matches.** Name specific kinds instead of wildcards. Use a namespace selector so the policy applies to the namespaces that matter. Exclude system namespaces you do not govern. Match the workload controller rather than the Pods it creates. **Turn background evaluation off where it makes no sense.** A policy has a background setting; when it is off, the rule decides only at admission and never appears in a scan. Any rule that reads the requesting user or the operation *must* be in that category, because a stored object carries no request — there is no user, no operation and no old object behind it. Rules guarding something meaningful only at creation time belong there too. **Filter at the door.** The Kyverno configuration carries resource filters that drop kinds and namespaces before the controllers process them at all. This is the blunt instrument for high-churn, low-value objects. **Lengthen the interval.** Real, but it buys less than the match changes and it costs freshness: the reports then describe the cluster as of the last pass, so anything driven off them lags by up to one interval. **Right-size the controller.** Give the reporting side its own resource budget so it is not competing with admission handling, whose latency is on the critical path of every request. **Stop using etcd as your results database.** If you want ninety days of history, trend lines, or a view across clusters, ship the findings to a store built for that and keep the in-cluster copy as a short-lived working set. ## The tradeoff to state out loud Every lever trades evidence for cost. A narrower match means you have no findings about what you excluded. A longer interval means your picture is older. Excluding a namespace means you cannot answer questions about it. That is fine as long as it is a decision rather than an accident — but say which population you are now blind to, because the next person to ask "does this rule hold across the cluster?" deserves an honest answer.
- Which rules cannot run in a background scan at all?Ones that depend on the admission request itself: the requesting username or groups, the operation, and the old object on an update. A scan reads a stored object, and there is no request behind it, so none of that data exists. Those rules must have background evaluation disabled and can only ever decide at admission.
- Why does matching Pods instead of Deployments multiply report volume?Each replica is its own object, so one bad workload produces one result per Pod, and every rollout replaces those Pods and rewrites the entries. Matching the controller gives one durable result per workload — cheaper, and also more useful, because the Pod is not the thing anyone can fix.
- What do you give up by lengthening the scan interval?Freshness. The reports then describe the cluster as of the last pass, so a resource fixed or broken since then is misrepresented, and anything you drive off the reports — dashboards, alerts, tickets — lags by up to one full interval. It is a cheap change but usually a smaller win than narrowing what matches.
- The operator wants to disable background scanning entirely. What is the cost?You lose all visibility into objects that already exist. Admission covers only what is submitted from now on, so anything created before the policy — or created while it was disabled — is invisible, and you can no longer answer whether the rule holds across the cluster. It is a defensible emergency measure, not a steady state.
saying these in an interview costs you the question
- Thinks reports live outside the cluster rather than in etcd
- Blames admission webhook latency for background scan load
- Raises the interval without narrowing what matches
- Scans every Pod when the controller object would do
- Assumes a background scan can read the requesting user
- Excludes namespaces without saying what visibility was lost