skip to content

A Gatekeeper Constraint's status lists 20 violating PVCs; why is that not the offender count?

level: seniorimportance: should knowfreq 48%

answer

  1. the list is capped, the count is not
  2. twenty is a default, not a finding
  3. look next to the list in the same status
  4. check when the sweep last ran
  5. totalViolations versus violations length

basics

~20 s

Because the violations list in a Constraint's status is truncated, not complete: Gatekeeper stores at most --constraint-violations-limit entries per Constraint, 20 by default. The real count sits beside it in status.totalViolations, as of the last audit timestamp.

solid answer

~50 s

`status.violations` is a sample, not an inventory. The audit controller keeps at most `--constraint-violations-limit` entries per Constraint — 20 by default — so on a rule with hundreds of hits you are reading the first page. The number you want is `status.totalViolations` in the same status block. Before reporting it, check three more things. `auditTimestamp`, because audit refreshes only on its interval and a slow, failed or disabled sweep leaves old numbers that look exactly like fresh ones. Coverage, because exempt namespaces, a narrow match block, or cache-based audit over an unreplicated kind can put part of the estate outside the sweep entirely. And `enforcementAction`, because a `dryrun` Constraint is recording these claims, not stopping them — the population is still growing while you read it. If you need the identities of all 412 claims, raise the limit for one investigation or list the claims yourself.

code

yaml · 15 lines
yaml
status:
  auditTimestamp: "2026-08-27T02:14:09Z"
  totalViolations: 412
  violations:
  - enforcementAction: dryrun
    kind: PersistentVolumeClaim
    name: data-payments-0
    namespace: payments
    message: 'storageClassName "standard" is not in the approved encrypted set'
  - enforcementAction: dryrun
    kind: PersistentVolumeClaim
    name: data-ledger-3
    namespace: ledger
    message: 'storageClassName "standard" is not in the approved encrypted set'
  # ... 18 more entries, then the list simply stops

go deeper

for a junior

Remember that the violations shown on a Gatekeeper Constraint are a capped list, not everything found, and that a separate total is recorded alongside them.

for a middle

Be able to name --constraint-violations-limit, its default of 20, and why the cap exists at all — status lives on the object in etcd and is watched by every client.

for a senior

Show the reflex of checking the denominator and the timestamp before acting on a sweep, and know the sane ways to get a full list: a temporary limit raise or a direct query on the property the rule checks.

for a principal

Own how a number leaves your team. State the scope, the as-of time and the enforcement mode with any figure you hand to leadership or a customer, and decide whether audit status is a good enough measurement surface for the estate at all.

## The list is a sample; the count is elsewhere When Gatekeeper's audit controller finishes a sweep it writes results onto each Constraint's own `status`. The part people read is `violations` — a list of offending objects with kind, name, namespace and the rule's message. It is capped. `--constraint-violations-limit` bounds how many entries are stored per Constraint, and its default is **20**. The cap exists for a good reason. Status is part of the Constraint object, stored in etcd and shipped to every watcher. A rule that matches four hundred PersistentVolumeClaims would otherwise write four hundred messages into one object on every sweep, repeatedly, for every noisy rule you have. The count you actually want is right there next to the list: `totalViolations`. A status that says `totalViolations: 412` with twenty entries beneath it is not inconsistent — it is the design. Reading the list's length as the population is the classic error, and it is the one that makes a security lead tell leadership that twenty claims are unencrypted when four hundred are. ```yaml status: auditTimestamp: "2026-08-27T02:14:09Z" totalViolations: 412 violations: - enforcementAction: dryrun kind: PersistentVolumeClaim name: data-payments-0 namespace: payments message: 'storageClassName "standard" is not in the approved encrypted set' # ... 19 more, then the list simply stops ``` ## Four more ways the number can still be wrong Even `totalViolations` is not a live population count. Interrogate it: **It is stale by design.** Audit refreshes on its interval. `auditTimestamp` tells you as of when. If the audit pod is unhealthy, a sweep errored partway, or someone set `--audit-interval=0`, the status keeps whatever it last wrote — an old number is visually identical to a current one apart from that timestamp. A status showing zero violations deserves the same suspicion: check that a sweep actually ran recently. **Part of the estate may be outside the sweep.** Namespaces exempted from Gatekeeper are not audited. If audit is reading a replicated cache rather than the live API, kinds nobody replicated contribute nothing. And the Constraint's own match block is a filter — if it selects a label or namespace set narrower than you assume, the objects you care about were never in scope. **Nothing was blocked.** If `enforcementAction` is `dryrun`, admission is still accepting new non-compliant claims while you read the status. The number is not a shrinking backlog; it is a moving one. Distinguishing "412 legacy claims to migrate" from "412 and rising" is the difference between a project plan and a leak. **A capped list biases what you see.** The twenty entries you got are not a random sample and should not be treated as representative of which teams or namespaces are affected. If you need to know where the four hundred live, you need all of them, not the page audit chose to keep. ## Getting the full list, carefully Two honest options. Raise `--constraint-violations-limit` deliberately for one investigation, read what you need, and put it back — leaving it at several thousand across many noisy Constraints bloats objects in etcd, makes every watcher carry the payload, and can push a Constraint toward the object size ceiling. Or bypass status entirely and enumerate the objects yourself with a direct query on the property the rule checks, which for the StorageClass rule is a one-line list of claims and their `storageClassName`. The second scales better and is what most teams end up doing, with audit's total as the cross-check. ## Reading a sweep as a population estimate The question behind the question is what you are entitled to conclude from one status block. A defensible statement sounds like: *as of 02:14 today, the audit sweep found 412 PersistentVolumeClaims outside the approved encrypted StorageClass set, across the namespaces audit covers; the rule is in dryrun so new ones are still being created; the twenty named in status are a truncated sample and the full list came from a direct query.* Every clause in that sentence is a hedge you can defend. The indefensible version is "we have twenty unencrypted volumes" — a number that is wrong by a factor of twenty, presented as a fact, sourced from a field that never claimed to be the total. The mechanical knowledge here (a cap, a default of 20, a `totalViolations` field) is small. The habit it protects — never quote a number out of a truncated view without checking the denominator and the timestamp — is the thing being tested.

  • Why not set the violations limit to 5000 and stop worrying about truncation?
    Because status is stored on the Constraint object in etcd and pushed to every watcher. Thousands of messages per Constraint, rewritten each sweep across many noisy rules, bloats objects, adds watch traffic and can approach the object size ceiling. Raise it for one investigation, read what you need, and put it back.
  • totalViolations says 412, but a direct query finds 430 non-compliant claims. What explains the gap?
    Audit's number is from its last sweep, so claims created since are missing. Beyond staleness: exempt namespaces, a match block narrower than your query, kinds absent from the replicated cache if audit reads from cache, or a sweep that errored partway. Reconcile by comparing scope and timestamp before assuming either number is wrong.
  • A Constraint's status shows no violations at all. What do you check before calling the estate clean?
    That audit is running — the interval is not zero and the audit pod is healthy — and that `auditTimestamp` is recent. Then that the Constraint's match block actually selects the objects you mean, that the relevant namespaces are not exempt, and that the kind is covered by whatever source audit reads. Empty and unexamined look the same.

The violations list is the first page of search results. The number of hits is printed separately, and nobody would claim there were twenty matches just because twenty links fit on the page.

saying these in an interview costs you the question

  • Reads the length of status.violations as the population
  • Quotes the number without checking auditTimestamp
  • Treats an empty violations list as proof of compliance
  • Raises the violations limit cluster-wide and leaves it there
  • Assumes dryrun findings are a fixed backlog, not a growing one
  • Treats the truncated sample as representative of affected teams

context