skip to content

Kyverno reports show thousands of failing resources that no team has fixed in months. What do you do?

level: principalimportance: nice to knowfreq 28%

answer

  1. a count is not a work queue
  2. group by rule and by owner
  3. a few rules produce most findings
  4. no owner and no date, no action
  5. delete rules nobody will ever act on

basics

~20 s

Treat the count as a symptom, not a work queue. Break it down by rule and by owning team, delete or narrow rules nobody will ever act on, and tie what remains to a dated, per-team list.

solid answer

~50 s

A four-figure violation count is not a backlog, it is evidence that nothing is routed to anyone. Start by characterising the population: group by rule, by namespace and by age, and you will usually find a handful of rules and a handful of namespaces produce most of it. Then ask, per rule, what happens if this is never fixed. If the answer is that a cluster upgrade breaks the workload, you have a genuine deadline and a work item. If it is a real risk, it needs a named owner and an agreed date. If nobody will ever pay for it, delete the rule or narrow it — a finding that produces only noise costs scan time, report storage and, worst of all, credibility. Finally, send each team its own dated list rather than the global dashboard, and track the age of the oldest open finding, not the total.

go deeper

for a junior

Understand that a recorded violation is not a fixed violation, and that a finding naming a resource but no person cannot become work. Being able to say who owns a namespace is more useful here than knowing the rule syntax.

for a middle

Be ready to slice the data — by rule, namespace, kind and age — and to explain why matching Pods instead of controllers inflates a count without adding information.

for a senior

Show the per-rule reasoning: what happens if this is never fixed, who owns it, what date it hangs on. Be willing to recommend deleting a rule and to defend that as strengthening the remaining controls.

for a principal

Own the standard that every reporting-only rule carries a decision: a date, an owner, or a deletion. Argue the credibility cost of a permanently red surface in terms leadership recognises, and replace the total count with an indicator that can actually improve.

## Why an unread report is worse than no report A report that nobody acts on has three costs and no benefit. It costs machine resources: every entry was produced by an evaluation and is stored as part of an API object that gets rewritten on every pass. It costs credibility: once people learn that the reporting surface is always red, they stop distinguishing the entry that matters from the four thousand that do not — including the next one you actually need them to read. And it is frequently claimed as coverage. "We have a policy for that" is true and useless when the policy has been failing silently for eight months. That gap between the control on paper and the control in effect is precisely what a serious reviewer probes. So the first move is not to attack the number. It is to accept that the number is telling you the reporting pipeline ends nowhere. ## Characterise before you triage A total is the least useful view of the data. Break it down: - **By rule.** Nearly always a Pareto: two or three rules produce most of the volume. - **By namespace or team.** Same shape, and this is the axis that determines who can do anything about it. - **By kind.** A rule matching Pods rather than their controllers inflates the count without adding information, and collapsing that alone can remove a zero from the total. - **By age.** A finding that appeared last week is a regression; one that has been there since the rule was written is a migration. That breakdown usually turns "thousands of violations" into something like "three rules, four teams, ninety workloads" — a sentence somebody can act on. ## Decide, per rule, what it is for For each rule that survives the breakdown, ask the only question that matters: **if this is never fixed, what actually happens?** There are three honest answers. **It breaks.** The clearest case, and the easiest to drive. A rule flagging a field that a coming API version removes has a date attached by someone other than you: the upgrade. That is not a policy ask, it is a compatibility work item with a deadline, and it should be handed over as such — "these nine workloads stop reconciling after the upgrade" moves far more than "you have nine policy violations". **It is a real risk with no deadline.** Then it needs the thing it currently lacks: a named owner and an agreed date. Without both it will look identical in six months. **Nothing happens.** Somebody once thought this was a good idea and no one will ever pay to fix it. Then delete the rule, or narrow it to the population where it does matter. This is the recommendation people find hardest to make, because removing a check feels like weakening security. It is the opposite: a rule that produces only ignored findings provides no control while consuming resources and eroding attention. Deleting it makes the remaining findings mean something. ## Route it, and make the ask finite A finding that names a resource but not a person is not actionable. Map namespaces to owning teams from whatever inventory you already have, and give each team only its own list. The global dashboard is for you, not for them — a team shown four thousand organisation-wide findings will correctly conclude that none of them are theirs. Then make the ask small and dated. "Nine Deployments, before the upgrade next quarter, here is the field and here is the replacement" is a task. "3,412 policy violations" is weather. Where you can, do the work rather than delegate it. If a fix is mechanical, offering the patch is often cheaper than the meetings needed to persuade four teams to write it themselves. ## Measure something that can improve Stop reporting the total. Better indicators: open findings per rule that has a committed date; the age of the oldest open finding; the trend per team; and how many rules are still in the reporting-only state with no decision attached. A total conflates a cosmetic rule with the one that breaks production, and it can rise because you added a good rule — which then punishes you for improving coverage. ## Say the honest thing about what reporting can and cannot do Be clear-eyed about the limits. Findings about stored objects only disappear when someone changes those objects; nothing about the reporting side edits a running workload for you. And a finding you cannot route is not a control at all — it is a note to yourself. The point of this exercise is to end up with fewer rules, each with an owner and a date, and a report that is worth opening.

  • What is the argument for deleting a rule rather than leaving it reporting quietly?
    It consumes evaluation time, report storage and attention while providing no control, and its permanent redness teaches people to ignore the whole surface — including the finding you need them to read next. Delete it, or narrow it to the population where the answer to "what happens if this is never fixed" is not "nothing".
  • How do you decide which failing rules become work items first?
    By consequence and deadline, not by count. A rule whose violations break at a known upgrade already has a date supplied by someone else, so it goes first. A genuine risk needs an owner and an agreed date. Everything else is a preference. Ranking by volume simply promotes whichever rule matched the most objects.
  • What would you show leadership instead of the total violation count?
    Open findings per rule that has a committed date, the age of the oldest open finding, and the trend per team. The total conflates a cosmetic rule with one that breaks production, and it rises when you add a good rule — an indicator that punishes you for improving coverage is the wrong indicator.
  • A team says the finding is intentional for their workload. How does that change the picture?
    Then it is a decision, and a decision that lives only as a permanent red line in a report is not recorded anywhere a reviewer can find it. Either the rule is wrong for that population and should be narrowed, or the acceptance needs to be written down deliberately with an owner and a review date rather than left as silent noise.

It is a smoke alarm that has been beeping for a year. The problem is not the beeping; it is that everyone in the building has learned to stop hearing it.

saying these in an interview costs you the question

  • Reports the raw violation count as if it were the risk
  • Sends every team the organisation-wide dashboard
  • Keeps a noisy rule because deleting one looks like weakening security
  • Assumes turning on enforcement repairs objects already stored
  • Counts installed policies as a measure of coverage
  • Opens a ticket per finding with no owner or date

context