skip to content

Your policy has run in audit mode for nine months with 400 open findings - what do you do?

level: principalimportance: nice to knowfreq 38%

answer

  1. a report with no reader
  2. ask what any row ever caused
  3. name the consumer and the decision
  4. the count is a price, not a backlog
  5. the default option is leaving it

basics

~20 s

Nine months of unread audit output is a log, not a control. Establish what any row ever caused, treat the 400 as the price of enforcing the rule, then give the report an owner, narrow its scope, or remove it.

solid answer

~50 s

First establish the honest position: ask what any row in that report has ever caused. If no namespace was fixed and no decision was made because of it, the rule has been producing evidence of its own existence and is not a control - it is a scheduled report. Then separate what the record does and does not prove. It proves the rule ran and is stable, and it prices the rule: 400 is the remediation bill, not a backlog. The exits are all uncomfortable and you should name them. Attach an owner and a date so somebody decides. Narrow the scope to where the violation genuinely matters, turning 400 into a number a team can carry. Or accept it will never be acted on and remove it. Refuse the default - leaving it as it is and continuing to claim coverage.

go deeper

for a junior

Understand that a record-only rule changes nothing by itself. Findings sit in a report until a person reads them, so a large open count means the work was never started, not that it is under way.

for a middle

Be able to say what the record supports: the rule ran, and this many resources violate it. Be equally clear that it does not show any reduction in risk, since no resource changed as a result.

for a senior

Show you would establish the consumer and the decision before anything else, then present the count as the cost of remediation and propose a scope that a team can realistically own.

for a principal

Own the organisational call: fund an owner, shrink the scope, or delete the rule and say so out loud. Be able to explain why an ignored finding stream is a liability that degrades every other rule the engine emits.

## Audit as a permanent resting state Record-only mode is chosen because it is free. It costs no engineer a blocked change today, it needs nobody's agreement, and it can be turned on in an afternoon. Precisely because it is free, nothing ever forces the next decision. Nine months later the rule is still recording, the finding count is a round number nobody can source, and the platform team is being asked what it is for. This is the most common end state for a policy rule, and recognising it is the point of the question. ## Start with the only question that matters **What changed because of a row in this report?** Not "is the rule correct", not "is the count going down" - what *happened*. Concretely: has any namespace acquired a default-deny NetworkPolicy because it appeared in this report? Has any team seen it? Has any decision - to accept, to remediate, to exempt - been recorded against any row? If the answer is none, say so plainly. A finding stream with no consumer and no decision attached is not a control in any meaningful sense; it is a logging job wearing a control's name. That distinction is worth being blunt about, because the organisation is very likely carrying it on a list of things that protect it. ## What nine months of audit output does prove Be fair to the record; it is not worthless. It supports three real claims: - **The rule runs and is stable.** It has evaluated continuously without breaking anything, which is a genuine input to any later decision about making it bite. - **It prices the rule.** Four hundred rows is what enforcing this rule would cost somebody: 400 namespaces that need a NetworkPolicy written, reviewed and applied by whoever owns them. Read as a *price tag*, the number is the most useful thing in the report. - **It describes the population.** Whether the count is growing, flat or shrinking tells you something. Growing means new namespaces are still being created non-compliant. Flat over nine months means nothing is remediating and nothing new is violating either - or that the sweep's scope has quietly stopped matching reality, which is worth checking before you conclude anything. And three claims it does **not** support: that the violations are accepted risk (nobody decided anything), that the risk went down (nothing changed), and that the rule is accurate (nobody looked at the rows closely enough to disagree with them). ## The exits, and who has to agree Every honest way out costs somebody something, which is why the situation persists: 1. **Give it an owner and a date.** The report gets a named reviewer, a cadence, and an expectation that each row ends in an action or a recorded decision. This is the option that converts it into a control, and it is real work for a person who currently has none of it allocated. If you cannot find that person, you have learned that the organisation does not value this rule as much as its presence on a list suggests. 2. **Narrow the scope.** Four hundred is unownable; forty might not be. If the rule genuinely matters most for namespaces handling regulated data or exposed traffic, scoping it there turns an ignorable number into a tractable one - and it makes the remaining rows mean something, because each one is now a namespace somebody agreed should be compliant. 3. **Remove it.** If nobody will own it and nobody will scope it, deleting the rule is the honest act. This feels like losing, and it is the option people resist hardest, but a permanently-ignored finding stream trains the whole organisation that the engine's output is noise. That cost is not paid by this rule; it is paid by the next rule that actually matters and is also ignored. 4. **Change how hard it bites.** Turning a rule from recording into something that stops changes is a real option, but it is a separate decision with its own rollout, its own negotiation and its own timing - and note that it does nothing about the 400, because a gate judges new requests, not resources that already exist. The fifth option is the one that happens by default: leave it. Name it explicitly when you answer, and say why you are refusing it. ## What to say to leadership The strongest framing is to stop reporting the rule as a control and start reporting it as a measurement. "We have measured the gap: 400 namespaces lack a default-deny NetworkPolicy. That measurement has been stable for nine months and nobody is assigned to close it. Here is what closing it costs, here is a smaller scope we could close instead, and here is what I recommend we stop claiming in the meantime." That converts an embarrassing backlog into a decision someone with a budget can make - which is the only thing that was ever missing. ## The habit worth carrying When a rule is deployed in record-only mode, write down at that moment who reads the output, on what cadence, and what a row obliges them to do. If those three answers do not exist on day one, they will not exist on day two hundred and seventy, and the report you are about to start generating already has no reader.

  • What does nine months of clean, continuous audit output actually prove?
    That the rule ran without breaking anything, and how large the violating population is. It proves nothing about whether risk fell, because nothing was changed by it, and nothing about whether the findings are correct, because nobody examined them closely enough to disagree.
  • How do you tell whether a record-only rule is a control or just a logging job?
    Name its consumer and its decision. Who reads the output, on what cadence, and what action does a row oblige them to take? If you cannot answer all three with a person and a timeframe, it is a scheduled report. Write those three answers down on the day you deploy the rule, not later.
  • The count has been flat at 400 for months. What do you read into that?
    Either nothing is remediating and nothing new is violating, or the evaluation's scope has drifted and it is no longer looking at everything it should. Check whether the population itself is changing - new namespaces created in that period should show up somewhere - before drawing any conclusion about behaviour.
  • The team says removing the rule looks like giving up on the control. How do you answer?
    That the control does not currently exist; only the report does. Keeping an unowned finding stream trains everyone to ignore the engine, and that cost lands on the rules that do matter. Either fund an owner, scope it to something a team can carry, or record the decision to accept the risk explicitly.

saying these in an interview costs you the question

  • Treats the audit report as proof the risk is managed
  • Says just switch it on without pricing the 400
  • Counts findings generated as work completed
  • Reads no complaints as agreement with the rule
  • Keeps it recording indefinitely to avoid the argument

context