skip to content

Half your org's blocking policy rules are now warn-only and nobody decided that. How do you fix it?

level: principalimportance: nice to knowfreq 33%

answer

  1. nobody meets to stop enforcing
  2. downgrade is cheap, reversal is not
  3. the repo shows intent, not state
  4. owner, reason, expiry on every downgrade
  5. findings with no consumer are logs

basics

~20 s

Give de-enforcement the same ceremony as enforcement: every downgrade carries an owner, a reason and an expiry, and each rule's live enforcement state is inventoried. Findings routed to a channel with no consumer are not a control.

solid answer

~50 s

Drift happens because downgrading is a cheap local action taken under pressure while reversing it is expensive work nobody is paid for, so I attack that asymmetry. First, visibility: an inventory built from the gates themselves rather than the policy repository, listing each rule, its mode per environment, who changed it, when, and how long since it last denied anything. Second, expiry: a downgrade is a dated decision, and past its date it escalates to a named owner or reverts. Third, leading indicators, because retry rate, timeout rate and waiver volume all climb before a rule gets downgraded. And I would be blunt about the warn tier: a rule whose findings land in an unowned channel enforces nothing while still counting as coverage. Give it a consumer, fix it so it can block, or retire it.

go deeper

for a junior

Know that a rule set to warn stops changing outcomes, and that where its findings are routed decides whether anyone ever acts on them.

for a middle

Be ready to describe how a downgrade should be recorded, with owner, reason and expiry, and why the policy source alone never tells you what is actually enforced today.

for a senior

Show how you detect drift: enforcement mode per rule per environment, time since a rule last denied anything, and the retry and timeout rates that precede a downgrade.

for a principal

Own the organisational answer: who may de-enforce and for how long, who funds the fix, and the willingness to retire a rule rather than carry it as coverage that is not real.

## Why gates drift rather than get switched off Nobody holds a meeting to stop enforcing a control. The sequence is always the same and always local: 1. A rule becomes slow, noisy or intermittently wrong. 2. It blocks people at a bad moment, so someone with access sets it to warn to unblock a release. 3. The pain disappears, because nothing is failing any more. 4. Restoring it requires fixing the underlying problem, which is real work with no visible payoff, and it is now nobody's priority. The result is a control that exists in the policy repository, appears in coverage reports, is described in onboarding docs, and changes nothing. The asymmetry between how cheap the downgrade is and how expensive the reversal is, is the actual mechanism, so any fix has to change that asymmetry rather than exhort people to be more careful. ## The visibility problem: the repository is not the state The first mistake is asking the policy source what is enforced. Source shows intent. Enforcement lives in the running gates, in per-pipeline configuration, in environment-specific overrides, in a flag someone set, in an exclusion list that grew. So build the inventory from the enforcing side and publish it: | Field | Why it matters | |---|---| | Rule, environment, current mode | The actual state, per place it runs | | Who changed the mode, and when | Makes the change a decision with an author | | Age of the current mode | Turns quiet drift into a visible number | | Time since the rule last denied anything | Distinguishes satisfied from not matching | That last row deserves care, because a rule that has denied nothing in six months is ambiguous. It may be genuinely satisfied, everyone complies now, which is the goal. Or it may have stopped matching because the surface it inspects changed shape. Or it may be advisory in the places that matter. You cannot tell those apart from the outside, which is exactly why the metric is worth publishing next to the mode. ## Making de-enforcement a dated decision Every downgrade should carry three things: a named owner, a reason, and an expiry. The expiry does the work. It converts an open-ended state into a scheduled question, and it forces the honest conversation at a moment when the incident pressure has passed. Past the expiry, the state escalates to the owner and their lead, or reverts automatically if the environment can tolerate that. A permanent downgrade should still be possible, but it should be a signed decision by whoever carries the risk, not an accumulation of forgotten temporary ones. Making people sign for it is usually enough to get one of the better outcomes instead: fund the fix, narrow the rule to the population that justifies its cost, or retire it deliberately. ## Leading indicators, so you intervene early By the time a rule is downgraded, the argument is already lost. The measurable precursors run weeks ahead: - **Retry rate per rule.** People re-running until green is the first workaround, and it is visible. - **Timeout and error rate per rule.** A rule that increasingly cannot answer is a rule that will soon be turned off. - **Waiver and exception volume.** A rising count means the rule is wrong for a population it keeps hitting, which is a scoping problem, not a compliance problem. A platform team that watches these can offer to narrow or fix a rule before anyone has to demand it be disabled, and that offer is usually accepted, because teams want the block to stop, not the control to die. ## Be honest about the warn tier Warn-only is a legitimate state when it has a consumer. It is a fiction when it does not. A rule producing findings into a channel with no owner has three costs and no benefit: it consumes evaluation time, it produces noise that trains people to ignore the channel, and it is counted as coverage in reports that then misrepresent the organisation's posture. For every rule in that state the choice is small and stark: give it a named consumer with a review cadence, fix it so it can block again, or delete it. Carrying it as false coverage is the worst of the three, because it is the only one that also misleads. ## What good looks like a quarter later Enforcement state is a published number, per rule and per environment, with an age. Every non-enforcing state has a person's name and a date on it. The count of rules past expiry is reviewed like any other operational backlog. And the organisation can answer, without an archaeology project, the question that started this: which of our rules are actually stopping anything today?

  • How do you find the rules that have quietly stopped enforcing?
    Query the gates rather than the policy repository: per rule and per environment, what mode is configured, who set it, and when did it last deny anything. A rule marked enforcing with zero denials for six months is either genuinely satisfied or no longer matching, and those look identical in source. Pair that inventory with the retry and timeout rates that typically precede a downgrade.
  • A team says the rule was downgraded because it was too slow, and that has not changed. Now what?
    Then the honest options are to fund the fix, narrow the rule to the cases that justify its cost, or retire it. Leaving it in warn indefinitely is not a fourth option, it is the absence of a decision. Put the choice in front of whoever carries the risk with the cost attached, and require a signature for permanent de-enforcement, which usually produces one of the first three.
  • Do warn-mode findings need an owner if they block nothing?
    Yes, or they are logs. Each warn-mode rule should have a named consumer, a review cadence and a date when its state is revisited. Without that, the findings train people to ignore the channel they land in, and the rule still gets counted as coverage in reports, which makes the organisation's picture of itself wrong in the flattering direction.

A smoke alarm with the battery taken out still hangs on the ceiling, and an inspection that only records alarm present will pass the building every time.

saying these in an interview costs you the question

  • Treats warn mode as a safe indefinite middle ground
  • Counts warn-only rules as coverage in reports
  • Reads the policy repository as the enforcement state
  • Re-enables every downgraded rule at once with no owners
  • Routes findings to a channel nobody is accountable for
  • Assumes zero denials proves everyone is compliant

context