skip to content

One rule accounts for most of your break-glass overrides this quarter - what do you conclude?

level: principalimportance: nice to knowfreq 36%

answer

  1. count against the rule's own denials
  2. a defect report on the rule
  3. segment by rule and by team
  4. somebody must fund the narrower rule
  5. silence is not proof of health

basics

~20 s

Read a concentrated override rate as a defect report on the rule, not on the teams. A rule routinely waved through is too broad, badly placed, or blocking work the organisation has already decided to accept.

solid answer

~50 s

I would split the overrides by rule, team and stated reason before concluding anything, and I would measure the rate against the number of times that rule fired, not against total deploys. Concentration on one rule usually means the rule encodes an intent it cannot express precisely - an egress rule denying any unfamiliar destination catches vendor failover alongside the thing it was written for - so every legitimate change has to buy its way through. The options are to narrow the rule until the legitimate shape passes, move it to a place where the distinguishing information is actually available, or accept it as advisory for that class of change. What I will not do is leave a rule blocking when most of its denials get overridden: it teaches everyone that a denial is a formality, and it spends the alarm value the break-glass path depends on.

go deeper

for a junior

Know that overrides are counted and reviewed rather than just filed, and that a rule people keep bypassing is a signal worth raising rather than a routine annoyance.

for a middle

Be ready to explain why the denominator is that rule's own denials rather than total deploys, and what a rate concentrated on a single rule most likely points at.

for a senior

Show how you would segment the data by rule, team and stated reason, and name the concrete change you would make - narrow the rule, move where it evaluates, or make it advisory for a class of change.

for a principal

Own the conclusion and its bill. Say who funds the ongoing work a narrower rule creates, what the interim position is while that work happens, and how you would defend either leaving the rule blocking or standing it down.

## Why override rate is the most informative number a gate produces A gate emits denials and passes. Neither tells you whether the gate is right - a denial only says the rule matched. An override says something much stronger: a human with authority looked at a specific denial and decided the rule was wrong about this change, or at least less important than shipping. That is a judgement, recorded, at scale. Treated as a stream of defect reports against your policy set, it is the closest thing a policy programme has to user feedback. ## Get the denominator right The usual mistake is overrides per deploy, which mixes every rule together and hides the interesting rule inside a large, healthy total. The number that means something is **overrides as a share of that rule's denials**. A rule that denied twice this quarter and was overridden once is a curiosity. A rule that denied four hundred times and was overridden two hundred is not a control; it is a speed bump with a recording device attached. Segment before you interpret: - **One rule, many teams.** The rule is the problem. Independent teams with different workloads reaching the same conclusion is about as clean a signal as this domain offers. - **Many rules, one team.** Look at that team's workload before you look at their discipline. They may own the one service that legitimately talks to the outside world - or they may be routing around everything, which is a different conversation with a different owner. - **One rule, one team, repeatedly.** Usually a single unmodelled legitimate case. Often the cheapest thing in the whole list to fix. ## The three explanations, and how to tell them apart **The rule is too broad.** It was written to catch a risk and it also catches a common legitimate shape, because the distinction was hard to express when it was authored. You can identify this by reading the stated reasons: they will all describe the same legitimate case in slightly different words. **The rule is asked a question it cannot answer where it sits.** The information that separates good from bad is not present at the point of evaluation - whether an external destination is a sanctioned vendor may be known to the service owner and to nobody at the gate. Rules like this cannot be tightened into correctness; they must move, acquire the missing input, or stop blocking. **The organisation has already decided to accept the behaviour.** Every override is granted, quickly, by people who agree it is fine. The rule expresses an intention nobody currently holds. That is a real finding, and the honest response is to change the rule to say what the organisation actually wants rather than leaving a control that is contradicted every week. ## The cost of leaving it alone A blocking rule that is usually overridden does three kinds of damage. It trains everyone that a denial is an administrative step rather than information, which bleeds into how they read every other rule. It consumes the break-glass path's alarm value - if the pull happens constantly, nobody looks at any individual pull, including the one that mattered. And it is very hard to describe as an enforced control to anyone who asks how it works, because the honest description is "it blocks, and then we let it through". ## The organisational layer, which is the actual principal question "Fix the rule" is easy to say and usually costs somebody who is not in the room. Narrowing an egress rule so that legitimate destinations pass generally means somebody must maintain the list of legitimate destinations, review additions, and answer questions about it - continuing work, on a team that did not plan for it. If that owner is not named and funded, the decision quietly becomes "keep overriding", which is the status quo with extra meetings. So the call has three parts, and a lead owns all of them: what the rule should become, who carries the ongoing cost of it being narrower, and what the interim position is while that work is done. An explicit interim - the rule is advisory for this class of change until the destination list has an owner - is far better than an implicit one where the rule blocks and everyone knows the override is automatic. The first is a decision; the second is a decision that nobody has to admit to making. ## The reverse signal A rule with no denials and no overrides is not evidence of health. It may be perfectly targeted, or it may not be matching anything at all - a selector that never applies, a surface that stopped being produced, a check that silently stopped running. Silence needs the same interrogation as noise: confirm the rule still fires and still denies a case you construct on purpose, before you count it as a working control. ## What to bring to the review Rate per rule with its denominator, the segmentation, the stated reasons in the overriders' own words, the trend after any change you made, and the list of overrides whose follow-ups were never closed. Bring the numbers rather than the anecdote - the anecdote is always about the one dramatic night, and the decision should be about the two hundred ordinary ones.

  • How do you tell a rule that is too broad from a team that is undisciplined?
    Segment. One rule overridden by many independent teams is a rule defect - unrelated people do not converge on the same wrong conclusion. Many different rules overridden by one team points at that team or its workload, and even then I would ask what is unusual about the work before asking what is wrong with the people doing it.
  • What is the risk of concluding that the rule is fine and the teams need training?
    Nothing changes, so the override becomes a ritual with a form attached. The real cost is the alarm: once pulling the handle is routine, nobody reads any individual pull, and the one that should have started a conversation looks exactly like the two hundred that should not have.
  • A rule recorded no denials and no overrides all quarter. Is it healthy?
    Unknown. It may be well targeted, or its selector may match nothing, or the check may have quietly stopped running. Before counting it as a working control, construct a case that should be denied and confirm it is. Silence is an absence of evidence, not evidence that the rule is doing its job.

saying these in an interview costs you the question

  • Reads a high override rate as a developer training problem
  • Measures overrides per deploy instead of per denial
  • Calls a routinely overridden rule fully enforced
  • Refuses to change a rule because teams complained
  • Treats zero overrides as proof a rule works
  • Decides to narrow the rule without naming who maintains it

context