Every firewall change passes per-ticket diff review, yet a shadowed deny let an intruder through — what does the diff never show, and what would catching it cost?
answer
- Correctness is not a property of one rule
- The diff has one rule; the defect needs two
- Position, not content
- Gate on the delta, not the backlog
basics
~20 sA diff shows one rule; reachability is a property of that rule's position against every other rule. Nothing in the ticket contains the rest of the ordered policy, so catching it means re-analysing the whole policy on every change.
solid answer
~60 sThe reviewer was not careless — the review's scope is wrong for the defect. A per-ticket diff contains one rule, and the rule is genuinely correct: right source, right destination, right port, a business justification, an approver. The defect lives in the relation between that rule and rules it was never shown, possibly hundreds of lines below. To see it a human would have to hold several thousand ordered rules in their head and compute, for the new rule, which existing rules it now covers and which now cover it. Nobody does that, and a second reviewer does not help because they have the same scope. The fix is to change what is reviewed, not who reviews it: run whole-policy analysis against the proposed policy before it is pushed, and report only anomalies this change introduces against the current baseline. That costs a policy export and an analysis pass per change, plus somebody to read the delta — and if you report the inherited backlog on every ticket instead, reviewers learn within a week to click past it.
go deeper
Know that a firewall rule which is correct on its own can still change what an earlier or later rule does, and that the change ticket does not contain the rest of the policy.
Be ready to explain the relation precisely: adding a rule can cover rules below it or be covered by rules above it, and only one of those two directions produces a symptom anybody reports.
Show that you would move the check from the diff to the exported policy and gate on the delta against a baseline, and be able to say why an ungated backlog dumped into every ticket gets overridden.
Own the process argument: a control whose correctness cannot be checked at the unit of change needs a different unit of review, and you should be able to defend the cost of that gate against the alternative of periodic audits.
## Why the review keeps passing Change review on a firewall is built around a ticket: a requested flow, a justification, an approver, and a proposed rule. The reviewer checks the rule against the request. Is the source what was asked for? Is the destination narrow? Is the port right? Is there an owner? Every one of those checks can pass on a rule that leaves the policy weaker than it was, because none of them are questions about the rule's *position*. Correctness in an ordered policy is not a property of a rule. It is a property of a rule and its predecessors. Formally, adding rule *r* to an ordered policy of *n* rules can change behaviour in two directions: - *r* may cover, wholly or partly, the match set of rules below it that carry a different action. If one of those is a deny, that deny stops being enforced. - *r* may itself be covered by rules above it, in which case *r* never fires. A reviewer would need to evaluate both directions against all *n* existing rules across every match dimension — source, destination, protocol, port range, zone or interface. At several thousand rules that is not a diligence problem, it is a computability-by-human problem. ## The asymmetry that decides which one hurts The two directions have wildly different consequences. | Fault | Symptom | Time to discovery | |---|---|---| | The new permit is itself shadowed | The requested flow does not work | Hours — the requester complains | | The new permit shadows an existing deny | None | Indefinite, or an incident | The loud one is self-correcting; the estate's own users are your detector. The silent one has no detector at all, and it is the one an intruder inherits: a permit that the policy's own author believes is denied, sitting in a rulebase that everyone has reviewed and nobody has read end to end. This is also why post-incident review of the change history is unsatisfying. The ticket that created the exposure looks perfect. There is no bad decision to point at, and blaming the reviewer produces a second reviewer and no improvement. ## What actually closes the gap Move the check from the diff to the policy. Whole-policy analysis takes the exported ordered ruleset, decomposes the match space into disjoint regions, determines which rule wins each region, and reports rules that win nothing (unreachable) or that win only part of what their author would expect (partial overlap with a different action). It is static — it needs the policy, not the traffic — and it is exactly the computation a human cannot do by hand. Run it as a pre-push gate on the *proposed* policy, not as a periodic audit of the live one. A quarterly audit tells you about exposure you have already been carrying for a quarter. ## The part everyone gets wrong the first time If you switch the gate on against an inherited rulebase, it will report thousands of anomalies on the very first ticket, none of which the requester caused and none of which they can fix. The predictable outcome is not a cleaner rulebase; it is a gate everybody overrides. Two design choices avoid it: 1. **Baseline the existing policy.** Record the current anomaly set as known, and fail the gate only on anomalies the proposed change *introduces*. That makes the reviewer's question answerable in a minute: did my rule make something new unreachable? 2. **Make the report actionable in the ticket.** The useful output is not a class name, it is the pair — this new permit at position *k* makes deny at position *m* unreachable, and here is what that deny names. That also changes what the gate costs. Reviewing a delta of zero-to-two findings is seconds; reviewing an inherited backlog on every ticket is a job nobody has. ## What it does not fix The gate stops the population growing. It does nothing about what is already there — the several thousand rules lifted wholesale from an old box during a migration, whose authors have left and whose anomalies are already baked into the baseline you just declared known. Those are a separate, larger, and much more expensive problem, and honest candidates say so rather than presenting the gate as a remediation. A final honest note about scope: this reasoning applies to an ordered, first-match policy. It is the classical firewall model and it is the model most large estates are still running, particularly where an on-premises rulebase was carried into a virtual appliance unchanged during a cloud migration.
- Would requiring the requester to list which existing rules their new rule interacts with fix this?Only where they already suspect an interaction. The rules that matter are ones they have never read, often hundreds of lines below, written by people who have left. At several thousand rules the honest answer is that a human cannot enumerate the interactions — it has to be computed from the exported policy.
- The reviewer approved a permit that turned out to be shadowed itself. Is that the same defect?Same mechanism, opposite consequence. A shadowed permit fails closed: the requested flow does not work, someone raises a ticket, and it is fixed within a day. A shadowed deny fails open and silently. The mechanism is symmetric; the risk is not, and that asymmetry is why only one of them shows up in incidents.
- Where in the change flow should the analysis run?Against the proposed policy before it is pushed, compared with the current baseline, so the reviewer sees only what this change introduces. Running it after the push turns it into a report of exposure you already have, and running it without a baseline buries the delta in inherited findings.
saying these in an interview costs you the question
- Blames reviewer diligence rather than the review's scope
- Says a second approver would have caught it
- Treats the rule's comment or justification field as evidence of its effect
- Assumes the appliance rejects a rule it renders unreachable
- Turns on a whole-policy gate with no baseline and expects it to be read