Whole-policy analysis of a migrated firewall rulebase returns thousands of findings — which can an intruder actually use, and what happens to a backlog you cannot staff?
answer
- Anomaly count is not risk count
- One class denies less than intended
- Specific exception above a broad deny is design
- Rank by what the deny names, not by range size
- Freeze the population before draining it
basics
~20 sAnomaly count is not risk count. Only one class means the policy permits what its author believes it denies: a deny made unreachable by an earlier permit. Rank those by what the deny names; record the rest as accepted rather than working them.
solid answer
~60 sStart by refusing to treat the report as a work queue. A whole-policy analyser reports several structurally different things and only one of them is an exposure. **A deny shadowed by an earlier permit** is the exploitable class — the policy allows a path its author, its reviewers and any auditor believe is closed. **Partial overlaps** with conflicting actions are second: order decides the intersection, and a human has to say which side was intended. **A broad rule below a narrower one with a different action** is usually deliberate design — the specific exception ahead of the general deny — and mass-normalising those is how you cause an outage. **Redundancy** changes nothing about what is permitted and is pure hygiene. So rank the first class by what the deny names and who can reach it, fix a handful in announced windows with rollbacks, and write the rest down as accepted findings with the reason. The binding constraint here is triage capacity and rule ownership, not detection — buying a better analyser makes the report longer, not the estate safer.
go deeper
Know that a policy analyser reports several kinds of order anomaly and that most of them are not security problems — the one that is, is a deny that can never fire.
Be able to distinguish the classes by what each means: unreachable deny, conflicting partial overlap, deliberate specific-before-general structure, and pure redundancy. Say which of them changes what the firewall actually permits.
Demonstrate triage under a capacity limit: rank unreachable denies by what they name and who can reach it, treat each fix as a change with a window and a rollback, and be honest that the rest is accepted rather than scheduled.
Be ready to defend an acceptance line in writing and to explain why a clean-up without a change gate is a permanent staffing commitment rather than a project.
## The situation Several thousand rules were lifted wholesale from an on-premises box into a virtual firewall appliance sitting in a hub transit network during a cloud migration. Nobody rewrote them; the migration's success criterion was that nothing broke. You now run whole-policy analysis for the first time and get thousands of findings. Your team owns the appliance but not the surrounding fabric's own flow records, so you cannot simply look at traffic to see which of these matter. ## The first move is a refusal The report is not a queue. Treating it as one is the failure this question exists to detect, because it guarantees you spend the year on the cheapest class and never reach the dangerous one. Sort by what each finding means about the *gap between the policy's behaviour and its author's belief*: **1. A deny made unreachable by an earlier, broader permit.** This is the only class where the policy permits something its author believed denied. It is the class an intruder inherits: they do not defeat a control, they use a path the control was written to close and cannot. Everything here is a candidate exposure. **2. Partial overlap with conflicting actions.** Two rules intersect but neither covers the other, and order decides the intersection. Some of that intersection is being treated in a way nobody chose. Second priority, and each needs a person to say which of the two rules meant to own the overlap. **3. A broad rule below a narrower rule with a different action.** This is nearly always intentional structure: a specific permitted exception placed ahead of a general deny, or vice versa. It is how policies are supposed to be written. An analyser flags it because it is technically an order dependency; treating it as a defect and normalising it in bulk is a reliable way to break production. **4. Redundancy.** A rule whose removal changes nothing about what is permitted. Zero security value, real readability value. Fix these only when you are already editing that block. ## Ranking inside class 1 Not every unreachable deny is worth an outage window. Rank by the answer to *who can reach the source side, and what is on the destination side*: - Denies covering management or administrative planes reachable from a lower-trust spoke rank highest — that is the path a foothold turns into control. - Denies covering spoke-to-spoke movement rank next: the hub exists to stop exactly that, and a shadowed deny there quietly re-flattens the network you paid to segment. - Denies on outbound paths from segments where nobody currently has a foothold rank lower — real, but not this quarter. Rank by what the deny names, never by the size of the address range or the number of rules involved in the finding. ## Why fixing one is not a small change The obvious remedy is to make the deny effective — move it above the shadowing permit, or narrow the permit so it stops covering the deny's traffic. Both change what the permit passes, and what it passes today may include a business flow that nobody documented and no analyser can see. Because your team does not own the fabric's flow records, you cannot confirm from traffic what is riding the overlap before you touch it. So each fix is a real change: an announced window, a named owner watching the affected flow, and a prepared rollback — not a bulk edit. That is the honest reason the exploitable class is worked in small numbers rather than eliminated, and saying so is the difference between a senior answer and a plan that will be abandoned in week three. ## The backlog you are not going to clear With thousands of findings and rules whose original owners left the organisation, triage capacity is the binding constraint. Two decisions make the problem finite: - **Set an acceptance line and write it down.** Redundancy: never worked as a project. Broad-below-narrow: only touched when someone is already editing that area. Partial overlaps: reviewed in batches by segment. Accepted findings recorded with their reason are a defensible position; the same findings sitting silently unread are a lapse, and the difference is only ever the write-up. - **Stop the population growing.** Gate new changes against the current baseline so no change may introduce a new unreachable deny. Without that, every rule you clean is replaced, and the clean-up becomes a permanent staffing line rather than a project with an end. ## What a weak answer looks like "We will remediate all findings" — you will not, and committing to it costs you the credibility you need for the twenty that matter. "We will normalise the rule order automatically" — that is a bulk change to a policy carrying undocumented business flows, made by a team that cannot see the traffic. "We will buy a better analyser" — detection was never the constraint; the report is already longer than the team's year.
- How do you rank the unreachable denies against each other?By reachability and target, not by size. A deny covering a management plane that a lower-trust spoke can already reach outranks a deny on an outbound path from a segment nobody has a foothold in. The question is what an intruder standing where they plausibly stand today gains from the permit that shadows it.
- Moving the deny above the permit is the fix. Why is that not a small change?Because it changes what the permit passes. The deny now fires on whatever lived in the overlap, and that may include a business flow nobody documented. Without access to the fabric's flow records you cannot confirm the overlap is empty, so it needs a window, a named owner watching, and a rollback.
- Is redundancy ever worth fixing?Only as a by-product. A redundant rule does not change what is permitted, so it carries no exposure; its cost is that the next reader has one more rule to understand. Fix it while you are already editing that block, and never fund it as a project while unreachable denies are open.
saying these in an interview costs you the question
- Treats every analyser finding as a vulnerability to remediate
- Mass-normalises rule order to clear all anomalies at once
- Ranks findings by address-range size or rule count
- Assumes clean-up can be validated from flow records the team does not own
- Asks for a better analyser when triage capacity is the constraint