Why can a firewall rule's zero hit count mean neither that the flow is dead nor that the reach is closed?
answer
- first match wins
- the rule above ate the traffic
- counters are memory, and memory clears
- which cluster member are you reading
- last hit plus counter epoch
basics
~20 sA counter can read zero because a broader permit above it matched the traffic first, or because the counter was reset by a reboot or failover. Deleting a shadowed rule closes nothing; the reach lives in the rule above.
solid answer
~50 sHit counters are weak evidence in both directions. Downward: in a first-match rulebase, a broad permit earlier in the list absorbs the traffic, so the specific rule below it reads zero forever even though the flow runs every day — delete it and the flow keeps working, which teaches you the wrong lesson, and delete it as a security win and you have closed nothing, because the reach was granted by the rule above. Counters are also volatile: on most platforms they are per-device and per-cluster-member and are cleared by a reboot, so "a year of zero" on a cluster that failed over in March is really four months. Upward: a rule with real hits tells you it matched, not who matched it or whether that use was legitimate. The usable artefact is the last-hit timestamp read together with the counter epoch — when the counters were last cleared — and even that only bounds the window you actually observed.
go deeper
Know that firewalls evaluate rules in order and stop at the first match, so a rule lower down can never see traffic an earlier rule already accepted.
Explain both failure directions — shadowing and counter volatility — and describe how the last-hit timestamp read against the counter epoch is more useful than the raw count.
Demonstrate that you check uptime, policy reloads and failover history before quoting an observation period, and that you look for topology evidence rather than relying on counters alone.
Be ready to argue why a cleanup measured in rules removed is the wrong metric, and what measure of reduced reach you would report instead.
## What the counter actually counts Most firewalls keep a per-rule hit counter and, usually, a last-hit timestamp. The counter increments when a packet or session is matched **by that rule**. That is a narrower statement than "this flow happened", and the gap is where deletion decisions go wrong. ## Downward failure: the rule is shadowed Ordered rulebases evaluate first match wins. Anything a broader earlier rule matches never reaches the specific rule below it: | # | Source | Destination | Port | Action | Hits | Last hit | | --- | --- | --- | --- | --- | --- | --- | | 10 | enterprise | plant-servers | any | permit | 41,900,214 | today | | ... | | | | | | | | 118 | eng-vlan | historian | 1433 | permit | 0 | never | Rule 118 reads zero. The engineering VLAN talks to the historian every shift; rule 10 matches it first. Two wrong conclusions are available and interviewers probe for both: - **"Zero hits, so the flow is dead — delete it."** You delete 118, nothing breaks, and you record a success. What you have proven is only that 118 was redundant *while 10 exists*. The day somebody narrows rule 10 — which is the correct security work — the flow breaks and nobody remembers why. - **"Zero hits, so nothing can use that path."** The reach from the engineering VLAN to the historian on 1433 is wide open. It is granted by rule 10, and deleting 118 changes nothing an intruder can do. A cleanup that removes only shadowed rules reduces the line count and reduces the exposure by nothing. The honest reading is that a zero on a shadowed rule is not evidence about the flow at all. It is evidence about the rulebase's order. ## Upward failure: the counter is volatile and partial - **Reboots clear it.** Counters normally live in memory. A device reboot for an upgrade resets them to zero, so an apparently long observation is only as long as the current uptime. - **Cluster members count separately.** In a high-availability pair, each member counts what it processed. The standby has near-zero counters for everything; after a failover the newly active member starts from its own history, not the pair's. Ask when the last failover was before you believe any figure. - **A counter cannot be replayed.** If you were not already collecting, you cannot go back. Whatever window you have started is the window you have. - **A hit does not identify a user.** A broad permit with forty million hits tells you it is used. It does not tell you by whom, from which hosts, or on which of the ports it allows, so it cannot by itself justify narrowing the rule to anything in particular. For that you need the device's per-rule session or connection records, which have to have been enabled beforehand. ## Reading the pair, not the number The combination that carries information is **last-hit timestamp plus counter epoch**: - Last hit *recent*: the rule is live. Whatever the paperwork says, someone depends on it. - Last hit *old but after the epoch*: genuinely dormant across an observed period. Now the question becomes whether that period spans the flow's real cadence. - Last hit *never, epoch recent*: you have observed almost nothing. This is the state most rulebases are in and the one most often mistaken for evidence. - Last hit *never, epoch old, rule not shadowed*: the strongest counter-based case you can build — and it is still an argument, not a proof. ## Evidence that does not come from the counter Because counters are this weak, the strong deletions usually come from somewhere else entirely: the destination address is in a range that was decommissioned and no longer routes; the address-object group referenced by the rule is empty; the destination sits in a subnet that was collapsed during a re-addressing project. Facts about the topology are independent of how long you watched, and they convince a change board in a way a counter does not. ## The mistake to avoid saying out loud "It has zero hits, so it is safe to remove" is the answer this question exists to catch. The better answer names both directions in one breath: zero can mean shadowed, reset, or simply not yet observed — and a zero-hit deletion that succeeds may have removed a line without removing any reach at all.
- You delete a zero-hit rule and nothing breaks. What have you actually proven?Only that it was redundant given the rest of the rulebase as it stands today. If a broader rule above was absorbing the traffic, the flow still works because that rule still exists. The dependency has not disappeared, it has become undocumented — and it will surface as an outage the day somebody correctly narrows the broad rule.
- A rule shows forty million hits. Why can you not use that to narrow it to the traffic that is really needed?Because the counter is a single number attached to the whole rule. It does not decompose into sources, destinations or ports, so it cannot tell you which subset of a broad permit is actually in use. Narrowing needs per-connection records from the device, and those have to have been enabled before you started, not after.
- How do you find out whether a rule's counters are trustworthy at all?Establish the counter epoch: device uptime, the date of the last policy reload that cleared counters, and the date of the last cluster failover — and check which member you are reading. If the epoch is more recent than the window you are claiming, your evidence is only as old as the epoch, whatever the rule table displays.
saying these in an interview costs you the question
- Says zero hits proves a rule is unused
- Forgets first-match ordering can starve a rule of hits
- Assumes hit counters survive reboots and failovers
- Reads one cluster member's counters as the pair's history
- Thinks removing shadowed rules reduces exposure