skip to content

You delete an allow rule from a workload's stateful rule set, yet the traffic it permitted keeps flowing — why?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the decision was already made
  2. when is the rule consulted
  3. first packet, then the flow table
  4. established flows are not re-decided
  5. force a reconnect to verify

basics

~20 s

The decision was already made. A stateful filter matches established flows against its connection table rather than re-deciding them against the rules, so on many platforms a long-lived connection outlives the rule that admitted it until it closes or its tracking entry expires.

solid answer

~40 s

Statefulness has a cost at change time. The rules are consulted on the **first packet of a flow**; after that the flow sits in a connection table and later packets are matched against the table. Remove the rule and new connections are refused immediately, while connections already established keep being matched by their existing entries — on many platforms, though some do re-evaluate live flows when the rule set changes. A tier with long-lived pooled connections can therefore keep talking for hours after a tightening that everyone watched succeed. The verification that actually proves the change is forcing reconnection on the caller and confirming the new connections fail, not observing that nothing broke. The stateless subnet filter is the opposite: with no table, it applies to the very next packet.

go deeper

for a junior

Recall that a stateful filter decides on the first packet of a connection and remembers it, so a rule change is about new connections rather than ones already running.

for a middle

Explain the connection table: later packets are matched against the entry rather than against the rules, which is why a removed rule can leave existing traffic untouched and why an idle flow can be forgotten and then dropped.

for a senior

Show the verification discipline — force reconnection, check flow age, do the destructive test in a lower environment — and state plainly that a removed rule is not an eviction if the concern was an unwanted caller.

for a principal

Make it part of the change standard: every boundary tightening records whether new connections were proven to fail and whether existing flows were evicted, because those are different outcomes and only one of them closes a security finding.

## Where the decision actually happens A stateful rule set attached to a workload evaluates the rules **once per flow**. The first packet is tested, a verdict is reached, and the flow — the address pair, the port pair, the protocol — is written into a connection table. Every subsequent packet of that flow, in either direction, is matched against the table. That is what makes the layer cheap and what removes the need for return-traffic rules. It also means the rules and the live traffic are only loosely coupled. Editing the rules changes what happens to **flows that have not started yet**. On many platforms it does not disturb flows already in the table; on others the table is re-evaluated against the new rule set and offending flows are torn down. Both designs are in the market, so the question to ask of your own platform is: after a rule is removed, are established flows dropped or kept? ## Why this matters during a tightening The search-tier review that starts with "every tier can reach every other tier" ends with a lot of allow rules being removed. The sequence that goes wrong looks like this: 1. The rule permitting a tier to reach the store is removed. 2. The change is applied and someone verifies the tier is still healthy — it is, because its pooled connections were established hours ago. 3. The change is recorded as complete. 4. Days later something restarts, the pool reconnects, and the tier fails in production, far from the change that caused it. The same property has a security shape. If the reason for removing the rule was a caller that should never have had reach, the removal does not evict it. A long-lived connection is exactly what something unwanted would be holding, and "the rule is gone" is not the same statement as "the traffic has stopped". ## How to verify a tightening properly - **Force reconnection** on the calling side rather than watching for breakage. New flows are the only ones the new rules govern. - **Look at flow age**, not just flow count. A boundary where nothing is younger than the change is a boundary nothing has tested. - **Do the destructive check in a lower environment** — apply the same rule removal and confirm the caller actually fails — so production is not the first place the new rules are exercised. - **Where eviction is required**, say so explicitly and use whatever the platform offers to reset the flows; if it offers nothing, restarting the caller's connections is the mechanism. - **Record what you verified**: refusing new connections and evicting existing ones are two different outcomes, and a change ticket that says "tightened" distinguishes neither. ## The opposite failure, from the same table Connection tracking also **expires** entries. A flow that has been idle longer than the tracking timeout is forgotten, and the next packet on it is treated as a mid-conversation packet belonging to no known flow — which is dropped. The symptom is a connection that worked, sat idle overnight, and then hung rather than failing cleanly, because neither side was told anything. Long-lived pooled connections and scheduled batch work are the usual victims. The mitigation is keeping the connection genuinely active at an interval shorter than the tracking timeout, or accepting reconnection and making the caller handle it. | Event | Stateful rule set on the workload | Stateless filter on the subnet | |---|---|---| | Rule removed, flow already established | often keeps flowing until it closes | next packet is dropped | | Rule added for new traffic | effective on the next new flow | effective on the next packet | | Flow idle for a long period | may be forgotten, then dropped | no state, so nothing to forget | ## The takeaway to say out loud A stateful boundary is not a continuously enforced predicate over traffic; it is a decision taken at the start of a conversation and remembered. That makes it efficient and makes rule changes lag reality in one direction and the connection table lag reality in the other. When you change one, verify with a **new** connection; when a long-idle connection hangs for no reason, suspect a **forgotten** one.

  • How do you verify a tightening that removes an allow rule?
    Force new connections rather than waiting for a symptom. Restart or drain the caller's connection pool, then confirm the new flows are refused. Watching a healthy tier after the change proves only that its old flows survived, which is the expected behaviour and not evidence of anything.
  • A connection that worked yesterday hangs after sitting idle overnight. What is the likely cause?
    The tracking entry expired. A stateful filter forgets flows that have been idle beyond its timeout, and the next packet then belongs to no known flow and is dropped silently. Keep pooled connections genuinely active below that interval, or let the caller reconnect and handle it.

saying these in an interview costs you the question

  • Assumes removing an allow rule immediately kills traffic already flowing
  • Declares a tightening verified because the tier is still healthy afterwards
  • Thinks a stateful filter re-checks every packet against the rule set
  • Believes a stateless subnet filter also lets established flows continue
  • Assumes a connection idle since yesterday is still tracked by the filter