Why is the wide firewall permit opened at 02:00 to end an outage still an intruder's route in two years later?
answer
- restoring service, not granting access
- the ticket closed, the rule did not
- no end date, no named person
- default outcome of nobody caring
- inherited, not exploited
basics
~20 sAn emergency permit survives because nothing removes it: the change was recorded as an outage fix rather than as a permit with an end date and a named owner, and later nobody can prove that deleting the rule is safe.
solid answer
~50 sAt 02:00 the goal is service, not policy. An engineer widens a rule on the boundary between a clinical imaging VLAN and the vendor-support zone, usually wider than the actual fault needed, because nobody yet knows which ports the scanner uses. The ticket then closes on "service restored" and the rule keeps running: it carries no expiry, no end date, and its owner is a team mailbox rather than a person. Two years on the permit is still exactly what it was, and an intruder who gets a foothold in the lower-trust vendor-support zone reaches the imaging archive along it. Nothing clever happens — no exploit, no bypass — the path is simply open. The defender's price is real too: the permit now sits in a rulebase nobody dares prune, because removing it might drop a clinical flow.
go deeper
Be ready to say plainly why a permit opened in an incident outlives the incident: nothing in the process removes it, and the rule carries no end date or named owner.
Explain the mechanics of expiry by default — a rule object with an expiry timestamp, a scheduled removal and a named renewer — and why review-based cleanup defaults to keeping everything.
Show that you have retired one. Talk about how you attributed an old permit to a flow, what evidence you needed, and how you avoided dropping something live when you closed it.
Own the trade: expiry creates renewal work and outage risk that somebody must fund or refuse. Be ready to say who signs for a permit nobody will renew and nobody will let you close.
## The night it is opened A scanner on a hospital's imaging VLAN stops reaching its archive. The vendor's support engineer is on the phone, the on-call network engineer is awake, and the firewall between the imaging VLAN and the vendor-support zone is in the path. Under time pressure, with an unknown fault and a clinician waiting on studies, the fastest safe-feeling move is to widen a rule until traffic flows again. That is a **temporary permit**: an access change made to buy back an outage, intended to be narrowed or removed once the real fault is understood. ## Why it is wider than the fault required At 02:00 you rarely know which ports and which hosts are involved; you know a direction and a rough source. So the permit is written to the widest thing that certainly works — a whole subnet instead of one host, any service instead of one port, an existing address group instead of a fresh, narrow object. Widening beyond the fault is not incompetence; it is what diagnosing while restoring looks like. The defect is not the width. The defect is that the width has no expiry attached to it. ## Why "temporary" turns out to have no mechanism behind it "Temporary" is almost always an intention held in one person's head. Look for the mechanism and there usually is not one: - The change record closes when service is restored, because that was the incident's success criterion. Nothing in that record forces a second visit. - The rule itself carries no expiry timestamp, so the firewall has no reason to ever stop honouring it. - Ownership, where it exists at all, is a team mailbox. A mailbox never renews anything and never feels the consequence of not renewing. - The engineer who opened it moves teams, and the context — what the fault was, how wide the permit had to be — leaves with them. What remains is a rule whose only remaining evidence is its own presence. Its presence proves that someone once needed *something*, not that the flow is still in use today. Establishing that a long-standing permit is genuinely dead is a separate exercise with its own evidence problem, and it is not what the 02:00 engineer can do. ## What the adversary actually gets The important part is how ordinary it is. An intruder who lands anywhere in the vendor-support zone — a support laptop, a jump host, a shared vendor account — does not have to defeat the segment boundary. The boundary already permits their source to reach the imaging archive, with whatever service range the permit was written to. They inherit a path that the network's own design documents do not allow, that no threat model accounted for, and that no alert fires on, because the traffic matches an approved rule. This is also why "nothing has broken, so it must be harmless" is the wrong reading. A standing permit is not harmful in the sense of causing failures; it is harmful in the sense of removing a boundary you believe you have. ## Why the annual review does not close it A review that asks "can we delete this?" defaults to keep, because keeping is free for the reviewer and deleting risks an outage they will be blamed for. Without evidence, the honest reviewer's answer is "I do not know", and "I do not know" renews the rule. That is why review-based cleanup produces a rulebase that only ever grows. ## The only mechanism that closes it Expiry by default. The permit is created as a time-bound object: an expiry timestamp on the rule, a scheduled removal, a named person (not a team) who must actively renew it, and a link back to the incident that justified it. Then the default outcome of nobody caring is *closed* rather than *open*, and the burden lands on the person who wants the access rather than on the person who wants the boundary. ## What that mechanism costs the defender It is not free, and pretending otherwise is how these programmes die. Expiry creates a register that someone must own, renewals somebody must perform, a pager that fires when a permit is about to lapse, and — the sharp one — the risk that an unannounced expiry drops a live clinical workflow mid-shift. On a hospital network that is a patient-safety event, not an inconvenience. So expiry has to be introduced with announcement, dry runs and a scheduled window; and the emergency re-open path itself needs an expiry, or the re-open simply becomes the next two-year permit.
- Why is a team mailbox a poor owner for a temporary permit?Because renewal requires someone to decide, and a mailbox never decides. A named person can be paged, can be asked what the permit was for, and carries the consequence of renewing something they do not understand. A mailbox absorbs the reminder and produces silence, which the rulebase reads as consent.
- Nothing has broken in two years. Why is that not evidence the permit is safe?Because a permit causes no failures by construction — it only removes a restriction. Its quiet existence tells you nobody removed it, not that nobody can abuse it and not that anything still uses it. Absence of breakage and absence of exposure are different claims.
- What should the 02:00 engineer do differently, given they still have to end the outage?Open the widening, but create it as a time-bound object in the same action: an expiry a few days out, their own name on it, and the incident reference. That keeps the outage short and makes the follow-up the default rather than a promise. Narrowing it later becomes a renewal decision, not an archaeology project.
A temporary permit is a fire door wedged open for one delivery at 02:00. Nobody comes back for the wedge, and the door has no way of knowing it was supposed to close.
saying these in an interview costs you the question
- Says an annual rulebase review closes temporary permits
- Treats the change ticket as if it were an expiry
- Assumes an old permit is harmless because nothing broke
- Claims an intruder must exploit something to use it
- Blames the on-call engineer instead of the missing expiry mechanism