300 emergency firewall permits never expired and each is a standing way in: how do you switch them to expiry-by-default without dropping a live clinical workflow?
answer
- make the deadline do the discovery
- enforce nothing on the first pass
- order the waves by what a mistake costs
- silence closes, but announce first
- the re-open needs its own expiry
basics
~20 sIn two stages. First attach expiry dates and enforce nothing, publishing what would drop and when so each flow gets claimed by a named person. Then enforce in waves, lowest blast radius first, never letting a clinical path lapse unannounced.
solid answer
~50 sStart by making the expiry visible before it is real. Attach an expiry date and an owner to every standing permit, enforce nothing, and publish a rolling list of what would drop and on which day. That dry run converts an unanswerable question — is this needed? — into a deadline that makes people identify themselves. Then enforce in waves ordered by blast radius: internal test and lab segments first, vendor-support and clinical paths last, each inside an agreed window with a rehearsed re-open. Anything renewed gets a short new expiry and a named person, never a team mailbox. Anything unclaimed expires, because the whole point is that the default outcome is closed. Two things are non-negotiable on a hospital network: no clinical path expires without announcement, and the emergency re-open path itself carries an expiry, or today's re-open becomes the next two-year permit.
go deeper
Understand why a permit with no end date survives forever, and why a bulk deletion on a live network is a different kind of risk from a bulk deletion in a lab.
Be able to describe the two-stage shape — expiry metadata with no enforcement, then enforcement in waves — and say what each stage is actually for.
Show the operational detail: how waves are ordered, what the change window and rehearsed re-open look like, and how you keep a breakage attributable to one permit.
Expect to defend the programme's cost and its risk appetite to people who own the affected services, and to explain what you will not do after the first breakage.
## What the problem actually is Three hundred standing permits are not three hundred technical objects; they are three hundred unanswered questions, and the unanswered ones default to open. Each was opened to end an outage, each is wider than the fault that caused it, and each is a path an intruder inherits without doing anything clever. The goal is to flip the default so that silence closes a permit instead of renewing it — without the flip itself becoming the incident. ## Stage 0: attribute, do not delete Before any enforcement, give every permit three attributes: what flow it plausibly serves, which zone pair it crosses, and which human is on the hook for it. Zone pair is the one you can derive mechanically and it is the one that ranks risk: a permit from a lower-trust vendor-support zone into a clinical imaging VLAN is not comparable to a permit between two test segments. Ownership is the hard part, and the honest position is that you will not find owners by asking; you find them by scheduling a change and letting the deadline surface them. Note what you are *not* doing here: you are not trying to prove each permit is dead from traffic evidence. That is a separate and harder exercise with its own evidence traps, and waiting for it is how these programmes stall for a year. ## Stage 1: log-only expiry Attach expiry metadata to every permit and enforce none of it. What this stage produces: - A calendar: "on the 14th, these 40 permits would have stopped permitting". - A publication path: that list goes to service owners, the vendor manager and the clinical systems team, weekly, with a date attached. - A claim mechanism: a permit is claimed by a named person accepting the renewal, not by a team saying it is probably needed. The dry run is where the political work happens. "Is this needed?" gets no answers; "this stops working on the 14th" gets answers within a day. It also finds the permits whose owner has left the organisation entirely, which is the population you most want to expire. ## Stage 2: enforce in waves, ordered by blast radius Order the waves by what a mistake costs, not by what is easiest: 1. Lab, test and build segments — a mistake here annoys engineers. 2. Corporate and administrative segments — a mistake here is a helpdesk queue. 3. Vendor-support zone permits — a mistake here means a support engineer cannot connect, which is recoverable during business hours. 4. Clinical and modality paths — a mistake here interrupts care, so it goes last, inside an announced window, with the re-open rehearsed. Each wave has a change window, a named engineer watching, and a pre-agreed re-open that does not require finding an approver at 03:00. Waves are small enough that if something breaks you know which permit did it — expiring 200 rules at once turns any breakage into an archaeology problem. ## The traps that actually bite **The widening is not always a rule.** If an emergency permit was implemented by adding a member to a shared address group, expiring the rule leaves the member behind. The register must track the object that actually carries the access. **Silence is not consent, and it is not refusal either.** An unclaimed permit expiring is the intended outcome, but on a clinical network you must be able to distinguish "nobody uses this" from "the person who uses it works nights and never saw your email". Announcement channels and window timing are how you buy that distinction. **The re-open becomes the new permanent permit.** The most common failure of a successful expiry programme is that everything lapses, half of it is re-opened within a week under pressure, and the re-opens carry no expiry. Every emergency re-open must be created as a time-bound object with a short expiry and a named person, or you have simply reset the clock on the same problem. **Rollback is not "put it all back".** Re-opening every expired permit after one breakage destroys the programme's credibility and rewards the loudest complainant. Roll back the wave, not the initiative, and record why. ## What you tell people afterwards The measure of success is not "300 rules deleted". It is that the standing count stops growing and every remaining permit has an expiry date and a person's name. That is the state in which the next 02:00 widening is temporary in fact rather than in intention.
- Why does the log-only stage find owners when asking for owners does not?Because a question about need is free to ignore and a dated interruption is not. "Is this permit still required?" has no cost attached, so it sits unread. "This stops working on the 14th" creates a consequence with a date, and the people who care identify themselves. It also cleanly separates permits with an owner from permits whose owner has left.
- A wave expires a permit and a clinical workflow breaks. What do you do, and what do you not do?Re-open that permit through the rehearsed path, with a short expiry and a named owner, and record what the flow actually was — you have just learned something the register did not know. What you do not do is suspend the programme or restore the whole wave; roll back the one permit, keep the schedule, and treat the breakage as the discovery the dry run failed to make.
- How do you stop the emergency re-open path from recreating the problem?Make time-bounding part of how a re-open is created, not a follow-up task: a short expiry, the opener's own name, and the incident reference, all set in the same action. If re-opening without an expiry requires more effort than re-opening with one, the default holds under pressure — which is the only time it matters.
saying these in an interview costs you the question
- Expires everything at once and calls the breakage feedback
- Waits for traffic evidence on all 300 before doing anything
- Orders the rollout by ease rather than by blast radius
- Re-opens permits with no expiry when something breaks
- Accepts a team mailbox as a renewing owner