How do you order the change windows to move a zone's enforcement onto the real traffic path, when an intruder keeps the unfiltered path until the last one?
answer
- discover, shadow, cut, clean up
- observation window must span a business cycle
- route change and policy change on different nights
- smallest blast radius cuts over first
- rollback is one step at a stated abort time
basics
~20 sObserve first, then move in the smallest reversible steps: build the permit set from real flows over a full business cycle, separate the routing change from the policy change, cut one zone pair or direction per night with a written rollback and a named owner, and accept that the gap stays open, with a dated end, until the last step.
solid answer
~50 sThe sequence is discovery, shadow, cut, cleanup. Discovery: collect the flows on the path that actually carries the traffic, over a window long enough to include the month-end and quarter-end jobs - the flow you never saw is the one that breaks the night. Shadow: put the policy at the new enforcement point in a non-blocking posture and reconcile what it would have denied against the observed set; every would-be deny needs an owner who confirms it is dead or an exception with an expiry. Cut: change routing and policy on separate nights, because doing both at once leaves you unable to say which one broke the application. Order pairs by blast radius, smallest and most reversible first, and make each night end with the estate working. Every window needs a change owner, an abort time and a rollback that is 'remove the policy' rather than 'debug live'. Cleanup - removing the old bypass path - is its own night, and until it happens nothing is actually enforced.
go deeper
Know that turning on a filter is a change with a window, an owner and a rollback, and that the flows it will block have to be discovered before the night, not during it.
Explain why the old rule base is an incomplete source for the new permit set, and why a discovery window has to cover the estate's longest business cycle.
Show the sequencing judgment: shadow before block, routing and policy on separate nights, smallest blast radius first, a one-step rollback at a stated abort time, and cleanup of the bypass as its own change.
Be ready to defend the programme's duration to a stakeholder who wants it done this quarter, and to state plainly that containment accrues per zone pair cut over, not per milestone closed.
## The shape of the problem You have established that a documented boundary is not on the path: an overlay stretch, a newer fabric path or an interconnect built to make two merged estates talk carries the traffic around it. Nothing about that is fixed by a decision. It is fixed by a sequence of maintenance nights, each of which must leave the estate working, and one of which will be the night an undocumented flow meets a policy for the first time. You are the change owner on that call. ## Stage 1 - discovery, and why the window length matters more than the tooling The permit set cannot be copied from the old rule base, because the old rule base only ever saw the flows that transited it - and by hypothesis the interesting flows did not. So the permit set is built from observation on the path that really carries the traffic. The single most common failure is a discovery window that is too short. Estates run on cycles: month-end batch, quarter-end reporting, an annual reconciliation, a disaster-recovery test that only replicates once a quarter. A two-week observation window plus a cut on the wrong weekend produces a finance outage with your name on the change record. Where you cannot observe a full cycle, you compensate: ask the application owners, and stage the cut so the first pairs to move are the ones with no batch behaviour. Record what you can and cannot see while doing it. Flow records tell you that a pair talked and on what ports; they do not tell you which business process depended on it. That mapping comes from owners, and getting it is a large part of the calendar cost. ## Stage 2 - shadow, and the honesty it demands Before blocking, run the intended policy in a posture where it logs what it would deny and denies nothing. Reconcile every would-be deny into one of three buckets: expected and permitted (add the rule), dead and to be removed (confirm with the owner), or unknown (an exception with an expiry date and a name against it, never an open-ended one). Say the uncomfortable part out loud in an interview: a shadow period is a window in which you know the gap exists and are deliberately not closing it. That is a defensible decision only if it is time-boxed, has an owner, and is recorded as an accepted risk with an end date. It is not defensible as an indefinite state, and a shadow policy that has been logging for eighteen months is a failure dressed as caution. ## Stage 3 - cut, in the smallest reversible steps Three rules do most of the work: 1. **Separate the routing change from the policy change.** Moving the path and turning on enforcement on the same night means that when something breaks you cannot tell which caused it, and your rollback has to undo both under time pressure. Move the path first, confirm the traffic now transits the enforcement point with the policy still permissive, then enforce on a later night. 2. **Order by blast radius, smallest first.** Start with a zone pair or a direction whose flow set is small, well-owned and easy to roll back. Each success buys credibility for the next window; a first-night outage on the payments path can stall the whole programme for a quarter. 3. **Make the rollback trivial.** The abort action must be one reversible step - withdraw the policy, restore the route - decided by a stated abort time, not by a debugging session at 03:00. What you watch on the night matters too. The device's deny counters are necessary but not sufficient: watch the application's own error rate and the owners' checks, because the flows that break most painfully are the ones whose failure looks like slowness rather than a refusal. ## Stage 4 - cleanup, and why it is not optional Until the bypass path is removed or re-scoped, the enforcement point is one route away from being irrelevant again. Cleanup is its own change, with its own window, and it is the step most likely to be dropped when the programme runs out of budget. A programme that stops after stage 3 has bought monitoring, not containment. ## What the intruder experiences through all of this Nothing, until the pair they would use is actually cut over. That is the honest framing for a status update: percentage of zone pairs enforced, not percentage of the project plan complete. An intruder in the estate does not get partial credit for a programme in flight. ## Interview signal Strong answers talk about calendars, owners and rollbacks as engineering artefacts, not project management overhead. Weak answers describe the target state in detail and treat getting there as an implementation detail - which is exactly the part that fails.
- An application owner refuses to confirm whether a flow is still needed. What do you do on the night?Do not guess in either direction. Either the pair moves to a later window, or the flow gets a named, dated exception carried into the cut - a permit with an expiry and an owner, reviewed on the date, not an open permit added quietly to the base. What you must avoid is an unowned permit surviving the programme, which is exactly how the next generation's undeletable rule is created.
- Why not enforce everywhere in one night and revert if it goes wrong?Because reverting requires knowing what broke, and a simultaneous cut over many pairs gives you a room full of failures with one cause each. Recovery time becomes unbounded, which is precisely what a maintenance window is not allowed to be. Small steps trade programme duration - during which the gap stays open - for a bounded worst case on each night, and that trade is the right one when the rollback path is what protects the business.
- How do you report progress to a security stakeholder who wants a single number?Report the share of zone pairs whose traffic now transits an enforcing device with policy active, plus the count of live exceptions and their expiry dates. Project-completion percentages overstate protection, because containment arrives step by step and only for the pairs already cut over.
saying these in an interview costs you the question
- Copies the old rule base as the permit set for the new enforcement point
- Observes for a week and cuts over before month-end
- Changes routing and enforcement on the same night
- Runs a shadow policy indefinitely with no expiry or owner
- Declares success before the bypass path is removed
- Plans no rollback beyond debugging live at 3am