skip to content

Forcing egress through the proxy on a 24/6 line gives you one window - how do you sequence it, and what justifies refusing a 03:00 revert that hands an implant the direct path out again?

level: seniorimportance: nice to knowfreq 32%

answer

  1. the window is not for discovery
  2. log-only mode builds the real inventory
  3. measurement beats vendor documentation
  4. criteria agreed before, not argued at 03:00
  5. prepare a narrow exception as the third option

basics

~20 s

Do not spend the window discovering what breaks. Run the policy in log-only mode first, build the destination inventory from records the gateway already produces, enforce in rings, and agree revert criteria in writing beforehand.

solid answer

~50 s

The window is for enforcement, not for discovery. Weeks before it, run the egress policy in log-only mode and build the real destination inventory from flow records and the server names visible in TLS handshakes - that measurement is what replaces the test environment a plant does not have. Then enforce in rings: office and non-production first, then one least-critical cell, then the line, so the failures you meet in the window are ones you have already seen a version of. Inside the window, sequence the riskiest change first and leave observation time after it. Most importantly, write the revert criteria before the window and get the line owner to agree them: which symptoms trigger a revert, how long you wait, and what the operator may do alone. Then hold a third option ready - a narrow, time-boxed, single-destination exception - because a blanket revert restores direct egress for every host at once, including whichever one is already compromised.

go deeper

for a junior

Know that a change on a production line is rehearsed in a monitoring-only mode first, and that a revert plan is written before the change rather than improvised during it.

for a middle

Explain how log-only enforcement produces a measured destination inventory, and why that measurement is more reliable than a vendor's documented list.

for a senior

Show the whole sequence - log-only, rings, riskiest-first ordering, agreed revert criteria - and articulate why a blanket revert is the widest and most expensive response available at 03:00.

for a principal

Own the negotiation that makes any of it possible: getting a line owner to agree criteria in advance, funding a standing exception review, and deciding what evidence would justify asking for a second shutdown.

## The constraint A line running 24/6 gives you one shutdown to enforce in. There is no environment that behaves like the plant, because the plant is a specific collection of machines, controllers, instruments and licence servers of varying ages, and no lab reproduces it. So the usual software answer - test it somewhere else first - is unavailable, and everything depends on how much you can learn without stopping anything. ## Learning without the window **Log-only enforcement.** Deploy the policy in a mode that evaluates and records but does not block, for long enough to cover the estate's slowest cycle - monthly licence check-ins, quarterly vendor diagnostics, the annual thing nobody remembers. What you are producing is a destination inventory measured from reality: every destination each device actually reached, from flow records and, where sessions use names, the server name in the TLS ClientHello. **Why the vendor's documentation is not a substitute.** Vendor lists are written for a generic install, are usually incomplete, and often name destinations that resolve differently in your region. Measurement beats documentation here, and where they disagree, the measurement is the thing that will break the line. **Ring enforcement.** Office segments and any non-production cell first; then the least critical production cell; then the line. Each ring is a rehearsal that costs you almost nothing and teaches you which classes of software fail and how the failure presents - silent telemetry, a stalled queue, an alarm on a panel. ## Sequencing inside the window Order the changes by risk, riskiest first, so the thing most likely to need investigation gets the most remaining time. Leave real observation time after each step rather than stacking every change and hoping. Have the rollback for each step prepared and tested as a step of its own, since a rollback you have never executed is not a rollback. And know in advance which machines you cannot restart, because a device that will not come back cleanly turns a two-minute change into the outage. ## The 03:00 call This is where the leaf's real question sits. At 03:00 someone will ask for the change to be reverted. The operator who owns the window must not be improvising a policy at that hour, so the criteria are written and agreed beforehand with the line owner: - **What symptom triggers a revert.** Named, observable, on a list: this alarm, this queue stopping, this licence check-in failing. Not "it feels wrong". - **How long you wait before acting.** Many failures are transient re-establishment behaviour and resolve in minutes. - **What the operator may do without waking anyone.** Usually: revert one step, or grant one pre-approved narrow exception. And there must be a third option prepared, because otherwise the choice at 03:00 is binary and the blanket revert wins by default. That third option is a narrow, time-boxed, single-destination exception for the device that broke - written from a template, logged as an exception with an owner and an expiry, and reviewed the next working day. ## What justifies refusing Refusing a revert is legitimate when the symptom is not on the agreed list, when the reported failure is not attributable to the change (a coincidence elsewhere in the estate, a device that was already failing before the window), or when the request is to disable everything because one device broke and a targeted exception would fix it. It is not legitimate to refuse when a named symptom on the list is present and production is being lost - that is what the list was for, and an engineer who overrides it at 03:00 has destroyed the trust the next window depends on. ## Why the blanket revert is the expensive answer Reverting the whole policy restores direct egress for every device at once. It is the widest possible response to a single failure. If anything in that estate is already compromised, the revert hands it back an unmonitored path out for however long the policy stays off - and "temporarily off" has a way of lasting until the next shutdown, which on a 24/6 line may be months away. That is the argument to have in daylight, with the line owner, when the criteria are being agreed, and it is precisely why the narrow exception must exist as a prepared option rather than an idea someone has at 03:00. ## Answering it well Lead with the reframing - the window is for enforcement, discovery happens in log-only mode beforehand - then rings, then sequencing, then the written revert criteria and the prepared third option. Finish on the cost of the blanket revert, in terms of the estate rather than the change: it is not one control off for one device, it is every device's direct path back for an unbounded period.

  • How long should the log-only period run before you enforce?
    Long enough to cover the slowest legitimate cycle in the estate, which is usually monthly licence check-ins and quarterly vendor diagnostics rather than daily traffic. A fortnight catches the routine and misses exactly the destinations that will break the line. If the calendar will not allow it, say what the residual risk is instead of pretending the inventory is complete.
  • The line owner refuses to agree revert criteria in advance. What then?
    Then the criteria default to whatever is said at 03:00, and the enforcement should not go ahead in that window. Record the reason, propose a smaller ring that does not touch the line, and use it to build the evidence that the change is survivable. Going ahead without agreed criteria means the operator carries a decision they were never authorised to make.
  • What do you do with the narrow exceptions granted during the window?
    Each one is logged as an exception with a source, a destination, an owner, a creation date and an expiry, and it is reviewed on the next working day rather than left to become permanent. The ones that survive review get the full treatment - measured destination list, compensating monitoring on destination novelty and volume. The point of the expiry is that the default is removal.

saying these in an interview costs you the question

  • Plans to discover what breaks during the outage window itself
  • Trusts the vendor's documented destination list over measurement
  • Enforces the line and the office segments at the same time
  • Treats the 03:00 choice as revert-everything or hold
  • Leaves window exceptions in place with no expiry or review

context