You want the firewalls reachable only from an out-of-band management path — who signs that nobody can reach them when that path fails?
answer
- the separate path needs its own way in
- who is woken when it fails
- security cannot accept someone else's outage
- bring a package, not a risk statement
- measure how often break-glass is used
basics
~20 sThe owner of the service whose availability the outage would hit, not the security team. A path that alone reaches the controls means its failure blocks every fix, so that risk needs a named owner's written acceptance and a tested break-glass.
solid answer
~60 sThe design is sound and the recursion is unavoidable: a separate management path — a dedicated management routing instance, its own circuit, serial console reach — must itself be reachable from somewhere, and that somewhere inherits the trust you just removed from the corporate network. Push the separation as far as it goes and you arrive at a posture where a failure of that path means no engineer can log into any firewall during an incident. That is not an engineering decision, because the cost does not land on engineering: it trades a security blast radius against an availability outage that lands on a service owner at three in the morning. So you take it to whoever owns those minutes with a costed package — a second path in a different failure domain, a sealed break-glass with vaulted per-device credentials and post-use rotation, a drill with a measured time-to-reach — and you get acceptance in writing. If they refuse, you build the weaker version and write down that console reach from the corporate estate is the residual, so nobody later believes a posture that does not exist.
go deeper
Know why management traffic is separated from ordinary user traffic at all, and that the separate path still has to be reachable by the engineers who use it.
Explain the constructions — dedicated management routing instance, separate interfaces, serial console reach, separate circuit — and why each one raises the question of where that path terminates.
Show that you would design the break-glass and drill it, and be able to state the residual risk and the measured time to first console access after a path failure.
Own the acceptance: identify who carries the availability consequence, bring a costed second path, a sealed break-glass and a drilled recovery number, and record the residual honestly when the strict posture is refused.
## The recursion at the centre of it Moving management off the ordinary network is the right instinct: it answers the reachability question that decides whether a workstation foothold is one hop from the policy plane. The usual constructions are a dedicated management routing instance (a management VRF), physically separate management interfaces and switching, serial console reach through console servers, and often a separate circuit so the path does not share fate with production. Every one of them has the same property: **the management path must itself be reachable from somewhere.** Engineers are not inside the rack. Whatever endpoint or network terminates that path inherits exactly the trust you removed from the corporate LAN, and if you attach it back to the corporate LAN for convenience you have rebuilt the thing you were trying to remove. Push the separation to its strict end and you get a posture in which a failure of the management path means nobody can reach any firewall — during precisely the events when you most want to. ## Why this is not an engineering decision The security benefit accrues to the estate; the failure cost accrues to whoever owns availability. During a path outage the escalation lands on a service owner whose customers are affected and who cannot be told "we chose not to be able to fix this." A security architect cannot accept a risk on someone else's behalf, and an on-call engineer certainly cannot — they are the person the posture blocks. So the question "who signs" has a real answer: the accountable owner of the service whose availability is exposed, with the security owner recommending and the operations owner confirming feasibility. In a provider relationship the same acceptance may also have to travel to the client, because their contracted recovery time is the number your posture constrains. ## What you take to the person who signs A bare risk statement gets refused, and a refusal is worse than a conversation, because it usually ends with the strict design being quietly built anyway and nobody knowing. Bring a package: 1. **A second path in a different failure domain.** Not a second link from the same provider through the same duct. Different carrier or medium, so the correlated failure that takes the first is unlikely to take both. This has a price and the price is the point of the meeting. 2. **A defined break-glass.** Per-device local credentials, unique, vaulted, sealed, alerted on retrieval, rotated after every use, and reviewed. Say out loud what it is: standing administrative power that exists only for this failure mode, deliberately accepted. 3. **A measured drill.** Simulate the path failure and record the actual time from decision to a working console session. An untested break-glass is an assumption, and the number you measure is what the signer is really accepting. 4. **The residual, in one sentence.** For example: with the primary path down, first console access takes N minutes via break-glass, and while that is in flight no policy change is possible anywhere. 5. **The cost of not doing it.** The alternative posture leaves console reach available from the general estate, which is the shortest path an intruder has to rewriting policy rather than evading it. ## When they refuse They sometimes will, and that is a legitimate outcome. The failure is not the refusal, it is pretending afterwards. Build the version they accepted — often console reach from a small, hardened set of sources rather than none — and record the residual explicitly: which network can reach the consoles, whose directory those consoles trust, and what an intruder there would rewrite. That record is what stops a future audit, a future architect, or a future incident review from operating on a claimed posture that was never funded. ## Keeping break-glass from becoming the front door Measure its use. A procedure invoked twice a month is not a procedure; it is the primary path, and the real primary path is broken. That measurement is also the cheapest ongoing evidence you have that the design works: rare, alerted, rotated retrievals with an incident behind each one. ## The wrong answer to aim at The answer a strong senior engineer gives is the full design — management VRF, out-of-band, console servers, no corporate reachability — with no mention of who accepts the failure mode or what it costs. It is a good design and an incomplete answer, because the reason estates do not have it is almost never that nobody knew how to build it. It is that the availability posture was never signed, the second path was never funded, and the break-glass was never drilled. The principal-level answer is the one that names the signer, prices the second path, and states the residual in a sentence somebody non-technical can act on.
- The executive refuses to sign the strict posture. What do you build instead?The version they accepted — typically console reach from a small hardened set of sources rather than none — and then you write down the residual explicitly: which networks reach the consoles, whose directory those consoles trust, and what an intruder landing there would rewrite. The defect to avoid is a documented posture that nobody funded and everybody believes.
- How do you stop break-glass credentials from becoming the normal way in?Make them sealed, unique per device, alerted on retrieval and rotated after every use, then measure the retrieval rate. If they are used monthly, the primary management path is the problem and the procedure is masking it. The rate is a design signal, not just an audit artefact.
- What single number makes this decision concrete for a non-technical signer?The measured time from a management-path failure to a working console session via break-glass, taken from an actual drill. It converts an architectural argument into a recovery-time commitment they already understand, and it is the number that changes if you fund a second independent path.
A door that can only be opened from the inside is excellent security right up to the night the fire is on your side of it, which is why someone senior has to decide where the key is kept.
saying these in an interview costs you the question
- Presents out-of-band management as pure upside with no availability cost
- Assumes the security team can accept an outage risk others carry
- Leaves break-glass credentials undefined, shared or never rotated
- Claims a dedicated management path needs no reachability of its own
- Builds the strict posture without ever naming who signed for it