Your host firewall agent can be stopped by an intruder and guests boot before policy lands — do you fail closed, and who signs that?
answer
- failure and sabotage are different events
- the watchdog dies with the agent
- the boot window is an ordering problem
- correlated failure means fleet-wide outage
- posture per zone, signed by an owner
basics
~20 sFail-closed answers the agent failing, not an intruder who can disable the fail-closed behaviour too. Set the posture per zone, close the boot window by attaching the network only after policy loads, and have a named risk owner sign the availability exposure.
solid answer
~50 sSeparate two failure modes that get confused. If the agent crashes or its policy service is unreachable, fail-closed means the guest keeps a persisted default-deny and loses connectivity — safe, and capable of taking out a whole tier when a bad push hits the fleet at once. If an intruder with local privilege stops the agent, fail-closed buys nothing: they control the watchdog too, so sabotage is answered only by an enforcement point outside the guest. The boot window is separate again — a guest carrying traffic before policy lands is unfiltered, and the fix is ordering, not posture. Then the organisational part: a provider will not underwrite fail-closed on their layer and an availability owner will not sign an estate-wide deny, so you set posture per zone, keep a break-glass path, and get the remainder accepted in writing.
go deeper
Know what fail-closed and fail-open mean for a packet filter: closed means traffic stops when the control is unhealthy, open means it keeps flowing unfiltered, and one of the two is always happening whether or not anyone chose it.
Explain the boot window and why it is an ordering problem: a guest that reaches the network before its policy loads is unfiltered, and the fix is a persisted default-deny or a deferred network attachment.
Show that you would distribute policy in rings, keep last-known-good with a bounded grace period, and guarantee an out-of-band path, because the realistic failure is a correlated fleet-wide push rather than one sick host.
Own the split: posture per zone, an availability exposure accepted in writing by a named owner, and the residual that no in-guest posture answers an adversary holding privilege inside the guest.
## Three failure events, three different answers The question sounds like one decision and is actually three, and separating them is most of the answer. **1. The agent fails on its own.** The process crashes, an upgrade goes wrong, or the policy distribution service is unreachable. This is the case fail-closed is genuinely for: the guest holds a persisted default-deny and stops passing traffic until policy is healthy again. It is also the case that hurts you, because agent bugs and bad policy pushes are correlated events — they hit the whole fleet in minutes, and "fail closed" then means an outage of everything that shares the agent, not of one host. **2. The guest boots before policy lands.** There is a window between the interface coming up and the agent applying its rules. During that window the guest is reachable with whatever the platform's default is. This is not a posture question at all, it is an ordering question, and it has real fixes: a default-deny ruleset baked into the image and applied by the boot sequence before the interface is attached, or the network attachment itself deferred until the agent reports policy loaded. If neither is possible, the window is a residual with a measurable size, and you should know that size. **3. An intruder with local privilege stops the agent.** This is the case that ruins the naive answer. A confident and wrong response is "we run fail-closed, so if an attacker kills the agent the guest goes dark". It does not. Whoever can stop the agent can also stop the watchdog that would have enforced the deny, unload the filtering path, or install a permit rule and leave the agent apparently healthy. A posture is a property of the software behaving as designed; sabotage is the case where it does not. **Failure posture answers failure. It does not answer an adversary at the same privilege level, and only an enforcement point outside the guest does.** ## Why nobody will sign the clean answer The organisational shape of this is the reason it is a leadership question rather than a configuration one. - **The provider will not underwrite fail-closed on their layer.** In shared tenancy, filtering you do not operate comes with the provider's own failure behaviour, and their contract is a best-effort service description rather than a promise to deny traffic when their control plane is unwell. You cannot make their layer fail the way your risk register wants. - **Your availability owner will not sign fail-open.** For a regulated or crown-jewel workload, "when the control is unhealthy, traffic flows unfiltered" is not acceptable to the person who signs the control statement. - Both positions are reasonable. That is what makes it a decision rather than a defect to fix. ## What a defensible decision looks like Set the posture **per zone rather than per estate**, and write down what each zone bought: | Zone | Posture when the agent is unhealthy | What it costs | Who signs | |---|---|---|---| | Regulated / crown jewels | Persisted default-deny; workload stops | Availability exposure to a fleet-wide agent or policy failure | Service owner accepts downtime risk in writing | | General production | Last-known-good policy retained, alert immediately, deny after a bounded grace period | A defined interval of stale policy | Security and service owner jointly | | Low-value / lab | Log and continue | Unfiltered traffic during outages | Accepted by default, reviewed annually | Three things make it real. **A break-glass path that does not share the failing component** — out-of-band console access, so a fail-closed estate does not lock out the people who must repair it. **Staged policy distribution**, because the outage you actually get is a bad push to everything at once, and a canary ring converts it into a bad push to one percent. **Evidence you can produce**: agent heartbeat coverage, count of guests currently on stale policy, size of the boot window, and deny counts — without them the posture is an intention rather than a control. ## The sentence that has to be signed Name the residual explicitly, because it is the one the naive answer hides: *an adversary who obtains local privilege in a guest can disable in-guest enforcement irrespective of failure posture; enforcement against that case exists only where we have a filter outside the guest, which in provider-operated tenancy we do not.* Attach an owner, a compensating control, a trigger that would change the decision — the provider shipping a per-guest filter, or the workload moving to dedicated tenancy — and a review date. An interviewer at this level is listening for exactly that: that you know which risk your posture actually addresses, and that you got the other one accepted by someone with the authority to accept it.
- Why does fail-closed not protect against an intruder who kills the agent?Because the intruder is operating at the privilege level that administers the enforcement, including anything that was supposed to react to the agent's absence. Fail-closed assumes the software stopped working; sabotage means it was made to work differently. Only a control outside the guest is unaffected.
- How do you make fail-closed survivable rather than an outage waiting to happen?Distribute policy in rings so a bad push hits a canary first, keep last-known-good policy with a bounded grace period before hard deny, and guarantee an out-of-band administrative path that does not traverse the control you just closed. Then measure how many guests are on stale policy at any moment.
- The provider will not commit to a failure behaviour on their filter. What do you do with that?Treat it as an unowned residual and escalate it, not as a technical gap to engineer around. Either the workload moves somewhere you control the layer, or a named risk owner accepts in writing that filtering below the guest has undefined behaviour under provider failure, with a review date.
- How would you size the boot window rather than guess at it?Measure the interval between the interface becoming reachable and the agent reporting policy applied, across a real sample of launches including the slowest images. If the number is uncomfortable, change the ordering — default-deny in the image or network attachment gated on policy load — rather than arguing about the number.
saying these in an interview costs you the question
- Claiming fail-closed contains an intruder who can stop the agent
- Choosing one posture for the whole estate with no owner named
- Ignoring the boot window between interface up and policy applied
- Assuming a provider will guarantee fail-closed on a layer they operate
- Designing a fail-closed estate with no out-of-band recovery path
- Treating a fleet-wide policy push as an uncorrelated risk