skip to content

Fail Open or Closed

When the engine times out or is unreachable the enforcer still has to answer. Fail-closed defends the rule and can take delivery down with it; fail-open ships and quietly enforces nothing.

on this pageshow

questions

4

A deploy tool calls a policy decision service and gets no answer - what do fail-open and fail-closed mean here?

level: juniorimportance: must knowfreq 66%

answer

  1. a missing answer is still an outcome
  2. who decides when nobody answered
  3. security against delivery availability
  4. the engine becomes a deploy dependency
  5. never log a fallback as a pass

basics

~20 s

Fail open means the deploy proceeds when the decision service cannot answer; fail closed means it is rejected. Open protects delivery and lets unassessed changes through; closed protects the guardrail and can halt every deploy in the organisation.

solid answer

~50 s

Fail open and fail closed describe what the *calling* tool does when the decision service gives it no usable answer - a timeout, a connection error, a 5xx, or a rule set the engine could not parse. Fail open treats the missing answer as allow: the change ships, and anything that would have been denied ships with it. Fail closed treats it as deny: the guardrail holds, but one unavailable service now blocks every deploy, including the one that fixes the incident. It is a security-versus-delivery trade, and it is not one global answer - it depends on what each rule was protecting and what a stopped pipeline costs. In both directions the discipline is the same: the fallback outcome must be visibly different from a real decision. A fail-open allow must be recorded and alerted on, never logged as if policy passed.

go deeper

for a junior

Be ready to state plainly what each default does when no answer comes back, and to name the risk each one accepts. The key idea to hold is that a missing answer is not the same thing as an allow.

for a middle

Explain the caller's view: timeout, connection error, 5xx and an unloadable rule set are one class of event, and the default is applied by the tool making the call rather than by the engine itself.

for a senior

Show you would not pick one global default. Talk about what happens after the fallback fires - the alert, the record of what shipped without a decision, and how those changes get reconciled once the service returns.

for a principal

Own the framing that this is a security-versus-delivery trade with real cost on both sides, and that the organisation should have a stated answer written down before an outage rather than improvising one during it.

## The missing answer is itself an outcome A policy gate has two halves: something that decides, and something that asked. Fail open and fail closed are properties of the **caller** - the deploy tool, the pipeline step, the enforcement point - not of the engine. They describe what that caller does with silence. ## What counts as "no answer" From the caller's seat these are all the same event: - the call times out; - the connection is refused or the name does not resolve; - the service answers 5xx; - the service answers, but with something the caller cannot interpret; - the engine started, but could not load or parse its rule set, so it has nothing to evaluate against. The last one matters because it does not look like an outage from the outside. Collapse all of these into one class - *no usable decision* - and apply one rule to that class. Splitting behaviour by error type is exactly how a gate ends up failing open on the single error nobody tested. Record the distinct cause for the operators; do not branch policy on it. ## Fail open, and why it is more dangerous than an outage On fail open the change proceeds and the pipeline goes green. Two properties make this the more expensive default. It is **silent**: nobody is paged because nothing appears broken. And it is **indistinguishable from a pass** unless you deliberately made it distinguishable: the same green tick, the same log line, the same "policy: ok" in the run summary. What you lose is not only the one change that would have been denied. You lose the knowledge of *which* changes were never assessed. A week later, asked which deploys went out without a decision, you can only answer if you emitted something at the time. ## Fail closed, and the dependency it creates On fail closed the guardrail holds - nothing unassessed ships. The price is that the decision service becomes a hard dependency of delivery: its availability is now your delivery availability. That includes the awkward case where the change that would fix the engine has to travel through the gate the engine is blocking, so the failure is self-sustaining until someone intervenes. Fail closed also has a social cost. A gate that stops the world when it is sick gets a reputation, and reputations turn into pressure to weaken the gate permanently. ## The trade is per rule, not per system The useful question is: *what does one unassessed change actually cost?* - A rule that denies privilege escalation on a workload: one miss puts something in production with more power than it should have, quietly, possibly for months. That is not recoverable by noticing later. - A rule that requires a cost-allocation label: one miss produces an untagged resource that someone fixes next week. Nobody is harmed by the gap. Same engine, same outage, opposite correct default. A single global switch forces you to be wrong about one of them, which is why mature designs sit between the two extremes rather than picking one. ## Make the three outcomes distinguishable Whichever default you choose, the caller has three outcome classes, not two: 1. allowed by a real decision; 2. denied by a real decision; 3. no decision, default applied. Emit the third as its own thing - its own status, its own metric, its own alert - and keep the list of what shipped under it so it can be reconciled once the service is back. A gate that cannot tell you class 3 happened has already lost the argument with an auditor, and more importantly has lost it with itself. ## What weak answers sound like "We fail open, it is fine, code review catches it" - review is a different control with different coverage, and it was not asked to evaluate the rule. "Fail closed is always the secure choice" - true only if you never have to ship anything, and it ignores that a blocked incident fix is itself a security outcome. "The timeout means it passed" - it means nothing was checked; those are not the same sentence. ## How this is asked Usually as a scenario, not a definition: your gate calls a service, the service is down, what happens and who finds out. The interviewer is listening for whether you treat the missing answer as a first-class outcome with its own handling, or as an edge case the config file happens to decide for you.

  • The service returned a 500 rather than timing out. Does that change which default applies?
    No. From the caller's view both are the same class of event: no usable decision. What changes is diagnosis, not policy. Treat timeout, connection refused, 5xx and an uninterpretable response identically at the decision point, and record the distinct cause for whoever is fixing the service. Branching the default on error type is how a gate ends up failing open on the one error nobody exercised in testing.
  • Why is a fail-open allow worse than a visible outage?
    Because it is silent and it looks like success. A fail-closed outage is loud: everyone is blocked, someone is paged, it gets fixed within the hour. A fail-open allow produces a green pipeline and a deployed change nobody assessed, and the window can stay open for a long time. There is also no artifact saying the rule never ran unless the caller deliberately emitted one.
  • Who actually applies the default - the engine or the caller?
    The caller. If the engine is unreachable it is by definition not deciding anything, so the behaviour lives in the tool making the call: the pipeline step, the deploy command, the enforcement point. That is why the default must be tested from the caller's side, by making the service genuinely unavailable, rather than assumed from a configuration setting.

A door with a card reader that has died: fail open unlocks it for everyone, fail closed leaves everyone standing outside. Neither is the right answer for every door in the building.

saying these in an interview costs you the question

  • Says fail open is fine because review catches it later
  • Claims fail closed is always right, names no cost
  • Treats a timeout as evidence the change passed policy
  • Logs a fallback allow identically to a real pass
  • Assumes engine unavailability is too rare to design for

context

open as a page

What middle designs sit between fail-open and fail-closed when a policy decision service is unavailable?

level: middleimportance: should knowfreq 48%

basics

~20 s

Split rules by what they protect: a small named subset stays fail-closed, everything else degrades to a static allow. Or answer from a last-known-good rule set with an expiry and a loud signal that the fallback is in use.

open as a page

At 02:00 your policy decision service is down and every deploy is blocked - do you flip the default to allow?

level: principalimportance: should knowfreq 38%

basics

~20 s

Only as a narrowly scoped, time-boxed change that expires by itself, and never for the small set of rules whose whole purpose is stopping unassessed privilege changes. Restoring service or serving last-known-good rules is the better move.

open as a page

Your policy service is healthy and answering fast but loaded an empty rule set - how do you detect that?

level: seniorimportance: nice to knowfreq 26%

basics

~10 s

Health checks only prove the process is up. Detect it behaviourally: evaluate a synthetic input that must always be denied, continuously, and treat an allow on that probe as an outage.

open as a page