skip to content

What middle designs sit between fail-open and fail-closed when a policy decision service is unavailable?

level: middleimportance: should knowfreq 48%

answer

  1. not one global switch
  2. classify each rule by harm per bypass
  3. small closed list, large degrade set
  4. answer from last-known-good rules
  5. degraded is its own outcome, and it expires

basics

~20 s

Split rules by what they protect: a small named subset stays fail-closed, everything else degrades to a static allow. Or answer from a last-known-good rule set with an expiry and a loud signal that the fallback is in use.

solid answer

~50 s

The global switch is the wrong granularity, because one outage should not force the same answer for a rule that stops privilege escalation on a workload and a rule that requires a cost-allocation label. Two designs sit in between. First, a **named fail-closed subset**: a short, explicit list of rules that deny when there is no decision, while every other rule degrades to a static allow. Second, a **static fallback answer** evaluated from the last-known-good rule set the caller already holds, so most decisions still come from real rules while the service is down. Both need the same three properties: the degraded result is its own outcome class rather than a normal pass; it is loudly signalled while it is in use; and it expires, so a fallback cannot quietly become the policy. Whatever shipped under it goes on a list to reconcile afterwards.

go deeper

for a junior

Know that real systems rarely pick one global default, and that the usual middle ground is denying for a small important set of rules while letting the rest through with a loud warning.

for a middle

Explain both designs concretely: how rules get classified into a closed subset, how a locally held last-known-good rule set produces a fallback answer, and why the degraded result needs its own status rather than reusing pass.

for a senior

Demonstrate the operating side - the expiry on the fallback, the alert while it is in use, and the reconciliation of everything that shipped degraded once the service is healthy again.

for a principal

Own the trade: graceful degradation costs a maintained list, a distribution path and a second evaluation site, and you accept that complexity because a single flag is wrong precisely when the system is under stress.

## Why the binary is the wrong granularity Fail open and fail closed are the two ends of one dial, and a single global setting forces every rule to share an answer. But the rules do not share a cost. Ask what a *single* unassessed change costs: - a rule denying privilege escalation on a workload: one miss puts an over-powered workload into production, silently, and noticing later does not undo it; - a rule requiring a cost-allocation label: one miss produces an untagged resource that gets fixed in the next sweep. Any global answer is wrong about one of these. So the design question is not "which default" but "how do I get different defaults for different rules, without building a second policy system to decide it". ## Design one: a named fail-closed subset Classify each rule at authoring time by the harm one bypass causes, and put the small, expensive set into an explicit fail-closed list. When there is no usable decision, the caller denies anything the closed set would have looked at and allows the rest. Three things keep this honest: - **The list is explicit and short.** Membership is stated per rule, not inferred from severity labels that everybody inflates. - **Default membership is the degrade set.** A new rule is not fail-closed unless someone argues it in, and the argument has to name the harm. - **It is reviewed.** A closed list that only grows becomes a fail-closed engine with extra steps, and inherits exactly the outage profile you were trying to avoid. There is a subtlety worth saying out loud: when the engine is unavailable the caller usually does not know *which* rules would have matched, because matching is the engine's job. In practice the subset is expressed in terms the caller can evaluate cheaply itself - a coarse predicate over the artifact being gated - which is why the closed set must stay small and simple enough to live outside the engine. ## Design two: a static fallback answer Instead of a fixed allow or deny, answer from the last rule set known to be good. The caller holds a copy, and when the service does not respond it evaluates locally against that copy. Most decisions are then still real decisions, made from real rules, and only rules added since the copy was taken are missing. This is strictly better than a blanket allow and strictly cheaper than a blanket deny, and it comes with obligations: - **An expiry.** A fallback served indefinitely *is* the policy. It drifts: rules written since the copy simply do not exist, and while everything stays green nobody feels pressure to restore the real service. Bound it in time, and when the bound lapses, degrade to the strict behaviour - the closed subset denies. - **Loudness.** Every decision served from the fallback is marked as such, in the result the caller reports and in the metrics. Serving a fallback quietly is the same defect as failing open quietly. - **A reconciliation list.** Record what was decided under fallback so those artifacts can be re-evaluated against the current rules once the service returns. ## Degraded is a third outcome, not an allow The single most common implementation mistake is folding the degraded result into the normal one. The caller has three outcome classes - decided-allow, decided-deny, and no-decision-with-a-default-applied - and only the third one tells anybody the system was running blind. Give it its own status, its own metric and its own alert, and make the pipeline surface say so rather than printing a green tick. ## Cost of the middle designs They are not free. The subset needs a curated list that people have to maintain and argue about. The fallback needs a distribution path for the rule set and a well-defined notion of "known good", and it introduces a second place a decision can be made, which someone has to test. Both add a failure mode of their own: a stale copy, or a subset list that no longer matches how the rules are written. The honest framing in an interview is that you are trading a simple, brittle behaviour for a slightly complex, graceful one, and you take that trade for the same reason you take it anywhere else - because the simple behaviour is wrong at exactly the moment it matters. ## What good sounds like "Most rules degrade to allow and say so loudly; a short list denies regardless; the fallback rules expire in hours, not weeks; and everything that shipped degraded is on a list I re-check on Monday." That answer shows you have thought past the config flag.

  • How do you stop the fail-closed subset growing until it is the whole policy?
    Make membership expensive and explicit. A rule joins the closed list only when someone can state the concrete harm of shipping one unassessed change, the list is short enough to read in a meeting, and it is reviewed on a schedule with the burden of proof on staying in. Default membership is the degrade set. If the list keeps growing you have rebuilt a fail-closed engine and inherited its outage profile.
  • Why must a last-known-good fallback rule set carry an expiry?
    Because otherwise it silently becomes the real policy. It drifts from current rules - anything authored since the copy was taken does not exist in it - and because everything stays green, nobody feels urgency to restore the actual service. A bounded lifetime forces the question: when it lapses, the small closed subset goes back to denying and someone has to deal with the outage rather than living inside it.
  • How does a caller decide the closed subset applies if the engine is the thing that does matching?
    It cannot use the engine's matching, so the subset has to be expressible as a coarse predicate the caller can evaluate itself over the artifact in front of it - a shape check, a field presence check, a small local copy of just those rules. That constraint is a feature: it keeps the closed list small, simple and cheap, which is what you wanted anyway.

saying these in an interview costs you the question

  • Picks one global default and stops there
  • Reports the degraded allow as a normal green pass
  • Puts every rule into the fail-closed subset
  • Serves a fallback rule set with no expiry
  • Assumes a stored fallback rule set stays current by itself

context