skip to content

When the verifier an admission gate calls is unreachable, what does failing open cost and what does failing closed cost?

level: seniorimportance: must knowfreq 50%

answer

  1. the gate itself can be unavailable
  2. open admits, closed blocks
  3. running replicas keep serving
  4. no new starts, no self-healing
  5. cache decisions against exact bytes

basics

~20 s

Failing open admits unverified images exactly when nobody is watching. Failing closed admits nothing new, so rollouts, replacements and scale-ups stop while replicas already running keep serving. Shared clusters usually choose closed and then make it cheap.

solid answer

~50 s

The gate itself is a dependency, and it can be unavailable. The rule has to say in advance which way it fails. **Fail open** admits requests it could not verify, which produces a silent window where unverified bytes enter — and the window opens during an incident, which is the worst possible time for it. **Fail closed** refuses everything it cannot verify: running replicas keep serving traffic, but no new start succeeds, so a rollout stalls, a scale-up does nothing, and a node that dies cannot have its replicas placed elsewhere. That last one is the cost people miss, because it turns one hardware failure into lost capacity. Most shared clusters choose closed and spend the engineering to narrow it: run the checker redundantly, and cache decisions against the exact bytes they were made about.

code

pseudocode · 12 lines
pseudocode
on startRequest(spec):
    imageDigest = resolve(spec.imageReference)
    try:
        result = verifier.check(imageDigest)
        return result.ok ? admit() : refuse("verification failed")
    catch unreachable:
        if cache.has(imageDigest):
            return admit("cached decision for these exact bytes")
        if rule.onVerifierUnreachable == "admit":
            alert("admitted without a verification result", imageDigest)
            return admit("failed open")
        return refuse("verifier unreachable, no cached decision")

go deeper

for a junior

Know that the checker an admission gate calls can itself be unavailable, and that the gate must then either admit images it could not verify or admit nothing new.

for a middle

Explain that failing closed stops new starts while replicas already running keep serving, and that failing open leaves an unattended window whose arrivals persist afterwards.

for a senior

Raise the cost people miss: with new starts refused, a node that dies cannot have its replicas placed elsewhere, so one hardware failure becomes lost capacity.

for a principal

Own the default in writing, then buy it down — redundancy for the checker, decisions cached against exact bytes, and an alarm whenever the fallback path fires at all.

## The two defaults, stated precisely An admission gate that consults a verifier has a dependency, and dependencies fail. The rule must therefore carry an answer to a question that has nothing to do with images: *what do we do when we cannot decide?* There are only two answers, and both cost something. - **Fail open** — admit the request without the verification result you wanted. - **Fail closed** — refuse the request because you could not get the result. This is not a detail to settle when it happens. Whichever way the rule is written is what will happen at three in the morning, so the choice is made in the rule text, in advance, by whoever owns the cluster. ## What failing closed actually stops The first thing to get right is what does **not** stop. Replicas already admitted keep running and keep serving traffic; their processes never consult the gate. The gate stands between a *request to start something* and the platform, so what stops is **new starts**: - a **rollout** halts partway, leaving the old and new versions both present in whatever proportions the replacement had reached; - a **scale-up** produces nothing, so load that would have been spread over new replicas lands on the existing ones; - a **replacement after a node is lost** cannot be placed, which is the consequential one: the platform's self-healing is exactly a new start, so a single node failure stops being self-correcting and becomes lost capacity until the verifier returns. Designs differ on one boundary here and it is worth naming rather than guessing: restarting the same container in place, on the node where it already is, typically does not re-enter admission, while creating a replacement elsewhere does, because that is a fresh request. If your incident plan leans on one of those behaviours, confirm which one your platform has before you need it. ## What failing open actually admits Failing open is attractive because its cost is invisible and deferred. What it does is admit, without a verification result, whatever is submitted while the verifier is down — and three things make that worse than it sounds. 1. **The window is unattended.** It opens when something is already broken and everyone is looking somewhere else. 2. **It is silent unless you make it loud.** A request admitted without a result looks exactly like one admitted with a good result, unless the gate records the difference and alarms on it. 3. **The population persists.** Whatever started during the window keeps running after the verifier returns, because admission does not revisit what is already running. Without a record naming the exact bytes admitted, there is nothing to go back and re-check. | | Fail open | Fail closed | |---|---|---| | New starts | Admitted unverified | Refused | | Replicas already running | Keep serving | Keep serving | | Self-healing after node loss | Works | Stops | | What an attacker gains | A predictable unchecked window | Nothing | | What you know afterwards | Only what you recorded | Exactly what was refused | ## Four ways to make the closed default cheap 1. **Cache decisions against the exact bytes.** A result computed for a given content-derived digest stays true while the verifier is unreachable, because those bytes have not changed. Replacements and scale-ups of what is already running then admit from cache, and only genuinely new content is refused — usually the smaller set during an incident. 2. **Run the checker redundantly.** A gate consulted on every start is production infrastructure and deserves the same availability treatment as anything else on that path. 3. **Resolve early.** Where the platform allows the gate to pin a verified reference to exact bytes at admission, later requests for the same workload describe content that has already been decided on. 4. **Alarm on the default firing.** Whichever way it fails, the fact that it fell back at all is an event worth paging on — it means the gate is no longer doing the job the rule claims it does. ## Choosing before the incident The defensible default for a shared cluster into which many teams deploy is **closed**, for a plain reason: fail-open converts an availability problem into a security problem silently, while fail-closed converts it into an availability problem loudly, and loud problems get fixed. Write the choice into the rule, measure how often the fallback path fires, and treat a gate whose unavailability would strand the estate as a design defect in the gate rather than an argument against having one.

  • Why does caching past decisions narrow the cost of failing closed?
    Because a decision keyed to a content-derived digest does not go stale the way a decision keyed to a name would — those exact bytes cannot become different bytes. While the verifier is unreachable, replacements and scale-ups of workloads already running resolve to content the gate has already decided on and still admit. What gets refused shrinks to genuinely new content, which is normally the smaller set in the middle of an incident.
  • If a gate is configured to fail open, what must it still do?
    Record and alarm. Every request admitted without a verification result should name the exact bytes it admitted and raise an event, so the window is visible while it is open and enumerable afterwards. Without that, the workloads started during the window are indistinguishable from verified ones forever, because admission never revisits what is already running.

A card reader that loses its network has to default one way or the other: unlock and let anyone through, or stay locked and strand the night shift outside. The building decides which before the cable is cut, not during.

saying these in an interview costs you the question

  • Assumes failing closed takes running workloads down.
  • Treats failing open as safe because the window is short.
  • Forgets that blocking new starts also blocks replacement after a node dies.
  • Thinks the fail direction can be decided during the incident.
  • Believes caching a past decision is the same as skipping verification.