skip to content

Your image verifier cannot reach its transparency log at 03:00 — how does that fail at each choke point?

level: seniorimportance: should knowfreq 41%

answer

  1. no attacker, just unreachable material
  2. a failed merge versus a page
  3. steady state survives until something moves
  4. recovery path is where admission bites
  5. verify once, decide locally

basics

~20 s

In a pipeline gate it fails a merge or promotion: delivery stops, nothing running is affected. At admission it fails every workload creation, so autoscaling, rescheduling and node replacement stop — the failure lands in your recovery path.

solid answer

~50 s

The same outage has completely different consequences depending on the seat. A pipeline gate turns it into a failed job: promotion stops, the estate keeps serving, and the blast radius is delivery velocity — painful only if you need to ship a fix during the same incident. At admission it turns into a page, because every new workload object is evaluated: autoscaling cannot add capacity, a replaced node cannot get its pods back, and anything that reschedules stays down. Steady state survives; the moment anything moves, it does not. Pull-time policy on hosts is worst to reason about, because outcomes diverge per node depending on what each cached. So the seat closest to execution is also the one most inside your availability path, and its inputs must live inside your boundary: mirror the material, or verify once at ingest and publish a local decision the seat can evaluate offline.

go deeper

for a junior

Know that verification needs external material — trust roots, identity or log data — and that if the verifier cannot reach it, the check cannot pass. Understand that this can happen with no attacker involved.

for a middle

Explain why the same unreachable dependency causes a failed build in one seat and blocked workload creation in another, and why already-running workloads are unaffected at first.

for a senior

Show that you would rehearse this, move the network dependency out of the workload-creation path, and treat verification material as a production dependency with owners, mirrors and expiry alarms.

for a principal

Own the trade-off that enforcement nearest execution puts your business downstream of the verifier's dependencies, and decide what availability commitment that seat's inputs must carry.

## The failure nobody models Verification is usually designed against an attacker. The outage that actually happens is the boring one: the verifier is fine, the policy is fine, and the **material it needs to make a decision is not reachable**. That material can be an append-only transparency log the verifier queries, an accepted-identity or revocation list, an intermediate certificate, or the trust root itself. This is a no-attacker failure, and it is where the honest cost of each choke point shows up. ## Same outage, three very different incidents **At a pipeline gate.** The verification step fails, the job fails, a merge or a promotion does not happen. Nothing that is running is affected. The failure is visible to a developer, in a log, attached to a specific change. You can wait until morning. The real cost is that delivery is blocked, which matters most in the worst case: you are already in an incident and the fix cannot ship because the gate that guards the fix depends on something that is also down. **At admission.** Every attempt to create a workload is evaluated, so the failure attaches to workload creation across the estate. Pods already running keep running — this is why the first minutes look calm. Then anything that moves stops: autoscaling cannot add replicas under load, a node that is drained or replaced cannot get its workloads back, an evicted or crashed pod that needs rescheduling does not come back. In other words, the failure lands precisely in the **recovery path**, at the moment you most need the estate to be able to move things around. The alternative configuration — skip the check when the verifier errors — is not free either: you have silently turned enforcement off fleet-wide, and nothing about the estate looks unusual while it is off. **At pull-time policy on hosts.** The outcome depends on what each host has cached and when it last refreshed. Existing nodes may keep working; new nodes fail; two hosts behave differently for reasons nobody can see centrally. This is the hardest of the three to diagnose at 03:00 and the hardest to reason about afterwards. ## The expired trust root is the same class, but worse A trust root or intermediate that expires during a long change freeze produces the same symptom with three nastier properties. It is **global** — no partial routes, no lucky region. It is **not fixed by retrying**, because chain validation fails deterministically. And **caching does not save you**, because a cached decision is still made against an expired chain when it is re-evaluated. It is also entirely predictable: the date was known when the material was issued. The controls are boring and effective — treat every root, intermediate and long-lived credential as an inventory item with an owner and an expiry, alarm well ahead of the date, and rehearse the rotation, because rotating a trust root that is already enforcing is a change you do not want to be doing for the first time under pressure. ## Designing so this does not page you Several moves, roughly in order of how much they buy: - **Move the network dependency out of the enforcement moment.** Do the expensive verification once, where a failure is cheap — at ingest into your registry, or at promotion — and publish the result as data the enforcement seat can evaluate with only local material. In practice that means the seat closest to execution decides against a locally replicated, signed set of approved digests rather than by calling out to the internet on every workload creation. - **Keep verification material inside the boundary.** Mirror the log or the identity material into the production network; distribute the trust root out of band as configuration you control; give the mirror the same availability treatment as any other production dependency, because that is what it now is. - **Cache with intent.** Decide, explicitly, how long a positive result stays valid and what happens when the cache is cold. A cold cache during a region rebuild is the scenario that matters. - **Make the dependency visible.** If your workload creation path depends on an external service, that service belongs on your dependency map and in your incident runbook, with the name of the person who can bypass it and the record that they did. - **Rehearse it.** Block the verifier's egress in a non-production cluster and watch what fails. Most teams discover the recovery-path problem this way rather than at 03:00. ## The general principle The more precise the enforcement seat — the closer it sits to the moment of execution — the more of your business is downstream of it. That is not a reason to enforce further away; it is the reason the seat closest to execution must have the fewest and most local dependencies. A check that is one network hop from a decision is an availability dependency of everything it guards.

  • Why is an expired trust root worse than a network partition to the same material?
    Because it is global, deterministic and immune to the usual mitigations. A partition affects some paths and clears; an expired root fails validation everywhere at once, retries do not help, and cached decisions do not survive re-evaluation against an expired chain. It is also entirely foreseeable, so the failure is really an inventory failure: nobody owned the expiry date. Track roots and intermediates with owners and alarms, and rehearse rotation before enforcement depends on it.
  • How would you keep the enforcement seat closest to execution from needing the network at all?
    Do the verification earlier, where failing is cheap, and hand the seat a decision rather than a task. Verify at ingest or promotion, then publish a signed, locally replicated list of approved digests that the enforcement point evaluates with material already on the host or in the cluster. The check keeps its coverage property, and the external dependency moves to a place where an outage costs you a delayed promotion rather than a stalled recovery.
  • During the outage, someone proposes disabling the check until morning. What do you need before agreeing?
    A record of who authorised it, a scope as narrow as the incident requires rather than fleet-wide, a hard expiry on the change, and a way to know afterwards what ran while it was off. The dangerous part is not the decision, it is that a disabled check looks exactly like a working one, so without an alert on the control's own state you will find it still off weeks later.

saying these in an interview costs you the question

  • Assumes running workloads stop the moment the verifier fails
  • Cannot say which seat blocks a merge versus a page
  • Treats skip-on-error as a free safety valve
  • Ignores that admission failures hit the recovery path
  • Never considers mirroring verification material internally

context