Your image-signing path is down mid-incident and the fix cannot be signed — what break-glass do you want in place?
answer
- designed bypass or undesigned one
- who may pull it, decided in advance
- one workload, one digest, short-lived
- alert while it happens, not only log
- reconcile the artifact next morning
basics
~10 sA pre-authorised bypass scoped to one workload and digest, expiring by itself, emitting an attributable record that alerts in real time. Without one, the real break-glass is someone disabling enforcement cluster-wide at 02:00.
solid answer
~50 sThe choice is not between bypassing and not bypassing — it is between a designed bypass and an undesigned one. With no break-glass, someone with cluster admin removes the enforcement rule mid-incident: the same bypass with unlimited scope, no expiry and no record. So design it. Name the roles who may invoke it; scope it to the specific workload and digest rather than the namespace; make it expire so enforcement returns without anyone remembering; and emit a record naming who, what digest and why, pushed as an alert rather than buried in a log. The next morning the artifact is reconciled — rebuilt and signed through the normal path, or replaced — and the record linked to the incident. Then watch the rate: repeated use means the signing path is a single point of failure needing redundancy, not a wider exception.
go deeper
Know that an emergency path has to exist and be defined in advance, and that the alternative is somebody turning the whole control off during an outage with no record of it.
Explain the four properties that make a bypass safe — narrow scope, automatic expiry, attributable record, real-time notification — and why each one addresses a specific failure of the ad-hoc alternative.
Demonstrate the operating loop: invoke, restore service, reconcile the artifact the next day, and read repeated invocations as a signal that the signing path itself is a single point of failure.
Own the trade explicitly. Break-glass spends non-repudiation to buy availability; you decide who is authorised to spend it, what evidence the organisation requires in return, and at what usage rate the platform investment beats the exception.
## Why refusing to have one is the worst option An enforcement control that can say no will, eventually, say no at the worst possible moment: the signing service is unavailable, the transparency or key backend is unreachable, or the only build that fixes a live outage cannot be produced through the normal path. At that point somebody restores service. The only question is whether they do it through a path you designed or by reaching for the biggest lever available. The undesigned path is an engineer with cluster admin deleting or relaxing the enforcement rule. Compare it to a designed one: | | Undesigned | Designed break-glass | | --- | --- | --- | | Scope | Whole cluster | One workload, one digest | | Duration | Until someone remembers | Expires on its own | | Record | A config change, maybe | Who, what, why, attributable | | Visibility | Discovered later | Alerts while it happens | Answering "we would never bypass" is therefore not the strict answer; it is the answer that guarantees the first column. ## The properties that matter **Pre-authorised.** Decide in advance who may invoke it — a named on-call role, not "whoever has admin". A two-person rule (one invokes, one acknowledges) is cheap during business hours and worth arguing about for 02:00; if it would meaningfully delay restoration, prefer strong after-the-fact review over a control people will route around. **Narrowly scoped.** The bypass should name the workload and the exact image digest being deployed. A namespace-wide or cluster-wide bypass turns one emergency into an open window through which anything can be deployed, including by someone who did not know an incident was in progress. **Self-expiring.** Enforcement must come back without human memory. The failure mode here is well known: the bypass is opened during an incident, the incident resolves at 05:00, everyone sleeps, and the bypass is still there a year later — at which point it is no longer break-glass, it is a permanent hole with an incident number attached. **Loud, not merely logged.** The record should name the human identity that invoked it, the digest admitted, the reason, and the incident reference — and it should notify a channel someone watches while it is happening. This is what makes the mechanism resistant to misuse: a break-glass path is the obvious route for an insider who wants to run something unverified, and the control that constrains that is real-time visibility plus tight scope, not the absence of a bypass. **Reconciled.** The next working day, the artifact that went out under the bypass is brought back into the normal regime — rebuilt through the signing path and re-deployed, or replaced — and the bypass record is attached to the incident review. Until that happens, something is running in production that nothing vouches for, and your inventory of "what is verified" is quietly wrong. ## What you are actually trading Break-glass spends **audit truth and non-repudiation** to buy back **availability**. That is often the right trade — an outage is a certain, present harm and the unverified artifact is a possible one — but it is a trade, and the reconciliation step is how you pay it back. Framing it that way is what separates a senior answer from a procedural one. ## Treat usage as a signal Break-glass invocations are data about your own platform. If the mechanism fires several times a quarter because the signing path is unavailable, the finding is not "we need an easier bypass" — it is that a security control has become an availability dependency for deployments. Fix that: redundancy in the signing path, a degraded mode you have actually tested, cached verification material, and clarity about whether verification fails open or closed when its backing services are unreachable. A break-glass that is used routinely has stopped being break-glass and has become the deployment process. ## Common weak answers - "We would just disable the policy for that namespace." That is the undesigned path with a smaller radius; it still has no expiry and no attribution. - "It is logged, so we are covered." Logging without notification means misuse is discovered during a review, if ever. - "Security approves each use." A synchronous approval gate at 02:00 either does not happen or delays restoration; approval belongs after the fact, with the invocation itself pre-authorised. - "We make the exception permanent so the same incident cannot recur." This converts one bad night into a standing hole, and it is the single most common way exception sets grow.
- Break-glass has fired four times this quarter. What is the finding?That a security control has become an availability dependency for deploys. The response is redundancy and a tested degraded mode for the signing path, plus a clear decision about whether verification fails open or closed when its backing services are unreachable — not a wider or easier bypass. Routine break-glass use means it is no longer break-glass; it is the deployment process.
- Should a break-glass deploy require a second person to approve before it proceeds?It depends on what the delay costs. Two-person control is a genuine deterrent against misuse and cheap in working hours. If it would add meaningful minutes to restoring a customer-facing outage at 02:00, prefer a single pre-authorised invoker with immediate notification to a watched channel and mandatory review afterwards. The control you route around protects nothing.
- What is the risk if the bypass has no automatic expiry?It survives the incident. The window opens at 02:00, the outage resolves, nobody revisits it, and months later that workload is exempt from verification with an incident number as its only justification. Expiry has to be a property of the mechanism, because the one thing you cannot rely on after a night incident is that someone remembers to close it.
It is the alarmed emergency exit rather than a propped-open fire door: both let you out fast, but only one tells the building who left and closes itself behind them.
saying these in an interview costs you the question
- Claims policy is absolute and no bypass is needed
- Break-glass means disabling enforcement cluster-wide
- Relies on someone remembering to re-enable enforcement
- Logs the bypass but notifies nobody
- Converts the emergency bypass into a standing exception