skip to content

Break-Glass Overrides

Every gate meets a change that must ship anyway, and what matters is who may say so and what the bypass leaves behind. Interviewers read the override rate as the honest signal about a rule.

on this pageshow

questions

4

What is a break-glass override of a blocking policy gate, and what must it leave behind?

level: juniorimportance: must knowfreq 66%

answer

  1. a planned exit, not a hole
  2. the bypass happens either way
  3. one change, not the whole rule
  4. loudness is the safeguard
  5. who, what, when, which rule

basics

~20 s

A break-glass override is a pre-authorised, deliberately loud way to ship one change that a policy gate denied, used when waiting is worse than the risk. It must leave a record: who pulled it, which rule, which change, and why.

solid answer

~40 s

A blocking gate that offers no legitimate way past it still gets passed - someone merges with elevated rights, edits the check, or deploys outside the pipeline. Break-glass makes that path explicit and observable instead of improvised. Three properties define it: it is authorised in advance so nobody invents it under pressure, it is scoped to a single change rather than to the rule, and it is recorded and alerted at the moment it is used. The record is the whole point - an override that leaves no trace is indistinguishable from a hole in the gate, and the deterrent is that someone else sees it, not that it is hard to do. It bypasses the gate, not the scrutiny: the change still gets reviewed, just afterwards.

go deeper

for a junior

Be ready to define it in one sentence and to say that a bypass must be recorded and attributable to a person. Knowing that it applies to one change rather than switching the rule off is enough at this level.

for a middle

An interviewer expects you to explain why a blocking gate needs a sanctioned bypass at all - that the bypass happens regardless, and the only choice is whether it happens inside the system. Contrast it with downgrading the rule to warn-only.

for a senior

Show that you would wire the alert to a third party and make the effect expire, and that you can tell an emergency from an inconvenience without a rulebook. Name the after-the-fact review as part of the mechanism, not an optional extra.

for a principal

Own the position that break-glass exists at all, and be able to defend it to someone who reads a bypass as a broken control. The argument is about attribution replacing false certainty, and about who is accountable when the pull rate climbs.

## The problem break-glass solves A policy gate that blocks turns a rule into a hard stop: the change does not merge, or the deploy does not run, until the rule is satisfied. That is the point of a blocking gate, and most of the time the answer "fix the change" is correct. But some fraction of denials arrive at a moment when fixing the change properly costs more than the risk the rule was written to prevent - a service is down, a customer-facing failure is spreading, and the correct fix happens to violate a rule. An organisation has exactly two options at that moment. Either there is a sanctioned way through, or people invent one. The invented ones are worse in every respect: an administrative merge, a hand-run deploy from a laptop, a quick edit that makes the check pass without fixing anything. All of them ship the change, and none of them tell you afterwards that a rule was violated. Break-glass exists because the bypass is going to happen; the only question is whether it happens inside the system or outside it. ## What makes it a break-glass override rather than a hole **It is authorised in advance.** Somebody decided, in daylight, that this gate has an emergency path and who is allowed to use it. Nothing is invented at 02:00 except the decision to use it. **It is scoped to one change.** Break-glass lets a specific change past a specific rule. It does not switch the rule off, and it does not apply to whatever else deploys that night. A bypass whose blast radius is "every pipeline until someone remembers" is a rule change wearing an override's clothes. **It is recorded and it is loud.** The pull produces an attributable record at the time it happens - who, which rule, which change, when, and the reason given - and it notifies somebody other than the person who pulled it. The exact fields the record must carry and how long it is kept are a compliance concern; what matters here is that the record exists at the moment of use and names a person, because a reconstruction assembled a week later from chat scrollback is not a record. **It is temporary.** The effect ends with the change. Whatever was loosened either expires, is reverted, or becomes a deliberate rule change through the normal path. ## What it is not It is not the same as switching the rule from blocking to warning. Downgrading a rule changes the default for everyone and keeps changing it until someone flips it back; a break-glass pull spends a single, visible exception and leaves the default deny intact for the next change. It is also not the same as a planned exception granted in advance for a known situation. A planned exception is negotiated with time to think and covers a class of changes. Break-glass is the unplanned case: no time, one change, and the record substitutes for the review that could not happen first. And it is not an admission that the gate is fake. Interviewers probe this, and the answer worth giving is that the deterrent in a bypassable gate is attribution, not difficulty. The rule still states what the organisation wants, the gate still stops the accidental case by default, and the override converts a silent violation into one with a name on it. ## The failure modes worth naming **The silent override.** A skip flag exists, everyone knows it, nothing alerts when it is used. This is the most common failure, and it is worse than having no gate because it produces a green history that nobody can trust. **The override that stops expiring.** Whatever was loosened for one deploy stays loosened, because reverting it was nobody's task once the incident closed. **The override that becomes the workflow.** If a team reaches for break-glass every week, the pull has stopped being an alarm. That is a signal about the rule or the process, and it needs to be read rather than tolerated. **The shared handle.** A bypass performed through a shared account or an unattributed token records that "someone" overrode the rule, which is the one thing the record was supposed to prevent. ## What to say in an interview Define it in one sentence, say why a blocking gate needs one at all, and then name the record and the alert as the parts that make it a control rather than a hole. If you can add that the follow-up review is mandatory and owned by a person, you have answered above the level the question is usually asked at.

  • If anyone can trigger the override, is the gate worth having at all?
    Yes. The gate still states the rule and still stops the accidental case by default - most denials are honest mistakes that get fixed. What the override changes is the small remainder: instead of someone finding an unmonitored way around, the violation ships with a name, a reason and an alert attached. The deterrent is visibility, not difficulty.
  • How does a break-glass override differ from switching the rule to warn-only?
    Switching to warn changes the default for every change and every team, and it stays changed until somebody remembers to flip it back. A break-glass pull leaves the default deny intact and spends one visible exception on one change. If your response to a blocked emergency is to downgrade the rule, you have permanently traded away the control to solve a five-minute problem.
  • Who should be notified when someone uses the break-glass path?
    Someone other than the person who used it, within the same shift - typically the rule's owner and whoever is on call for security. A notification that lands in a monthly report is not an alert; it is filing. The point of notifying a third party is that the pull is observed while the context is still recoverable.

It is the glass on a fire alarm: pulling it is allowed and sometimes right, but it breaks visibly, everyone hears it, and someone comes to ask what happened.

saying these in an interview costs you the question

  • Says break-glass means turning the rule off for the day
  • Thinks urgency removes the need for a record
  • Claims a gate that can be bypassed is not a real control
  • Treats a verbal approval in a chat channel as the record
  • Confuses break-glass with a pre-agreed standing exception

context

open as a page

When you design a break-glass path around a blocking gate, should the override cost a click or a ticket?

level: middleimportance: should knowfreq 50%

basics

~20 s

Make the override cheap to pull and expensive to hide. A ticket-priced bypass nobody can reach at 02:00 gets routed around; a one-click bypass that names the user, alerts others and covers one change keeps the event visible.

open as a page

As the approver holding break-glass authority at 02:00, how do you decide whether to override an egress rule blocking an outage fix?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Decide whether the rule is right about this change, not whether the incident is urgent. Grant the narrowest override that restores service, record it as you grant it, and book the review that either reverts the loosening or changes the rule.

open as a page

One rule accounts for most of your break-glass overrides this quarter - what do you conclude?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Read a concentrated override rate as a defect report on the rule, not on the teams. A rule routinely waved through is too broad, badly placed, or blocking work the organisation has already decided to accept.

open as a page