skip to content

When you design a break-glass path around a blocking gate, should the override cost a click or a ticket?

level: middleimportance: should knowfreq 50%

answer

  1. friction is the wrong dial
  2. cheap to pull, loud to use
  3. compare it against the workaround
  4. attribution costs no minutes
  5. emergency and planned are different paths

basics

~20 s

Make the override cheap to pull and expensive to hide. A ticket-priced bypass nobody can reach at 02:00 gets routed around; a one-click bypass that names the user, alerts others and covers one change keeps the event visible.

solid answer

~40 s

Friction is the wrong dial. What you actually want from an override is that it is available when it is needed and impossible to use quietly, so you charge for it in attribution and noise rather than in wait time. In practice that means the pull is self-service for people who could already deploy the service, it names them, it alerts somebody else immediately, it affects exactly one change, and it creates a follow-up owed by a person. The test I apply is whether the override is cheaper than the workaround: if filing a ticket and waiting for an approver is slower than an administrative merge, you have not built an override, you have built a detour sign. Approval-gated, ticket-priced paths belong to non-emergency exceptions, which have time to be negotiated properly.

go deeper

for a junior

Know the two shapes an override can take - self-service versus approval-gated - and that even the cheap one has to name the person who used it. Recognise that a bypass with no record is not an override.

for a middle

Be ready to explain the trade directly: friction buys deterrence at the cost of availability, and an override that is unavailable during an emergency is routed around rather than obeyed. Say what you would charge instead.

for a senior

Demonstrate the instrumentation you would insist on - alert to a third party, effect scoped to one run, automatic expiry, a follow-up owed by name - and explain how you would notice the path becoming routine.

for a principal

Own the organisational decision about who is trusted to bypass which gates, and be able to defend a self-service emergency path to someone who wants an approval queue, without leaving the control unusable at 02:00.

## Two different things you might be buying When people argue about how hard a bypass should be, they are usually conflating two goals. One is *deterrence*: making the override unattractive so that it is used rarely. The other is *availability*: making sure that when the override is genuinely needed, it works. Friction buys the first at the direct expense of the second, and the trade is worse than it looks, because the deterrent effect of friction is largely illusory. An engineer at 02:00 with a customer-facing outage and a blocked deploy has options that do not involve your override. They can merge with elevated rights, run the deploy by hand, or take the change through a path where the gate does not run. Every one of those is faster than a ticket queue. So an expensive override does not stop the bypass; it stops the bypass from being *visible*. You end up with a control that looks pristine in its own logs and is regularly circumvented outside them. ## Charge in the right currency The cost that actually deters is being seen. A useful override is: - **Attributed.** It records the individual, not a shared account, not a service identity, not "the pipeline". - **Loud.** It notifies someone other than the person pulling it, at the time of the pull - the rule's owner and the security on-call, in a channel a human reads that shift. - **Narrow.** It applies to one change, one service, one run. A flag that persists silences the rule for every later run of that job, and then the record shows one bypass while many changes shipped. - **Reasoned.** A free-text justification captured at the moment, even a bad one, is worth more than a reconstruction later - it tells you what the person believed at the time. - **Followed up.** The pull creates work owed by a name, not a note in a channel that scrolls away. None of those add minutes to an outage. All of them make routine use uncomfortable, which is the deterrent you actually wanted. ## Who may pull it The instinct is to require a second person. It is worth interrogating. During an incident, the responder can already deploy the service - that is what makes them the responder. Requiring an approver who will almost certainly say yes adds a wake-up, a delay, and a rubber stamp that adds no information. In that context, a self-service pull that shouts is more honest than an approval that is theatre. Outside an incident the calculus flips. There is time, the pressure that justified self-service is absent, and a second pair of eyes catches the case where someone is using break-glass to avoid doing the work. So a reasonable split is: emergency path self-service and instrumented, non-emergency path approval-gated and slow. What you must not do is have only the slow path and pretend the emergency case does not exist. The worst arrangement is a shared break-glass credential that several people know. It converts the strongest property of the record - naming a person - into "the break-glass account did it", and it is nearly always chosen because it was easier to set up than per-person authorisation. ## Keeping the cheap path from becoming the normal path The honest objection to a one-click override is that it becomes how the team ships. Three things hold it back. The alert goes to a human other than the user, so someone notices the third time this week. Every pull creates a follow-up assigned by name, so the cost is deferred, not waived. And the pull rate per rule and per team is reviewed on a regular cadence, so the number is looked at by people who can change something. When the number stops being embarrassing to anyone, the control is already gone - the click was never the problem. ## Failure modes **The persistent flag.** A skip variable set once in a deploy job's configuration keeps skipping. Scope the effect to a single run and let it expire by default. **The unreachable approver.** An override that depends on a person who is not on call is unavailable exactly when it is needed. If you require approval, you must roster it. **The unattributed pull.** Shared accounts, shared tokens, and "we all use the same label" all defeat the record. **Cost measured in delay.** Any design whose safety argument is "it takes long enough that people will not bother" is measuring the wrong thing and will lose to the workaround. ## What interviewers listen for The strong answer resists the false dichotomy in the question. It does not pick click or ticket; it says that emergency and non-emergency bypasses are different mechanisms with different costs, and that the emergency one is priced in visibility. Mentioning the workaround - that an override more expensive than the alternative simply moves the bypass out of view - is the observation that separates a candidate who has run a gate from one who has only read about them.

  • Can the engineer whose change was blocked pull the override themselves?
    During an incident, usually yes - they already hold the access to deploy that service, so a self-service pull that alerts other people is more honest than an approval that will be granted anyway. Outside an incident, no: there is time for a second person, and the pressure that justified self-service is absent. The rule is that self-service and loudness travel together.
  • What stops a one-click override from becoming the normal way to ship?
    Three things. The alert reaches a human other than the person pulling it, in the same shift. Every pull creates a follow-up owed by name, so the work is deferred rather than waived. And the pull rate per rule and per team is reviewed by people with the authority to change the rule. Remove any one of the three and the click becomes routine.
  • The bypass is a variable set in the deploy job's configuration. What is the risk?
    It persists. Set once, it silences the rule on every later run of that job, so the record shows a single bypass while an unknown number of changes shipped unchecked. Scope the effect to one run - a value passed to a single invocation that expires when it finishes - rather than to a configuration that has to be remembered and removed.

saying these in an interview costs you the question

  • Wants bypassing made as painful as possible
  • Accepts a shared break-glass account as attribution
  • Insists two approvers at 3am is always safer
  • Assumes a ticket requirement prevents abuse
  • Sees a skip flag left set as harmless

context