skip to content

As the approver holding break-glass authority at 02:00, how do you decide whether to override an egress rule blocking an outage fix?

level: seniorimportance: should knowfreq 60%

answer

  1. urgency is not evidence
  2. ask what the rule actually denied
  3. look for the smaller change first
  4. narrowest grant, with an expiry
  5. book the review before standing down

basics

~20 s

Decide whether the rule is right about this change, not whether the incident is urgent. Grant the narrowest override that restores service, record it as you grant it, and book the review that either reverts the loosening or changes the rule.

solid answer

~50 s

First I read what the rule actually denied. A change opening traffic to a new external destination has exactly the shape the egress rule was written to catch, so urgency does not settle the question: I ask the responder what the endpoint is, who owns it, and whether a smaller change restores service - an already-approved destination, disabling the failing feature, or routing through an existing path. If the override is the right call, I grant the narrowest one available: this change, this service, this window, never a global skip that exempts everything else deploying tonight. The pull is recorded and alerted as it happens, because the incident channel is not an audit trail. Then the part people forget - a review within days, owned by a name, that either removes the temporary allowance or converts it into a rule change through the normal path.

go deeper

for a junior

Know that overriding a gate in an emergency is allowed, that you say clearly what you overrode, and that the bypass is recorded rather than mentioned in passing.

for a middle

Be ready to explain scoping: why a per-change override differs from disabling the rule for the pipeline or the whole organisation, and why any temporary loosening needs an expiry rather than a reminder.

for a senior

Show the decision itself - interrogating the denial, hunting for a smaller change that violates nothing, granting the narrowest form, and owning the review that follows. Say plainly that urgency is not evidence about risk.

for a principal

Own the standing arrangement: who may make this call at 02:00, how the organisation distinguishes a real emergency from schedule pressure, and what it costs to be the approver who is either always available or always bypassed.

## Why this is the hardest chair in the room The approver at 02:00 has incomplete information, a colleague under pressure, and a rule that is doing exactly what it was written to do. The rule denies a change that opens outbound traffic to a destination nobody has approved. That is the shape of a legitimate vendor failover endpoint, and it is also the shape of an exfiltration path. Nothing about the urgency of the outage distinguishes the two, which is the first thing to say out loud: **urgency is not evidence about the risk.** It changes the cost of saying no; it does not change the probability that the rule is right. ## The three questions worth asking before you answer **What exactly did the rule deny?** Not "the deploy failed policy" but which rule, on which field of which change. Half the time the denial is narrower than the responder thinks, and satisfying it is a two-minute edit rather than a bypass. This is also the moment you find out whether the responder has read the denial at all. **Is there a smaller change that restores service?** Options in rough order of preference: use a destination already on the approved list; turn off the feature that needs the new destination and accept degraded service; ship the fix without the egress change and follow with the rest later. A partial restoration that violates nothing often beats a full one that opens a hole, and the responder under pressure has usually not enumerated these. **If I grant it, what else does it open?** This is the scoping question, and it is where the damage is usually done. A per-change override affects one deploy. A flag on the pipeline exempts everything that pipeline ships tonight. Turning the rule off globally exempts every team, including the ones whose changes have nothing to do with this incident and who will never know the gate was down. ## Grant narrowly, and grant with an expiry If the answer is yes, the shape of the yes matters more than the yes. Prefer, in order: an override attached to this single change; a temporary allowance for this one service and this one destination with a defined end; and only as a last resort anything that turns a rule off for a class of changes. Whatever you grant should end by default - expire, or revert on the next reconciliation - so that forgetting the follow-up restores the deny rather than preserving the hole. Loosenings that require someone to remember to remove them do not get removed. Record it as you grant it. Who asked, what was granted, on what scope, and the reason given at the time. This takes thirty seconds and it is the difference between an override and a mystery. Reconstructing it later from a chat channel produces a story, not a record, and the version people remember is the one that makes the decision look obvious. ## The part that decides whether any of this was worth it An override creates an obligation that outlives the incident. Within days, and while people still remember, the rule's owner and the responder reconvene on one question: was the rule right? There are only three honest outcomes. - **The rule was right.** The change should not have shipped in that form. Remove the allowance, do the work properly, and note that the override bought time at a real cost. - **The rule was too broad.** It denied something legitimate because it could not express the distinction that mattered. The rule changes, and the change is tested against the case that broke it. - **The destination is legitimate and permanent.** It joins the approved set through the normal path, and the temporary allowance disappears. What is not an outcome is silence. A temporary egress allowance that survives the incident is a standing hole with no owner and no record of why it exists - and it will be found by someone six months later who cannot tell whether removing it breaks production. ## Failure modes an interviewer is listening for **Refusing on principle.** "Policy is policy" trades a certain, present outage for a hypothetical risk, and it teaches the organisation to stop asking you. The approver who never says yes is quickly not consulted. **Granting globally because it is faster.** The broad switch is always the quickest thing to do at 02:00, and it is the one that produces the incident nobody notices for months. **Treating the incident channel as the record.** Chat is where the conversation happened; it is not where the decision is recorded, and it will not be searchable when it matters. **Skipping the review.** If the reconvene is optional it does not happen, because the people who owe it are the people who were up all night. ## What a good answer sounds like Interrogate the denial, look for a smaller change, grant the narrowest override that restores service, record it at the time, make it expire, and own the review. Say explicitly that you would not let the severity of the outage substitute for a judgement about the risk - that sentence is the one that shows you have actually held the pager and the authority at the same time.

  • The responder says the endpoint is a vendor failover host they cannot verify at 02:00. Do you grant it?
    Usually yes, but scoped: one service, that one destination, time-boxed, with the traffic watched while it is open. Refusing on unverifiability trades a certain outage for an uncertain risk. Granting broadly trades it for a permanent one. Verification then becomes an explicit item in the review rather than something that quietly never happens.
  • How can the override itself become the next incident?
    Two ways. A temporary egress allowance that is never removed becomes a standing hole nobody owns, and by the time it is found nobody can say whether removing it breaks production. And a global skip granted for speed exempts every other change deploying that night - including ones nobody was watching, because everyone's attention was on the outage.
  • How do you stop the post-incident review from being skipped?
    Create the follow-up automatically when the override is pulled and assign it to a person, not a team. Make the loosening expire by default, so ignoring the review restores the deny instead of preserving the exception. And put the open follow-ups in front of whoever reviews gate health, so an unclosed one is visible to someone other than the exhausted responder.

saying these in an interview costs you the question

  • Skips all checks because it is an incident
  • Disables the rule globally and cleans up later
  • Treats the incident chat as the official record
  • Refuses any override during an incident on principle
  • Considers the matter closed once service is restored
  • Lets outage severity substitute for judging the risk

context