skip to content

At 02:00 your policy decision service is down and every deploy is blocked - do you flip the default to allow?

level: principalimportance: should knowfreq 38%

answer

  1. refuse the single-switch framing
  2. which rules, which pipelines, how long
  3. the open window is the blast radius
  4. the flip must expire by itself
  5. reconcile everything that shipped degraded

basics

~20 s

Only as a narrowly scoped, time-boxed change that expires by itself, and never for the small set of rules whose whole purpose is stopping unassessed privilege changes. Restoring service or serving last-known-good rules is the better move.

solid answer

~50 s

Start by refusing the framing that it is one switch. The question is which rules, for which pipelines, for how long. If a last-known-good rule set can still be evaluated, use it - most decisions stay real. If not, keep denying for the short fail-closed subset (rules that stop privilege escalation on a workload are the archetype) and let the rest through, marked as degraded rather than green. A global flip to allow is the last option, it is not one on-call engineer's call to make silently, and it must carry its own expiry so the default returns without anyone remembering. The window is the blast radius: keep the list of what shipped during it and re-evaluate those artifacts once the service is back. And the postmortem action is the design, not the outage - being forced to improvise this at 02:00 means the system had no degraded mode.

go deeper

for a junior

Understand the shape of the dilemma: an unavailable policy engine can block all delivery, and the tempting fix is to turn the gate off, which quietly lets unassessed changes into production.

for a middle

Be able to describe the safer moves - degrade most rules while keeping a small set denying, scope any exception to the pipeline that needs it, and make the change announce itself while it is active.

for a senior

Show the operational mechanics: a self-expiring change, a record of everything that shipped in the window, reconciliation once the service returns, and a repair path that does not run through the gate you just opened.

for a principal

Own the organisational call. Decide in advance what degrades and what never does, who is authorised to widen it and for how long, and drive the postmortem toward the missing degraded mode rather than the crashed service.

## What is actually being asked A policy decision service is down. The gate is fail closed, so every deploy in the organisation is stuck, and it is the middle of the night. Somebody wants the flag flipped. The interviewer is not testing whether you know the flag exists; they are testing whether you can make a security decision under delivery pressure without either capitulating or grandstanding. ## Refuse the binary framing first "Do we open the gate?" is the wrong question because it assumes one dial. The decomposition is: - **Which rules?** Nearly all of them can degrade to allow for a few hours with acceptable risk. A small set cannot - the rules that stop something entering production with more privilege than it should have, where one miss is not recoverable by noticing later. - **Which pipelines?** Usually one team needs to move right now. Opening the default for everything to unblock one change multiplies the number of artifacts you will have to go back and assess. - **For how long?** An hour is a different decision from a week, and the answer must be encoded, not intended. - **Is there a better option?** If callers can evaluate a last-known-good rule set locally, most decisions are still real decisions and the question largely dissolves. ## The blast radius is the window, not the change The cost of flipping is not "one risky deploy". It is every change that ships while the default is flipped, none of which will have been evaluated, and - if you did not record them - which you will not be able to enumerate afterwards. That enumerability is the difference between a bounded incident and a permanent unknown. Before the flip goes in, decide where the list of affected deploys comes from. ## Make the reversion automatic The most common way this goes wrong is not the flip; it is that the flip outlives the outage. Somebody changes it at 02:00, the service comes back at 04:00, and the gate is still open in November. So: - the change carries an **expiry** and restores the previous default on its own; - it is **visible** - announced, alerted, on the dashboard - for as long as it is in force, because a temporary weakening that nobody can see is a permanent one; - it is **as easy to revert as it was to make**, so restoring is not itself a deploy through the gate you just opened. The circular case is worth calling out: if fixing the engine requires shipping a change through the gate the engine is blocking, you need a path that does not depend on the engine at all. Discovering that at 02:00 is a bad time to discover it. ## What you keep denying regardless The short fail-closed subset exists exactly so that this conversation has a floor. Whatever the pressure, rules protecting against privilege escalation on a workload keep denying, because the argument "we will look at it in the morning" does not undo a privileged workload that ran for eight hours. Having that floor written down before the incident is what lets an on-call engineer be firm at 02:00 without negotiating security architecture while sleep-deprived. ## The incident-fix case Often the blocked deploy *is* the fix for another incident. That sharpens the scoping rather than changing the principle: move that one change through a narrow path instead of opening the default for everyone, and have a human look at it, because the reason the machine is not looking at it is that the machine is down. ## Reconciliation When the service returns: 1. take the list of what shipped during the window; 2. evaluate those artifacts against the current rule set; 3. treat any violation as a normal finding with an owner and a date - not as a crisis, and not as something to quietly drop because it is already in production. Without step 1 the window stays unknown forever, and that gap is precisely the thing anyone reviewing the control will ask about. ## The postmortem writes the real answer The finding is not "the service went down". It is that an engine outage forced a security decision to be improvised, under pressure, by whoever happened to be on call. The actions are the design: the named fail-closed subset, the last-known-good fallback with an expiry, an out-of-band path to fix the engine, and a written statement of what degrades and what does not, agreed while everyone is calm. Treat the flip as the symptom. ## What a weak answer sounds like Two failure modes, and both are common. "Security first - we stay blocked until it is fixed" ignores that a frozen organisation is itself a risk and that the incident fix may be in the queue. "Just open it, we will tighten up tomorrow" ignores that tomorrow's tightening never gets scheduled and that nobody will know what shipped. The good answer scopes, time-boxes, records, and then fixes the design that made the choice necessary.

  • The blocked deploy is the fix for the incident the gate itself detected. Does that change your answer?
    It sharpens the scope, not the principle. Move that single change through a narrow path rather than opening the default for the whole organisation, because the smaller the exception, the smaller the set you must reassess afterwards. And a human still reviews it - the reason no machine is checking it is that the machine is unavailable, which is not the same as the change being safe.
  • How do you reconcile the changes that shipped while the default was flipped?
    Keep the list as the window happens, not afterwards - which deploys ran, which artifacts they produced, between which timestamps. Once the service is restored, evaluate those artifacts against the current rule set and treat each violation as a normal finding with an owner and a due date. Without the list the window is permanently unknown, and that is the part any reviewer of the control will press on.
  • What is the postmortem action item here?
    Not just fixing what broke the service. The real finding is that its unavailability forced a security decision to be improvised at 02:00, which means the system has no degraded mode. The actions are the design: a named fail-closed subset, a last-known-good fallback with an expiry, a path to fix the engine that does not run through the gate it is blocking, and an agreement written down while everyone is calm.

saying these in an interview costs you the question

  • Flips the global default and plans to revert it manually
  • Says security wins, leaves the organisation blocked indefinitely
  • Opens the gate for everyone to unblock one deploy
  • Keeps no record of what shipped during the open window
  • Treats the flip as the fix rather than a symptom of missing design

context