skip to content

A cluster's admission gate is blocking the fix for a worsening outage, and someone asks you to switch it off — how do you answer?

level: principalimportance: should knowfreq 32%

answer

  1. ask what is refused first
  2. three reasons, three different fixes
  3. narrowest scope, shortest life
  4. expires by construction, not memory
  5. a gate that lengthens incidents is a defect

basics

~20 s

Find what the gate is refusing first: an unapproved source, a missing verification result and a spec asking for privilege are three different problems. Then unblock the narrowest one, time-boxed — not the whole gate.

solid answer

~50 s

Start by reading the refusal rather than arguing about the switch. The three common reasons have three different narrow fixes, and one of them — a spec asking for host privilege — is a reason to ask what the fix is really doing. If the gate is genuinely in the way, the instrument is an exception scoped to **one workload and one exact set of bytes**, with an expiry the gate itself enforces and an approver who is not the person deploying. The cluster-wide switch is the wrong instrument: it removes the gate for every team, at the moment an attacker would most want it, and it is rarely put back. The exception is when the gate has stopped answering at all, and is refusing everything — then the gate is the outage, and taking it out of the path is the fix rather than a bypass.

go deeper

for a junior

Know that a gate refuses for a stated reason, and that reading that reason comes before any conversation about switching something off.

for a middle

Explain that an unapproved source, a missing verification result and a spec asking for host privilege are three refusals with three different narrow fixes.

for a senior

Argue for the narrowest instrument under pressure: one workload, one exact set of bytes, an expiry the gate enforces, and an approver who is not the deployer.

for a principal

Hold the position that a gate able to lengthen an incident is a design defect, and that a rehearsed narrow bypass is what stops anyone reaching for the wide one.

## First: what is actually being refused? "The gate is blocking us" is not yet a diagnosis. A gate refuses for a stated reason, and the reasons are not interchangeable — each one is a different sentence about the fix you are trying to ship. | Refusal reason | What it is telling you | The narrow fix | |---|---|---| | The reference names an unapproved source | The fix was built or published somewhere the cluster does not accept | Publish it where the cluster accepts, or admit that one set of bytes | | No verification result for these bytes | Nothing the cluster trusts has vouched for the content | Produce the result, or admit those exact bytes under an exception | | The spec asks for host privilege or a host path | The fix itself is asking for more than the workload had | Stop and ask why the fix needs it | The third row is the one to slow down on. Under incident pressure, a spec that suddenly wants the host's privilege set or a host directory mounted is not a gate problem; it is a change nobody has reviewed, arriving at the moment review capacity is lowest. That is precisely the case the gate exists for. ## Why the cluster-wide switch is the wrong instrument - **It is unscoped in space.** One team needs one workload admitted; turning the gate off admits everything, for every team, including whatever else is being deployed or retried at that moment. - **It is unscoped in time.** A switch flipped by a person is unflipped by a person remembering. Under incident conditions, the remembering happens days later or not at all, and "temporary" becomes the running configuration. - **It is exactly the state an attacker would engineer.** Anything that can cause or prolong an outage can cause the gate to be switched off, which makes the switch a documented path to unverified execution. - **It destroys the record.** With the gate out of the path there is no list of what was admitted without checks, so afterwards nobody can say which workloads to re-examine. ## What a narrow exception must have 1. **A scope of one.** One workload, and one exact content-derived digest — not a source, not a team, not a pattern that will match tomorrow's images too. 2. **An expiry the gate enforces.** The exception lapses on its own, and the rule is back whether or not anyone remembered. An exception that ends when someone edits it back is permanent in practice. 3. **An approver who is not the deployer.** Not because the deployer is untrusted, but because an incident removes the pause that review normally supplies, and a second person is the cheapest way to restore it. 4. **A record with the reason.** What was admitted, which check it skipped, who asked, who approved, when it expires — so the follow-up is a task with inputs instead of an argument from memory. ## When taking the gate out of the path is the fix, not a bypass There is an honest case for the wide instrument, and refusing it on principle is its own failure. If the gate has stopped answering — refusing every request, including ones it passed an hour ago, for workloads that have not changed — then the gate is not enforcing a rule, it is failing. That is a fail-closed default firing, and the symptom is uniform: every start refused, across teams, with the same reason. Taking a broken gate out of the request path is repairing the platform, not evading a control. The distinguishing test is simple enough to apply under pressure: **is it refusing one thing for a specific reason, or everything for the same reason?** The first is the gate doing its job and you should use the narrow instrument. The second is the gate being the outage. ## The argument to have before the incident - A gate on the start path is **production infrastructure**, and if its unavailability can lengthen an incident, that is a design defect in the gate — answered with redundancy and with decisions cached against exact bytes, not with a switch. - If the narrow path does not exist and has never been rehearsed, people will reach for the wide one, because at four in the morning nobody invents a scoping mechanism. **Building and drilling the narrow instrument is what keeps the wide one unused.** - Every firing of a bypass is evidence about the rules, not just about the incident. A rule bypassed repeatedly by honest teams is a rule that is wrong, or a pipeline that cannot satisfy it, and that is the fix with the longest half-life.

  • How would you tell whether the gate is blocking your fix or is itself failing?
    Look at the shape of the refusals. A gate doing its job refuses one thing for a specific, stated reason while other teams' starts keep succeeding. A gate that has failed refuses everything for the same reason, including workloads that passed an hour ago and have not changed. The first calls for a scoped exception; the second means the gate is the outage and belongs out of the request path until it works.
  • What makes a time-boxed exception actually expire?
    An expiry the gate itself enforces, after which the rule applies again with no human action. Anything that depends on someone remembering to reverse an edit is not time-boxed, it is indefinite with good intentions — and an incident is the worst possible moment to be creating follow-up work that only a memory will complete.
  • The same rule is bypassed by three different teams in a quarter. What does that tell you?
    That the rule, or the path to satisfying it, is the problem — not the teams. Either the rule refuses something legitimate, or the pipeline cannot produce what it demands within the time teams actually have. Repeated honest bypasses are the most useful signal a gate emits, and treating each one as a one-off incident is how a control quietly becomes decorative.

saying these in an interview costs you the question

  • Turns the gate off cluster-wide and plans to re-enable it later.
  • Treats every refusal as the same problem with one fix.
  • Assumes a bypass someone must remember to remove will be removed.
  • Argues a gate should never block, so nothing ever needs an exception.
  • Lets the person deploying authorise their own exception.
  • Refuses any bypass on principle, even when the gate itself is failing.