skip to content

The policy gate is your incident: jobs are queuing and every retry goes green on the third attempt. What do you do in the first hour?

level: seniorimportance: should knowfreq 46%

answer

  1. is the gate deciding at all?
  2. denials falling, errors climbing
  3. one rule, one dependency
  4. narrow and time-box, do not disable
  5. list what shipped un-evaluated

basics

~20 s

Make an explicit, time-boxed decision instead of letting timeouts make it. Use the gate's telemetry to find the slow rule and its dependency, narrow that one rule under a named owner and expiry, and record what shipped un-evaluated.

solid answer

~50 s

First I establish whether the gate is deciding at all: duration percentiles per rule, timeout counts, retry counts, and the split between deny outcomes and error outcomes. Green-on-the-third-attempt is the signature of one rule's outbound dependency being slow or throttled, so I isolate which rule and what it calls rather than treating the whole gate as sick. Then I pick one narrow lever and announce it: scope the expensive rule to the changes that actually need it, cut concurrency against the throttled dependency, or downgrade that single rule with a named owner and an expiry timestamp. Not the whole gate, and not open-ended. I tell the blocked teams what changed and what they will be asked to re-run. Finally I capture the list of commits that passed while the gate was degraded, because if I cannot enumerate them the incident permanently hides whatever slipped through.

go deeper

for a junior

Know what to capture and who to tell: which step is slow, how often it times out, and which changes went through while it was degraded.

for a middle

Be able to read the gate's telemetry, durations, timeouts, retries, and the error-versus-deny split, and point at the single rule and dependency generating the queue.

for a senior

Demonstrate a narrow, time-boxed mitigation with a named owner, clear communication to the blocked teams, and a concrete plan to re-evaluate what shipped while the gate was not deciding.

for a principal

Own the standing rule for degraded gates before an incident forces one: who may downgrade what, for how long, what the teams are told, and what the organisation owes afterwards.

## Reading the situation before touching anything Two very different failures look the same from the outside. Either the rules are genuinely catching a lot today, or the gate cannot answer and is producing noise. The instrumentation that separates them is the gate's own telemetry, and if you do not have it the first hour will be guesswork: - **Evaluation duration percentiles, per rule.** A single rule at the tail is the common case; a uniform slowdown across all rules points at capacity or a shared dependency instead. - **Timeout count and retry count, per rule and per pipeline.** These are the numbers that tell you the gate is failing rather than working. - **The split between deny outcomes and error outcomes.** Denials falling towards zero while errors and timeouts climb is the fingerprint of a gate that has stopped enforcing without anyone deciding it should. The pattern in the question, green on the third attempt, is diagnostic on its own. A rule whose verdict depends on an outbound call is being throttled or is timing out intermittently, and every retry is another draw. That points you at one rule and one dependency. ## Deciding, rather than letting the timer decide The worst outcome of this hour is not that the gate blocks people. It is that nobody decides anything, timeouts quietly pass changes through, and afterwards there is no record of what was and was not evaluated. So the deliverable of the first hour is an explicit decision, taken by a person, that is written down and has an end. The levers, roughly in order of preference: 1. **Narrow the expensive rule's scope.** Run it only on the changes it exists for. This keeps enforcement for the population that matters and takes most of the load off immediately. 2. **Reduce concurrency against the struggling dependency.** If the gate is being rate-limited, more parallelism makes it worse; a queue with a smaller in-flight window can go faster end to end. 3. **Downgrade that one rule, with an owner and an expiry.** The precise form matters: one named rule, not the gate; a named person accountable for restoring it; a timestamp after which the state escalates or reverts. Announce it in the same channel the blocked teams are complaining in. What you should not do is raise the timeout until the slow call fits inside it. That trades a visible failure for a longer queue and a bigger blast radius, and it hides the dependency problem instead of surfacing it. ## The list nobody remembers to keep While the gate is degraded, changes are shipping. Some fraction of them were never evaluated. The single most valuable artifact of the incident is a list of those commits, captured at the time, so the gate can be re-run over exactly them once it is healthy. Building that list afterwards, from logs that may have rolled, is far harder, and skipping it means the incident's real cost, whatever it let through, is permanently unknown. ## Communication is part of the mitigation Every minute the queue is stuck, teams are deciding on their own how to get around the gate: opening exception requests, re-running until green, or asking for it to be switched off entirely. A short, concrete message saying which rule is affected, what has been changed, when it expires, and what re-work to expect, keeps that pressure in one channel and out of ad-hoc workarounds. If you say nothing, the workarounds happen anyway and you find them later. ## Knowing when it is over A drained queue is not a recovery signal, because that is also what passing-by-timeout looks like. The incident is over when timeout and retry rates are back inside their normal band and denials are being produced again. And it is not closed until two things have happened: the changes that shipped un-evaluated have been re-run through the gate, and the temporary downgrade has actually expired or been deliberately renewed by the person who owns it. ## The trap that follows this incident The most common sequel to an hour like this is not a repeat incident, it is silence. The rule stays downgraded, the pressure disappears because nothing is blocking anyone, and six months later the control exists only on paper. That is why the expiry and the named owner are part of the mitigation and not paperwork bolted on afterwards.

  • Teams are blocked and asking you to just switch the gate off. What do you say?
    That I will do something narrower and time-boxed instead: name the one rule that is slow, scope or downgrade that rule, attach an owner and an expiry, and tell them what they will be asked to re-run afterwards. A blanket disable is the same action with no end date and no record of what it let through, and it is the version nobody ever reverses.
  • What telemetry do you want in place before an incident like this?
    Per-rule evaluation duration percentiles, timeout counts, retry counts, and the split between deny outcomes and error outcomes, all attributable to a rule and a pipeline. Without that split you cannot tell the rules are catching a lot today from the gate cannot answer, and the symptom, a long queue and a lot of red, is identical in both cases.
  • How do you decide the incident is over?
    When timeout and retry rates are back in their normal band and denials are being produced again, not when the queue drains, because a drained queue is also what passing-by-timeout looks like. And it is not closed until the changes that shipped un-evaluated have been re-run and the temporary downgrade has expired or been explicitly renewed.

saying these in an interview costs you the question

  • Switches the whole gate off and moves on
  • Raises the timeout until the slow call fits
  • Adds retries and calls the incident resolved
  • Downgrades a rule with no owner and no expiry
  • Never re-checks what shipped while degraded
  • Treats a drained queue as proof of recovery

context