skip to content

A CI policy check times out and the pipeline records it as passed. Why is that unsafe?

level: juniorimportance: must knowfreq 62%

answer

  1. count the possible outcomes
  2. silence is not consent
  3. when do timeouts cluster?
  4. unevaluated must not read green
  5. trace exit code to check result

basics

~20 s

A timeout means the rule produced no decision at all, not that the change is compliant. Recording it as a pass turns an overload into a silent approval, and that happens most often exactly when the system is busiest.

solid answer

~40 s

A gate has three possible outcomes, not two: allowed, denied, and could-not-evaluate. A timeout is the third one, so collapsing it into `passed` makes the pipeline report a guarantee it never obtained. The failure is not randomly distributed either: timeouts cluster under load, on big release days, during incident recovery, or when a slow upstream the rule has to call is throttling, so the changes that slip through un-evaluated are exactly the ones going out under pressure. I would render the undecided case as its own visible result, with a reason like `policy evaluation did not complete`, count it in the gate's own telemetry alongside timeouts and retries, and escalate it to a person rather than letting a timer make the call. Otherwise nobody downstream can tell an approved change from an unexamined one.

go deeper

for a junior

Be ready to say plainly that the check did not finish and the check approved it are different results, and that only one of them is evidence about the change.

for a middle

Expect to explain how a step's exit status becomes a reported check result, and where a timeout gets swallowed, in wrapper scripts, caught errors, or continue-on-error settings, so an unevaluated run reports green.

for a senior

Show that you instrument the gate itself: timeout and retry counts, duration percentiles per rule, and a way to enumerate and re-evaluate the changes that shipped while the gate was not deciding.

for a principal

Own the position that the undecided outcome is a policy question rather than a timeout setting. Someone must own what happens to work arriving while the gate cannot answer, and that choice should be written down rather than emergent.

## Two outcomes is the wrong model Most people picture a gate as a boolean: the change is compliant, or it is not. A real gate has a third state, and every serious failure in this area starts by ignoring it: | Outcome | What it means | What it proves about the change | |---|---|---| | Allowed | The rules ran and found nothing to block | The change satisfies the rules that ran | | Denied | The rules ran and something matched | The change violates a named rule | | Undecided | The evaluation did not complete | Nothing at all | A timeout, an out-of-memory kill, a crashed evaluator, an unreachable data source the rule needed: all of these land in the third row. They are not verdicts. When a pipeline maps that row onto the first one, it is not making a risk tradeoff, it is publishing a false statement, and every consumer downstream, the reviewer glancing at the green tick, the release dashboard, the quarterly coverage report, reads it as evidence. ## Why the mapping happens by accident The collapse is rarely a decision anyone made. It is usually mechanical: - A wrapper script catches the timeout and exits zero so the job does not look broken. - The check is wrapped in a `|| true` or an equivalent continue-on-error setting that was added months ago to stop one flaky rule from paging anyone. - The step has a time limit, and the pipeline treats an expired step as skipped, and skipped steps do not fail the job. - The evaluation runs in a background step whose result is never joined back into the job's status. In each case the pass is a property of the plumbing, not of the policy. That is the first thing to check when the log says the evaluation was cut short but the job is green: trace the path from the evaluator's exit code to the reported check result and find where the non-zero was swallowed. ## Why the un-evaluated changes are the risky ones If timeouts were uniformly random you could reason about them as a small, tolerable sampling gap. They are not. Evaluation time goes up when the queue is deep, when a rule's outbound dependency is slow or rate-limiting, when many pipelines run at once, and when a change is unusually large. Those conditions correlate with release days, with hotfixes going out during an incident, and with the end of a quarter. So a gate that quietly passes on timeout is biased towards approving the traffic that was moving fastest and getting the least human attention. Saying `it only happens under load` is not a mitigation, it is a statement of the exposure. ## What to do instead Make the undecided outcome first class: 1. **Emit it as a distinct result.** Not green, and ideally not the same result as a violation either, because reusing the deny result hides the difference and teaches people to re-run denials. A separate status with a reason string is what lets you count it. 2. **Instrument the gate itself.** Evaluation duration percentiles per rule, timeout counts, retry counts, and the split between deny outcomes and error outcomes. Without that split, a spike in red looks the same whether the rules are catching a lot today or the gate cannot answer at all. 3. **Put a human in the loop for the undecided case, not a timer.** Whether work is allowed to continue while the gate cannot answer is a decision somebody owns; a step timeout is a number somebody picked to keep jobs from hanging. Those should not be the same lever. 4. **Keep a list of what went through undecided.** If ten changes shipped while the evaluator was failing, you want to be able to re-run the gate over exactly those ten afterwards. If you cannot enumerate them, the outage permanently hides whatever it let through. ## The distinction to hold on to `No violations were found` and `no answer was produced` are different sentences. Only the first is about the change. A gate that cannot tell them apart is not enforcing a policy; it is producing a green tick whose meaning depends on how loaded the system was at the time, which is the same as no meaning at all.

  • How would you make the undecided outcome visible without halting every pipeline the moment the gate is slow?
    Emit it as a distinct, non-green result with its own reason, alert on the rate rather than on each occurrence, and queue the affected commits for a re-check once the gate recovers. Work can keep moving, but the decision to let it move is a person's, taken with a count in front of them, instead of a step timeout's.
  • The step is green but its log ends with a deadline-exceeded message. What do you check first?
    Whether the reported result is actually derived from the evaluator's exit status, or from a wrapper that caught the timeout and exited zero. Swallowed non-zero exits, a trailing `|| true`, and continue-on-error settings are the usual cause. A rule that never ran should not be able to produce a pass at all.
  • Is a timeout the same as a rule that returned no violations?
    No. An empty violation list is a result: the rules ran over the input and matched nothing. A timeout is the absence of a result, so it carries no information about the input. Systems that model both as an empty list cannot distinguish a clean change from an unexamined one, and every report built on top inherits the confusion.

A metal detector that is switched off does not mean the passenger is unarmed. A queue moving quickly because the scanner died is not a security result.

saying these in an interview costs you the question

  • Says a timeout is fine because the check usually passes
  • Reads a green check as proof the change was evaluated
  • Models the gate as pass or fail only, with no third state
  • Blames flaky CI and re-runs without recording anything
  • Assumes timeouts hit random changes rather than busy ones

context