skip to content

When the Gate Bends

A gate that is slow, noisy or wrong gets routed around, downgraded to a warning, or retried until green. Interviewers ask because the failure mode is social long before it is technical.

on this pageshow

explore

questions

11

In a CI pipeline, why does a green policy check not prove the gate actually evaluated that commit?

level: juniorimportance: must knowfreq 55%

answer

  1. green summarises failures, not runs
  2. a skipped job still reports success
  3. no rule matched looks like compliant
  4. absence of a failure is not a result
  5. a pass must be asserted, not inferred

basics

~20 s

Green means nothing failed, not that something ran. A skipped job, a path filter or a check nobody required all look successful. Without a record asserting an evaluation happened, passed and never invoked are indistinguishable.

solid answer

~50 s

A pipeline status summarises what failed; it is not an inventory of what ran. The policy job can be skipped by a path filter, excluded by a branch condition, deleted from the pipeline definition in the very change it was meant to gate, or it can run, load zero rules and exit successfully. Every one of those ends green. The same ambiguity exists inside the engine: when no rule matches a resource, most engines emit no violation, and no violation renders exactly like compliant. So green supports one claim only, that nothing reported a failure. To get further, the evaluation itself has to emit a record tied to the commit saying which rules ran against which input and what they decided. Then the pipeline status is a cross-check rather than the evidence, and a commit with no record is visibly unevaluated instead of quietly assumed fine.

go deeper

for a junior

Be ready to say what green actually asserts: that nothing reported a failure. Name at least two ways a check reports success without evaluating anything, such as a skipped job or a path filter.

for a middle

An interviewer expects the mechanics of the false green: conditions and filters in the pipeline, a rule set that loaded empty, a rule whose selector matched nothing. Explain why each renders identically to a clean pass.

for a senior

Show how you would make absence detectable in a running system: emit a decision per evaluation, require its presence downstream, fail closed on an empty rule set. Say what you would tell an incident review that holds only a green build.

for a principal

Own the framing that a build status is not evidence. Decide what your organisation treats as the authoritative record of a gate decision, and accept the cost of every pipeline emitting one instead of inferring outcomes from build history.

## The four states, and the one nobody records When you go back to a commit months later and ask what the gate did, there are four possible answers: it ran and passed, it ran and blocked, it ran and was overridden, or **it never ran at all**. Continuous integration systems are built to surface the second one loudly and are indifferent to the difference between the first and the fourth. A build result is an aggregation of failures. If nothing failed, the build is green — and *nothing ran* is a perfectly good way for nothing to fail. That fourth state deserves its own name. Call it **indeterminate**: the evidence you hold cannot distinguish a passing evaluation from an evaluation that never happened. Indeterminate is not a rare pathology; it is the default state of most pipelines, because almost nothing in a normal setup writes down the positive fact that a rule was applied. ## Concrete ways a check goes green without checking anything - **The job was skipped.** A path filter says the job only runs when files under `infra/` change; the datastore was created by a module bump elsewhere, so the job never started. A skipped job contributes no failure. - **A condition excluded it.** The job runs only on the default branch, or only when a label is present, or was disabled for a release freeze and never re-enabled. - **The check was never required.** The job ran and failed, but nothing was configured to block the merge on it, so the change went in with a red check that no one was watching. - **The gate was removed in the same change.** Pipeline definitions usually live in the repository they gate, so the change that introduces a violation can also be the change that stops the violation being checked. - **The engine loaded no rules.** A moved directory, a bad path, a rule bundle that failed to download and was skipped with a warning: the engine evaluates the input against an empty rule set, finds nothing to complain about, and exits zero. - **No rule matched.** The rule set is fine, but nothing in it selects the resource type in front of it. An unmatched rule produces no result, and an absence of results is not a pass. That last one is the subtle one, and it is where policy engines differ from tests. A test that does not run is usually reported as skipped. A rule that does not match is usually reported as nothing at all. ## Why this matters to someone reading the record later Suppose an unencrypted datastore turns up in production and the encryption-at-rest rule has existed for a year. The post-incident question is narrow: on the commit that created that datastore, did the rule run? If the only artifact is a green build, the honest answer is that you do not know. You cannot credit the gate, you cannot blame it, and you cannot say whether the same hole is open on the other four hundred changes that shipped that quarter. The gate might be working perfectly and simply not have covered this path; it might have been silently inert since February. Both stories fit the evidence equally well. The same applies to a minimum retention rule on an audit log. A year of green builds is consistent with a year of enforcement and equally consistent with a rule that stopped being invoked in week three. ## What turns green into evidence The fix is to stop *inferring* outcomes from build history and start *asserting* them. In practice that means three things. 1. **Emit a decision, not a status.** Every evaluation writes a record keyed to the change it judged, stating which rules were evaluated, against which input, and what the outcome was. Passed becomes a positive statement rather than the absence of a complaint. 2. **Make absence loud.** Something downstream — the merge check, the deploy step, a reconciliation job that walks merged commits — requires that record to exist. The moment a missing record fails something, unevaluated changes stop being invisible. 3. **Fail closed on a degenerate run.** If the engine loaded zero rules, or evaluated zero resources, that is an error, not a pass. A gate that cannot detect its own emptiness will report green forever. None of this requires a heavyweight system; it requires that the evaluation say what it did rather than leaving a reader to guess from the colour of a build. ## The replay trap A tempting shortcut is to re-run the check on the old commit today. That answers a different question. A replay evaluates that code against today's rules, today's engine version and today's surrounding data, and tells you whether the change would pass **now**. It cannot tell you whether anything was applied then. Reconstruction has to come from a record made at the time; anything else is a new decision wearing an old commit's clothes.

  • Your engine loaded zero rules because of a bad path and the job exited successfully. How would anyone notice?
    Only if the run asserts what it did. The record should carry how many rules loaded and how many resources were evaluated, and the job should fail when either is zero. A gate that cannot fail closed on an empty rule set is decorative: it will report green for every change until someone happens to feed it a known-bad input.
  • Six months on, can you just re-run the pipeline on that commit to find out whether it would have passed?
    That answers a different question. Re-running evaluates the old commit against today's rules, today's engine and today's surrounding data, so it tells you whether the change would pass now, not whether it was gated then. Reconstruction has to come from a record written at the time; a replay is a new decision, not the old one.
  • What is the cheapest change that makes never invoked look different from passed?
    Have each evaluation emit one record keyed to the change, and make something downstream require that record to be present. Absence stops being silent the moment a missing record fails the merge or the deploy, and you get an enumerable list of changes with no decision behind them.

saying these in an interview costs you the question

  • Treats a green pipeline as proof the rule was evaluated
  • Assumes no violations reported means the change complied
  • Thinks a required status check guarantees the job actually ran
  • Believes re-running the check today reconstructs the past decision
  • Confuses a build status with a decision record

context

open as a page

A CI policy check times out and the pipeline records it as passed. Why is that unsafe?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A timeout means the rule produced no decision at all, not that the change is compliant. Recording it as a pass turns an overload into a silent approval, and that happens most often exactly when the system is busiest.

open as a page

What is a severity threshold in a pipeline security gate, and what does it actually decide?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A severity threshold is the rule that collapses a scan report into pass or fail: block when a finding sits at or above a chosen severity, such as critical. It decides what stops a change, not what is genuinely risky.

open as a page

A policy rule that calls a registry and live API discovery fails one run in ten. How do you make it deterministic?

level: middleimportance: must knowfreq 54%

basics

~20 s

Split acquiring the external data from deciding on it. One step fetches and pins the registry and discovery facts, the rule then evaluates that pinned document offline, and a failed fetch is reported as an error, never as a verdict.

open as a page

In a security findings gate, how does blocking on new findings differ from blocking on total findings?

level: middleimportance: must knowfreq 60%

basics

~20 s

Blocking on total findings fails every build until the entire backlog is clean. Blocking on new findings compares each report against a recorded baseline of what already existed and fails only on what the current change added.

open as a page

The only proof your encryption-at-rest gate passed is a file in the gated team's own repo. What does it prove?

level: seniorimportance: must knowfreq 45%

basics

~20 s

That someone committed a file saying pass. A record the gated party can write, edit, regenerate or delete carries almost no weight about whether the gate ran, and a missing file cannot be told apart from a deleted one.

open as a page

The policy gate is your incident: jobs are queuing and every retry goes green on the third attempt. What do you do in the first hour?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Make an explicit, time-boxed decision instead of letting timeouts make it. Use the gate's telemetry to find the slow rule and its dependency, narrow that one rule under a named owner and expiry, and record what shipped un-evaluated.

open as a page

Your zero-criticals gate blocked a one-line change over a pre-existing finding. What do you change?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Rescope the rule from the absolute state to what the change introduced: fail on findings absent from a recorded baseline, and move the pre-existing critical into owned, dated backlog work. Do not simply raise the threshold until the pain stops.

open as a page

A record says your audit-log retention rule passed six months ago. What must it capture to still mean anything?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

The identity of what was judged and what judged it: the input the engine actually saw, the rule set version in force, and an explicit decision. Passed means little unless you know which rule passed it.

open as a page

A regenerated findings baseline let a real regression pass a green gate. How did that happen?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Regenerating a baseline overwrites it with whatever the current scan reports, so a finding the branch just introduced is recorded as pre-existing and stops counting as new. The gate then passes truthfully while a real regression ships.

open as a page

Half your org's blocking policy rules are now warn-only and nobody decided that. How do you fix it?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Give de-enforcement the same ceremony as enforcement: every downgrade carries an owner, a reason and an expiry, and each rule's live enforcement state is inventoried. Findings routed to a channel with no consumer are not a control.

open as a page