A policy rule that calls a registry and live API discovery fails one run in ten. How do you make it deterministic?
answer
- same commit, different answer
- the world changed, not the change
- split I/O from the decision
- pin the facts, evaluate offline
- retry the fetch, not the verdict
basics
~20 sSplit acquiring the external data from deciding on it. One step fetches and pins the registry and discovery facts, the rule then evaluates that pinned document offline, and a failed fetch is reported as an error, never as a verdict.
solid answer
~50 sThe rule is not deciding on the change alone, it is deciding on the change plus whatever a network call returned at that instant, so the same commit can be allowed, denied, or error out from run to run. I would make evaluation a pure function of its inputs: one step resolves the external facts, the digest and signature metadata from the registry, or the list of API versions removed in the target cluster version, with its own retries and timeout, and writes them into a pinned input document; the rule then evaluates that document offline. A failed resolve surfaces as an infrastructure error, never as an allow and never as a deny. The moment errors and denials look alike, re-running until green becomes rational for developers, and a real violation clears on the third attempt just as a transient failure does.
go deeper
Know that a check which passes and fails on the same unchanged commit is telling you something about the check, not about the commit.
Be ready to name every external call a rule makes and to split acquiring data from deciding on it, so the verdict becomes a function of inputs you can show and replay.
Explain how the pinned data stays honest: how it is refreshed, how its age is measured and alerted, and how you would demonstrate that a given commit evaluates the same way twice.
Own the standard that a verdict must be reproducible before a rule is allowed to block, and the consequences: rules that cannot meet it get funded, narrowed, or kept out of the blocking set.
## Where the non-determinism comes from A gate rule is only trustworthy if the same input yields the same verdict. Two common rules break that on purpose: - **Image verification** has to reach a registry to resolve a tag to a digest, or to read the metadata it will judge. That is an outbound call over a network, to a service that rate-limits, has outages, and answers differently over time. - **A deprecated API version check** wants to know which API versions are removed in the cluster version being targeted, which tempts the author to query live discovery from inside the rule. In both cases the verdict depends on the state of a remote system at evaluation time. The change did not vary; the world did. So the check flips between red and green on identical commits, which is the definition of flaky. ## Why flakiness is worse for a gate than for a test A flaky test costs time. A flaky gate costs the control, for three reasons. **It trains people to retry.** Once a red result is sometimes meaningless, the cheapest response to any red result is to press the button again. That is a rational habit and you cannot argue people out of it. It also applies to genuine denials, which now clear on the third attempt exactly like transient failures do. **It destroys the meaning of the pass rate.** If a run's outcome is partly a coin flip, the fraction of green runs stops measuring compliance and starts measuring dependency availability. **It creates the pressure that leads to a downgrade.** Nobody argues to stop enforcing a rule that works. They argue to stop enforcing one that blocks them at random, and they usually win. ## Separating acquisition from decision The fix is structural, and it is the same shape in every engine and every pipeline: 1. **Resolve step.** A step whose only job is to gather the external facts: resolve the tag to a digest and record what the registry says about it, or load a table of removed API versions for the target platform version. This step is allowed to be flaky, because it is I/O. Give it bounded retries with backoff, a timeout, and its own visible failure result. 2. **Pinned input document.** The resolve step writes a document, and that document is the rule's input alongside the change itself. Store it or attach it to the run so the evaluation can be reproduced later with the same facts. 3. **Pure evaluation.** The rule performs no I/O. It compares fields in the change against fields in the pinned document. Given the same two documents it returns the same verdict tomorrow, on another machine, offline. 4. **Three distinct outcomes.** Allow, deny, and could-not-acquire. The third is an error owned by the platform team, not a policy verdict about the change, and it must not look like either of the others in the pipeline's reporting. ## What you give up, and how to manage it Pinning trades freshness for reproducibility. A table of removed API versions refreshed nightly can be up to a day stale. That is usually the right trade, because staleness is measurable and alertable, in a way that a live call that failed and was retried away is not. Publish the age of the pinned data, alert when it exceeds a threshold, and make refreshing it a reviewed change rather than an invisible side effect of whichever run happened to succeed. Where freshness genuinely cannot be traded away, for example a revocation-style signal that must be current, then accept that acquisition failure is a real operational state and design what happens in it deliberately, with an owner, rather than letting a timeout decide. ## Retries: which layer gets them Retry the fetch, never the verdict. - Retrying the resolve step with backoff is ordinary engineering, and it should be capped and counted. - Re-running a completed evaluation is not a retry, it is a second draw from a distribution, and it lets the first transient failure hide behind a later success, or a genuine deny hide behind a later error. Count retries per rule and alert on the rate. A climbing retry rate is the earliest measurable sign that a gate is about to be routed around, and it shows up weeks before anyone proposes downgrading the rule. ## The one-line diagnosis If a check gives different answers for the same commit, the defect is in the gate, not in the change, and no amount of re-running will convert it into a control. Make the decision a function of inputs you can show, and make the inability to gather those inputs a visible, separately named outcome.
- Teams have learned to re-run the job until it goes green. What does that habit cost you?It erases the difference between a rule that denied a change and a rule that could not evaluate one, because both render as red and one of them clears on retry. So genuine denials get retried too, and the gate's pass rate stops meaning anything. Fix the signal first by separating the error outcome from the deny outcome, so that a retry can only ever clear an error.
- Should the pipeline retry the check automatically?Retry the data-acquisition step, never the verdict. Bounded retries with backoff around a fetch are ordinary engineering; automatically re-evaluating a rule that already returned a deny just hides the deny behind whichever attempt succeeds. Cap the retries, count them per rule, and alert when the rate climbs, because that is the earliest sign the gate is being routed around.
- How do you keep a deprecated-API-version rule accurate without calling live discovery on every evaluation?Ship the removal data as a versioned artifact keyed to the target platform version and refresh it on a schedule through a reviewed change. The rule then compares manifests against a fixed table. You trade a bounded, measurable staleness for a verdict that reproduces, and stale data is visible as an age, whereas a failed live call disappears the moment someone retries.
A referee who has to phone head office before every whistle will make different calls depending on the phone line. Hand them the rulebook printed out and the calls become repeatable.
saying these in an interview costs you the question
- Calls it flaky CI and adds a blanket retry
- Treats an acquisition failure as an allow
- Treats an acquisition failure identically to a violation
- Queries live data inside the rule to stay fresh
- Says the rule is fine because it eventually passes