Your local replay allows a CI job the gate denied. What could differ between the two evaluations?
answer
- evaluation is deterministic
- so an input must differ
- recorded document versus regenerated one
- rules ship continuously; pin the revision
- live reference data and request context
basics
~20 sOne of the decision's inputs differs. Either you replayed a rebuilt document rather than the recorded one, a different rule revision, or reference data that has since changed. Chase the difference; do not call the gate flaky.
solid answer
~40 sA policy verdict is deterministic given its inputs, so a disagreement means an input differed, and that difference is the finding rather than a nuisance. I check in order. First the document: did I replay the recorded evaluated input, or regenerate it, in which case rendering, expansion or defaults may have produced a different shape. Second the rules: the set ships continuously, so pin the revision the decision used rather than whatever is deployed now. Third any external data the rule consults, such as the approved-destination list, which is looked up live and may have moved. Fourth, context that only exists at the gate, such as the requesting identity or the target environment, which my laptop supplies differently or not at all. Only if all four match would I suspect the engine itself.
go deeper
Know that the same input under the same rules always gives the same answer, so a replay that disagrees with the gate means something you fed it was different. Do not reach for the re-run button.
Be able to list the candidate differences in likelihood order - regenerated document, unpinned rule revision, live reference data, missing request context, engine configuration - and say how you would test each one.
Demonstrate that you convert each cause into a platform fix: publish the evaluated input, pin revisions in the replay tool, snapshot the data the rules consult, and let the harness take request attributes.
Be ready to argue that tolerating unexplained disagreements is how retry-until-green becomes culture, and to fund the reproducibility work before the first argument about whether a gate can be trusted.
## Treat the disagreement as the signal When a local replay allows what the gate refused, the tempting conclusion is that the gate is flaky and a retry will fix it. That conclusion is almost always wrong, and acting on it is how teams learn to retry-until-green. Policy evaluation is deterministic: the same document, the same rules and the same reference data produce the same verdict every time. A disagreement is therefore evidence that one of those is not the same, and finding out which one is the whole diagnosis. ## Difference one: you replayed a different document This is the most common cause by a wide margin. There is a difference between *the input the engine evaluated* and *an input you generated that resembles it*. If you re-rendered the pipeline configuration on your laptop, you may have gotten a different document: a matrix expanded differently, a template resolved to another version, a variable interpolated from your local environment rather than the pipeline's, a default filled in by a different tool version, or a field that only exists once the platform has annotated the job. The test is mechanical: diff the document you replayed against the document the run recorded. If you cannot get the recorded document at all, that is the actual defect to fix - a gate that cannot show what it judged cannot be diagnosed by anyone. ## Difference two: you replayed a different rule set Rules are software and ship continuously. Between the denial and your investigation an hour later, the set may have been updated - possibly by someone fixing exactly the bug you are chasing. Replaying against "current" therefore answers a different question than "why was I denied". Pin the revision the decision used. When the pinned revision denies and the current one allows, you have not found a flake; you have found that the rule was already corrected, which is a useful and reportable outcome. The mirror case matters too: the gate may not have been running the revision you assume. Rule distribution is asynchronous, so different enforcement points can be a few minutes - occasionally much longer - apart, and a stale or partially rolled-out set is a real cause of "it denies here and allows there". ## Difference three: reference data moved Many rules consult data rather than hard-coding values: the list of approved egress destinations, the registry of service owners, the current exception list. That data is usually fetched live at decision time. If the list changed between the denial and the replay - a destination approved, an exception granted, an owner added - the rule is the same but its inputs are not. Reproducing faithfully means snapshotting that data alongside the document, which is a design requirement on the platform, not something the blocked engineer can retrofit at 5pm. ## Difference four: context that only exists at the gate Some rules read attributes that are properties of the request rather than of the document: who is making the change, which environment is targeted, whether this is a dry run, what time it is. On a laptop those are absent or supplied by you, and absent is not the same as false - a rule branch that only fires for the production environment simply will not fire when the environment attribute is missing, and your replay allows quietly. If your replay harness cannot supply that context, the replay is not faithful and you must say so rather than trusting the green result. ## Difference five, last and least likely: the engine Different engine version, different configuration, a different entry point evaluated, a decision short-circuited by a timeout or an error path that the gate treated as allow. These do happen, and a gate that treats an evaluation error as a pass is worth knowing about, but they belong at the bottom of the list because they are rare relative to the four above. ## The discipline Work the list top-down and stop at the first difference that explains the outcome. Write down which of the five it was, because the answer determines what you fix: a document difference means the gate must publish its evaluated input; a rule-revision difference means your replay tooling must pin; a data difference means snapshots; a context difference means the harness needs to accept those attributes. "It was flaky" fixes nothing and quietly licenses the next person to retry until green.
- Why is 'the gate is flaky, just re-run it' a dangerous conclusion here?Because it is usually false and it is self-reinforcing. Evaluation is deterministic, so a disagreement means an input differed; declaring it flakiness leaves that difference in place and teaches the team that retrying is a legitimate response to a denial. Once retry-until-green is normal, a real violation eventually gets through and nobody notices, because the same behaviour covers both cases.
- How does a rule that reads request context, such as the target environment, break a naive replay?That attribute is a property of the request, not of the document, so it is simply absent on a laptop. A branch guarded on it never fires, the replay allows, and you conclude the rule is fine. A faithful harness has to accept those attributes explicitly, and if it cannot, you report the replay as inconclusive rather than as a pass.
saying these in an interview costs you the question
- Declares the gate flaky and re-runs until it passes
- Replays a regenerated document instead of the recorded one
- Evaluates against the current rules rather than the revision that denied
- Forgets that rules consult live external data
- Treats missing request context in a replay as equivalent to false