skip to content

gator test in your rule repo's CI exits 0 on a manifest you know violates the constraint — why?

level: middleimportance: nice to knowfreq 30%

answer

  1. green does not mean enforced
  2. check what was actually loaded
  3. the constraint may be in dryrun
  4. the match block may select nothing
  5. did the step even keep the status

basics

~20 s

Usual causes: the constraint is set to dryrun or warn, so violations print without failing the exit code; the rule files were never in the input set; the match block selects nothing; or the CI step discards the status.

solid answer

~50 s

Work outward from the exit code. First, `enforcementAction`: only a deny-enforced violation makes `gator test` exit non-zero, so a constraint left at `dryrun` or `warn` reports the violation on stdout while the job goes green — read the output, not just the status. Second, the input set: if a glob missed `template.yaml` or `constraint.yaml`, no rule was loaded and there was nothing to violate. Third, matching: check `spec.match` for the fixture's kind and apiGroup, and remember a `namespaceSelector` can only match if the Namespace object is itself in the input, and a `namespaces` list will not match a fixture with no namespace set. Fourth, the CI wiring — a pipe or a tolerated failure can swallow the status. Structured output (`-o json`) makes the difference between "no violations" and "no rules loaded" visible instead of guessed at.

go deeper

for a junior

Know that a green run can mean 'nothing was checked' as easily as 'everything passed', and that the first thing to confirm is that the rule files were actually included in the input.

for a middle

Be able to give the ordered checklist — enforcement action, loaded rules, match block, rule result, propagated exit status — and explain why each one produces exactly the same green output.

for a senior

Show how you make the failure modes observable: structured output asserted on by the job, a canary fixture that must always violate, and a deliberate red run before you trust the gate at all.

for a principal

Own the standard that a control counts as coverage only once it has been demonstrated to fail. A gate nobody has watched go red is a reporting line, not a guardrail, and that distinction belongs in how you count controls.

## Why this is worth being systematic about A gate that silently passes is worse than no gate, because it is recorded as coverage. The specific hazard with an offline policy runner is that four completely different situations look identical from the outside: everything is fine, nothing was loaded, nothing matched, and nothing was allowed to fail. Diagnose them in order. ## 1. The constraint is not enforcing deny `gator test` returns a non-zero exit status for violations of a constraint whose `enforcementAction` is `deny` — the default. A constraint set to `dryrun` or `warn` still has its violations reported in the output, but they do not by themselves fail the command. This is exactly right in a cluster, where dry-run is how you roll a rule out without blocking anyone, and exactly surprising in CI, where the same file gets tested. If the rule repository ships constraints in dry-run and expects the pipeline to fail on them anyway, the pipeline has to read the reported violations rather than the exit code. ## 2. Nothing was loaded The policy is whatever YAML you passed. A `-f` pointing at a directory that does not contain the template, a glob that missed a file, a file with the wrong extension, a document that is present but is not the `ConstraintTemplate` you thought — any of these leave you evaluating an object against an empty rule set. There is no violation because there is no rule, and "no rule" produces the same green as "compliant". This is the single most common cause when a check has *never* failed, as opposed to one that stopped failing. ## 3. Nothing matched The constraint's `spec.match` decides which objects are in scope, and a mismatch is silent by design: - `kinds` with the wrong `apiGroups` — a `Deployment` is in group `apps`, not the core group, and an entry that gets that wrong matches nothing. - `namespaces` or `excludedNamespaces` against a fixture whose manifest sets no `metadata.namespace` at all — a very common shape for a hand-written test fixture. - `namespaceSelector`, which can only be evaluated if the `Namespace` object with those labels is part of the input, because offline there is no API server to look it up. - `labelSelector` against a fixture that never had the labels your production workloads carry. And the broadest version of this: the rule matches `Pod` while the fixture is a `Deployment`, which is a matching failure with its own remedy. ## 4. The rule produced no result A violation rule that never yields a result is indistinguishable from a compliant object, and a rule body can stop yielding for reasons that are not a decision — a field it indexes is absent from this particular object, or a parameter it reads was not set on the constraint. If the template and constraint are loaded and the match block is right, this is where to look next. ## 5. The job swallowed the status Finally, the least interesting and most common cause of all: the command's exit status never reached the job. It was piped into something else, wrapped in a tolerant construct, or run in a step configured to continue on failure. Verify by deliberately breaking a fixture and confirming the pipeline goes red — a gate you have never watched fail is a gate you have not tested. ## Making the difference visible Two habits remove most of this class of problem: - **Emit structured output** (`-o json` or `-o yaml`) and have the job assert on the content, not only the status. "Zero violations from two loaded constraints" and "zero violations from zero loaded constraints" then look different. - **Keep a canary fixture.** Commit one manifest that must always violate, and assert that it does. If the input set, the match block or the rule body quietly stops working, the canary case fails and tells you so, instead of the whole suite passing for the wrong reason. A test harness with expected-failure cases — a suite with deny cases — is the same idea expressed properly. ## The shape of the answer an interviewer wants Not a list of guesses, but an ordered narrowing: prove a rule was loaded, prove the object was in scope, prove the verdict, prove the status was propagated. Each step has a cheap check, and each corresponds to a different bug with a different fix.

  • How do you tell 'no violations' apart from 'no rules were loaded'?
    Ask for structured output and inspect it rather than reading only the exit status, and keep a committed canary fixture that must always violate. If the canary stops failing, something in the input set, the match block or the rule body has quietly stopped working, and you learn it on that commit rather than months later.
  • Why does a constraint set to dryrun not fail the command?
    Because the exit status mirrors enforcement: only a deny-enforced violation is a rejection, and dryrun exists precisely to observe without blocking. The violations are still reported in the output, so a pipeline that wants to fail on dry-run findings has to read the report rather than the status code.
  • What match-block mistake bites fixtures specifically rather than real objects?
    Namespace-based matching. Hand-written fixtures often set no `metadata.namespace`, so a `namespaces` list never selects them, and a `namespaceSelector` cannot be evaluated at all unless the labelled Namespace object is itself part of the input — there is no API server offline to look one up.

saying these in an interview costs you the question

  • Reads exit code 0 as proof no violation was reported
  • Forgets dryrun and warn constraints do not fail the run
  • Never checks whether the template was in the input set
  • Assumes a fixture with no namespace matches every match block
  • Ships a gate without once watching it go red

context