A Helm chart passes conftest on your laptop but the CI policy gate blocks the same commit - how do you diagnose it?
answer
- the rule is deterministic, the input is not
- same chart, different rendered output
- values files, set flags, subchart versions
- publish the rendered manifests from CI
- one render command for both runs
basics
~20 sThe two runs evaluated different documents. Compare the rendered output, not the chart: the values files, set flags, release name and subchart versions used by each render decide whether the fields the rule wants exist at all.
solid answer
~50 sThe rule did not change between the two runs; the artifact did. conftest judges whatever manifest stream it is handed, so the first move is to get both rendered outputs and diff them - have CI publish the rendered manifests and the policy report as build artifacts, then render locally with the exact same command. The usual causes are all render inputs: CI renders with the production values file while you rendered with the chart defaults, so a `resources` block, a readiness probe or a PodDisruptionBudget that your values enable is simply absent in theirs; CI passes extra `--set` flags or a different release name and namespace; or the subchart versions differ. A close cousin: you never rendered at all, so the rules matched nothing and exited green. The fix is structural: one render command, defined once, used by both the pipeline and the local check.
go deeper
Know that the check runs over rendered manifests, so the values file used decides the result. If local and CI disagree, the outputs differ - not the rule.
Be able to list the render inputs that change the output: values files and their order, set flags, release name and namespace, and resolved subchart versions. Reproduce the pipeline's exact command.
Demonstrate the full diagnosis: publish and diff both rendered streams, then collapse the two renders into one shared command. Show you also fix the developer experience - the judged document, the rule message, the reproduction command.
Own where the finding is routed when the violation comes from values another team owns, and defend a deliberate warn-on-incomplete, block-on-rendered split rather than letting a green partial run pass for assurance.
## Start by rejecting the framing "The gate is flaky" is almost never true here, and neither is "the rule is wrong". A Rego rule over a fixed document is deterministic. If the same commit passes in one place and fails in another, **the two runs did not evaluate the same document**. The chart is shared; the rendered output is not. So the whole diagnosis is: reconstruct both inputs and diff them. ## The diff, concretely Ask the pipeline for two things it should already be publishing: the **rendered manifest stream** and the **policy report** (which rule, which document, which field). Then render locally with the pipeline's exact command and `diff` the two streams. Nine times out of ten the answer is visible in the diff within seconds - a container with no `resources` block, a missing `readinessProbe`, no `PodDisruptionBudget` document at all. If the outputs differ, walk the render inputs in this order: - **Which values files, and in what order.** `-f values.yaml -f values-prod.yaml` is not the same as either alone, and later files win. A developer testing with dev values gets probes and limits that the production values file leaves out - or the reverse. - **`--set` flags injected by the pipeline.** Image tags, replica counts, feature toggles. These are frequently generated by the CI job and invisible in the repo. - **Release name and namespace.** Templates branch on them more often than people expect, and names derived from the release appear in selectors and labels. - **Subchart dependencies.** The rendered output includes whatever versions were pulled into `charts/`. A stale local `charts/` directory versus a fresh `helm dependency build` in CI is a real and very confusing source of divergence. - **Capability-dependent templates.** Charts that branch on the target cluster's API versions render differently depending on what the render was told about the cluster. If the outputs are *identical* and the results still differ, then it is the policy side: a different policy version or bundle, a rule that reads an environment variable or a data document, or - most often - a local run that pointed at the chart directory instead of the rendered stream and therefore matched nothing. Undefined rules produce no messages, so a run over the wrong input is green, not red. ## The developer's chair Whoever is blocked at 5pm did not write the values file that made the rendered output non-compliant, and telling them "policy failure: memory limit missing" over a document they cannot see is what makes a gate hated rather than trusted. A gate that blocks should hand back three things, every time: 1. **The rendered document it judged**, published as an artifact, with the offending path (`spec.template.spec.containers[0].resources`). 2. **The rule identity and its message**, in the developer's vocabulary, not the engine's. 3. **The command to reproduce it locally**, byte-identical to what the pipeline ran. If those three exist, the 5pm block costs ten minutes. If they do not, it costs the evening and the next argument about whether policy should exist at all. ## The structural fix Stop letting two places construct the render. Put the render behind one command - a script or a make target that names the values files, the release name and the dependency build - and call it from both the pipeline and the local pre-commit path. Now "works on my machine" is a claim someone can falsify in one step, and a change to the render inputs is a reviewable diff rather than a pipeline edit nobody sees. There is a legitimate two-run pattern worth naming. A cheap un-rendered pass - schema and lint over the chart's own files - can run early and only **warn**, because it is provably incomplete. The rendered pass is the one that **blocks**, because it evaluates the documents that will actually be applied. What is not acceptable is the accidental version of that split, where the un-rendered run is green and everyone believes it means something. ## When the values file is not yours to see Sometimes production values live in another repository owned by the platform team, and the developer genuinely cannot reproduce the render. That is not a reason to gate on the dev render instead - the dev render tests nothing about what deploys. It is a reason for the pipeline to publish the rendered output (with secrets excluded) and for the platform team to own the fix when the violation comes from their values, not the developer's chart change. Route the finding to whoever owns the input that caused it; blocking the person who happened to push is what turns a control into a tax.
- What do you change so the two runs cannot diverge again?Define the render once - a script or make target that pins the values files and their order, the release name and namespace, and runs a dependency build - and call it from both the pipeline and the local check. Then have CI publish the rendered stream and the policy report as artifacts, so any future disagreement is a diff rather than a debate.
- The production values file belongs to another team and the developer cannot render with it. How do you handle the block?Keep gating the production render, because that is the document that deploys, but publish it as an artifact so the developer can read what failed. If the violation comes from the values rather than the chart change, route the finding to the team that owns the values. Blocking whoever pushed for a defect in someone else's input is what makes gates resented.
- Is it defensible to run the check twice - once un-rendered and once rendered - with different outcomes?Yes, if the split is deliberate. The un-rendered pass is provably incomplete, so it warns; the rendered pass evaluates the documents that will be applied, so it blocks. The failure mode to avoid is the accidental version, where a green un-rendered run is quietly treated as coverage.
saying these in an interview costs you the question
- Calls the policy engine non-deterministic or flaky
- Re-runs the pipeline hoping the result changes
- Adds an exemption before comparing the two rendered outputs
- Switches CI to render with dev values to get green
- Assumes CI renders with the same values by default
- Never asks CI to publish the manifest that was judged