skip to content

One Engine, Many Formats

One engine that parses YAML, JSON, HCL and Dockerfiles gives every artifact type a single rule language, at the price of a language nobody else on the team reads.

on this pageshow

questions

4

Why does one conftest rule banning plaintext secret env values need per-format logic?

level: middleimportance: must knowfreq 50%

answer

  1. policies read the parse, not the text
  2. instructions array versus nested object
  3. env as a list of objects, or a map
  4. a missing path is undefined, not false
  5. run conftest parse before writing

basics

~20 s

Because conftest policies read the parsed document, not the file text, and each parser produces a different shape. A Dockerfile becomes an array of instruction objects, a Kubernetes container env is a list of name/value objects, and a pipeline env is a plain map - three different paths for one logical field.

solid answer

~50 s

conftest gives every format a single generic document model, but generic does not mean identical: the parser decides the shape. `ENV API_TOKEN=s3cr3t` in a Dockerfile arrives as an element in a flat array of instructions with `Cmd` set to `env` and `Value` holding alternating key/value strings. The same secret in a Kubernetes Deployment lives at `spec.template.spec.containers[_].env[_].value` - and in a bare Pod it is one level shallower. In a CI job definition it is usually just a map of names to values. A rule that only knows one of those shapes passes the other files silently, because a path that does not exist is undefined rather than false. The practical answer is small per-format adapters that normalise each input into a common list of name/value pairs, plus one shared rule that decides whether a pair looks like a secret - and running `conftest parse` on a real file before writing anything.

code

json · 9 lines
json
[
  { "Cmd": "from", "Value": ["node:20"] },
  {
    "Cmd": "env",
    "Value": ["API_TOKEN", "s3cr3t"],
    "Original": "ENV API_TOKEN=s3cr3t"
  },
  { "Cmd": "run", "Value": ["npm ci"] }
]

go deeper

for a junior

Recall that policies see parsed data, not the original text, and that a Dockerfile and a manifest do not look alike once parsed. Knowing there is a command that prints the parsed document is enough at this level.

for a middle

Explain the three concrete shapes - instruction array, list of name/value objects, plain map - and why a wrong path passes silently instead of erroring. Describe the normalise-then-decide split.

for a senior

Demonstrate how you keep one security decision from drifting into three divergent implementations, and insist on must-fail fixtures for every format a rule claims to cover.

for a principal

Own the boundary between the shared judgement and the per-format extraction, so format knowledge can be fixed by anyone while what counts as a secret stays a single reviewed decision.

## The trap "One engine, many formats" is the selling point, and it is true - but it is easy to hear it as "one document shape, many formats", which is false. conftest parses each format into generic JSON-like data, and the **shape of that data is whatever the parser produced**. The logical thing you care about - a secret handed to a process as a plaintext environment variable - lands in a different place in every format it can appear in. ## Three shapes for one idea **Dockerfile.** The Dockerfile parser emits a *flat array of instruction objects*, not a map keyed by instruction. Each element carries `Cmd` (the lowercased instruction, `from`, `env`, `run`) and `Value` (an array of strings). For `ENV API_TOKEN=s3cr3t` the value array holds the key and the value as consecutive entries, so a rule iterates the array, filters on `Cmd == "env"` and walks `Value` two entries at a time. There is no `.env.API_TOKEN` to reach for. **Kubernetes.** A container's environment is a *list of objects* with `name` and either `value` or `valueFrom`, and the path to that list depends on the kind: a Pod has `spec.containers`, a Deployment or Job wraps it in a pod template at `spec.template.spec.containers`, and both may also have `initContainers`. Same API concept, three or more paths, inside the same file format. **A CI job definition.** Here the environment is usually a plain *map* of name to value hanging off a job. Now the key is a map key rather than a `name` field, so even the way you get at the variable's name is different. **Infrastructure code.** An HCL parse gives you blocks nested under their block type and labels. Note that a raw HCL parse shows the *expressions as written*: a value produced by a function call or a variable reference has not been resolved into a literal, which is one reason teams evaluate a machine-readable plan for infrastructure rather than the source. ## Why a wrong path fails silently In Rego, referencing a field that does not exist makes the expression **undefined**, and an undefined expression is not false - it simply produces no result. So a rule written only for the Kubernetes shape, run against a Dockerfile, does not error and does not fail: it produces nothing, the file passes, and the summary happily reports a success. This is the single most common reason a policy author believes a rule is protecting far more than it is. ## How to write it The pattern that survives contact with a mixed repository is **normalise, then decide**: 1. Write one small helper per format that extracts `{name, value}` pairs from that format's shape. Three tiny adapters, each obviously correct by inspection. 2. Write one shared rule that answers the actual policy question: does this name look like a credential (a `TOKEN`, `SECRET`, `PASSWORD`, `KEY` suffix or similar), and is a literal value present rather than a reference to a secret store? 3. Keep those adapters in format-specific namespaces so a repository with no Dockerfiles is not evaluating Dockerfile rules at all. Splitting it this way means the security decision - what counts as a secret, and what the acceptable alternative is - is written once and reviewed once, while the fiddly shape knowledge is isolated where it can be fixed without touching the judgement. ## Find the shape, do not guess it `conftest parse <file>` prints exactly the document the policy will see, for any supported format. Running it on a real example is the first step of writing any rule, and it is the fastest way to settle an argument about whether a field is a list or a map. Pair every rule with fixtures: a file that must fail and a file that must pass, for **each** format the rule claims to cover. A rule with only a passing fixture proves nothing, because a rule that matches nothing also passes it. ## The wider lesson The rule family here - "no plaintext credentials in environment variables" - is a good illustration precisely because it is one sentence of policy that must hold over an image build, a workload spec and a pipeline definition. When a control is expressed once in English and three times in code, the three implementations drift. Keeping the decision shared and the extraction per-format is what stops the Dockerfile version from quietly meaning something weaker than the Kubernetes one.

  • How do you find out what shape a file actually has before writing the rule?
    Run `conftest parse` on a representative file - it prints the exact document the policy will be evaluated against, for any supported format. Guessing the shape from the source text is how authors end up referencing fields that do not exist, which produces an undefined expression and a silently passing rule rather than an error.
  • Two Kubernetes kinds carry the same env field. Do you really need two rules?
    You need two extractions, not two decisions. A Pod's containers sit at `spec.containers` while a Deployment wraps them in a pod template, and initContainers are a separate list again. Collect containers from all of those into one set with a helper, then apply a single rule to the result so the judgement is written once.
  • How would you prove the Dockerfile half of this rule works?
    With fixtures per format: a Dockerfile that must produce a failure and one that must pass, alongside the equivalent manifest pair, executed as part of the policy repository's own tests. A passing fixture alone is worthless here, since a rule that matches nothing also passes it.

saying these in an interview costs you the question

  • Assumes every format lands the field at the same path
  • Writes the rule against raw file text or a regex
  • Expects a Dockerfile to parse into a map of ENV keys
  • Tests one format and assumes the others are covered
  • Thinks a wrong path errors out rather than passing

context

open as a page

What file formats can conftest test, and how does a CI job learn it failed?

level: juniorimportance: should knowfreq 58%

basics

~20 s

conftest parses structured configuration - YAML, JSON, HCL, Dockerfiles, TOML, INI and more - into one JSON-like document and evaluates policies against it. Failures print to the console and the process exits non-zero, which is the signal CI acts on.

open as a page

Your conftest gate has been green for months, but the rule never ran. How?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A conftest run with nothing to evaluate exits zero. The usual causes are a namespace never selected on the command line, a renamed policy package, a glob that missed the files, or an extension with no parser attached - all of which look identical to a clean pass.

open as a page

What does conftest's --combine flag change about the input a rule sees?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Without it, each file is evaluated on its own and the input is that file's document. With --combine, all files are merged into one evaluation where the input is an array of elements carrying a path and the file's contents, so rules must iterate rather than address fields directly.

open as a page