skip to content

Your org already runs an IaC scanner — when do you extend it with custom checks rather than adopt a second policy engine?

level: principalimportance: should knowfreq 34%

answer

  1. you inherit more than you write
  2. parser, graph, report, ids, one gate
  3. the ceiling is the scanner's model
  4. which artifacts does the rule need to read
  5. worst outcome is quietly running both

basics

~20 s

Extend the scanner while your rules fit its model and its authors are the people already on the rota: you inherit its parsing, reporting, ids and pipeline placement for free. Adopt a separate engine when rules must span artifacts the scanner cannot parse, or when several gates should share one rule language.

solid answer

~60 s

Default to extending what you already run. A custom check inside the scanner inherits everything that is expensive to rebuild: the parser, the resource graph, the report format your dashboards consume, the id and severity conventions, the pipeline step teams have already been trained on, and one place where suppressions and exceptions are handled. Standing up a second engine buys you rule-language power you probably do not need and costs you two of everything — two sets of findings, two exception paths, two things to keep running, two things a developer has to understand when a build goes red. The case for the second engine is real when the rules you must write are not about the scanner's domain at all: a rule that has to read a manifest, a CI job definition and an attestation together, or the same rule enforced at more than one gate where duplicating it in two dialects is worse than adopting one language. Judge it on where the rule's inputs live and who will maintain it in a year — not on which language is nicer to write.

go deeper

for a junior

Know that a custom rule can live inside the scanner your team already runs, and that adding a whole second policy tool is a bigger decision than it looks because it duplicates reports and pipeline steps.

for a middle

Explain what a custom check inherits from its host scanner — parsing, the resource graph, the report format, the id and severity conventions — and what the ceiling is: you can only express what its check API and parsed model support.

for a senior

Argue the call from operational reality: one gate teams already understand, one exception path, one set of findings, versus the genuine cases where rules must read artifacts the scanner never parses or the same rule must run at more than one gate.

for a principal

Own the organisational consequences. Decide who maintains the rules and who operates any new engine, prove the need with rules you actually tried to write, and prevent the real failure — both approaches running half-maintained with nobody able to say which controls are enforced.

## The decision, stated properly Your organisation runs an infrastructure-as-code scanner in CI. You now need rules its catalog does not contain: only approved base images and AMIs, only these instance families and sizes, only these regions. A team proposes standing up a general-purpose policy engine alongside it, because the rule language is more expressive. You have to decide. The decision is not about language expressiveness. It is about what you inherit, what you must build, and who is on the hook a year from now. ## What extending the scanner gives you for free A custom check inside the scanner inherits its whole apparatus: - **Parsing.** Someone has already solved turning your configuration language into a queryable structure, including modules and nested blocks. That is not a small piece of work, and it is not a piece of work your team gets credit for redoing. - **The resource graph.** Cross-resource questions — is this instance attached to that security group — are answerable because the scanner already built the relationships. - **One report.** Your custom findings arrive in the same output, with the same ids and severities, into the same dashboard, alongside the built-in ones. Nobody has to correlate two reports to know whether a change is clean. - **One exception path.** However your organisation grants an exception, custom checks use the same one. Two engines means two, and the second one is always the one people forget. - **One pipeline step.** Teams have already been taught what this step is and what to do when it goes red. Every additional gate is a new thing to explain, a new source of build latency, and a new thing that can be the outage. - **The authors you already have.** The people who will write and maintain these rules are the ones already on the rota. A rule language nobody else on the team reads is a bus-factor problem dressed up as elegance. ## What extending costs you - **The scanner's model is the ceiling.** You can only express what its check API and condition operators support, over the resource types it parses. A rule that needs to correlate information the scanner does not carry is not writable, and discovering that after you have written six checks is expensive. - **You are coupled to its data shape.** Custom checks read the scanner's parsed representation, so an upgrade that changes that representation can quietly change what your rules match. - **Rules live inside a tool, not in a policy repository.** Distribution — getting the rules to every pipeline that needs them, versioned — is something you now have to solve within that tool's mechanisms. ## When the second engine is genuinely the right call Three conditions, and you want at least one of them to be strongly true: 1. **The inputs are not the scanner's domain.** The rules you must write read manifests, CI job definitions, image metadata, scanner output or attestations — several document kinds, sometimes together. A scanner extension cannot reach outside what the scanner parses. 2. **The same rule must be enforced at more than one gate.** If the rule that runs in CI must also run at another decision point, writing it twice in two dialects guarantees they drift. One rule language across gates is a real, durable win. 3. **You have someone to own the engine.** A policy engine is production infrastructure with an operational story of its own. If nobody owns it, it becomes a gate that fails open when it is unhealthy and nobody notices. Notice that "the rule language is nicer" is not on the list, and neither is "our first custom check was awkward to write". ## Running the decision as a lead Make it concrete rather than architectural. Write the three or four rules you actually need in the scanner you already run. If they come out fine, you are done and you saved the organisation a system. If one of them cannot be expressed, you now have a specific, demonstrable reason for the second engine rather than a preference — and that reason is what gets you the headcount to operate it. Whichever you land on, own the conventions before the rules multiply: an id namespace of your own so custom rules never collide with a vendor's, a declared severity on every rule so findings can be triaged, a name and a remediation pointer written for the developer who gets blocked, one rule per intent so exceptions can be granted narrowly, and a test fixture per rule so a rule that stops firing fails a build instead of failing silently. Those conventions matter more to whether the guardrail survives than the choice of engine does — and unlike the engine choice, they are entirely in your control. ## The failure to avoid The common bad outcome is not choosing wrong. It is choosing both by accident: a few custom checks in the scanner, a few rules in a new engine that one enthusiastic engineer stood up, no clear rule about which gets used for what, and eighteen months later two half-maintained rule sets, two exception paths and no one able to answer which controls are actually enforced.

  • What is the single strongest technical signal that the scanner cannot carry your rules?
    The rule needs inputs the scanner does not parse — correlating a manifest with a CI job definition, or a configuration with a scanner report or attestation. A scanner extension is bounded by what its parser produces; no amount of cleverness in a custom check reaches outside that.
  • A team says custom checks are too awkward to write, so they want a different engine. How do you respond?
    Take the complaint seriously but test it concretely: write the three rules actually needed and see. Awkwardness in the first check is usually unfamiliarity with the parsed data model, which is a day of learning; an unwritable rule is a real architectural reason. Fund the second engine on the second finding, not the first.
  • If you do extend the scanner, what conventions do you fix before writing more than a few rules?
    An id namespace of your own so custom rules never shadow vendor ids; a declared severity on every rule; a name and remediation pointer written for the blocked developer; one rule per intent so exceptions stay narrow; and a test fixture per rule so a rule that stops firing breaks a build rather than going quiet.

Extending the scanner is adding a room to the house you live in; a second engine is buying a second house. Both are defensible, but you should know you are now heating two.

saying these in an interview costs you the question

  • Chooses the engine on rule-language aesthetics
  • Ignores that a second gate needs an owner and an on-call story
  • Forgets that two engines means two exception paths
  • Assumes custom checks cannot do cross-resource logic
  • Lets both approaches run with no rule about which is used

context