skip to content

Infrastructure-as-Code repositories are usually tested at several levels, from cheap file checks up to real deployments. What are those levels, and what does each one catch that the cheaper level below it cannot?

level: middleimportance: must knowfreq 62%

answer

  1. cheap checks first, real cloud last
  2. each level sees more reality
  3. the file, then the diff, then the API
  4. only an apply hears the provider say no

basics

~20 s

IaC testing layers, cheap to expensive: static checks on the files, assertions on the generated plan or diff, unit tests of modules against faked providers, then a real apply into a throwaway account. Each layer sees more reality at more cost.

solid answer

~50 s

I think of it as a pyramid. At the bottom are static checks that only read the files: parse and schema validation, linting, formatting, and security or secret scanning. Above that are checks on the **generated plan** — the tool has resolved variables and read existing infrastructure, so I can assert on the intended diff: nothing is being destroyed, tags are present, this instance class is not appearing in production. Above that sit module unit tests, where a faked or mocked provider lets me run the module across many input combinations in seconds without touching a cloud. At the top, a real apply into a throwaway account or project is the only level where the provider actually answers: quotas, permissions, invalid combinations the schema allows, and slow eventual consistency all surface only here. The rule of thumb is that each level up sees more reality, takes longer, and costs money, so you run fewer of them and less often.

go deeper

for a junior

Be able to name the levels in order and say plainly that cheap checks run on every commit while real deployments run rarely. Knowing that a linter reads only the file is enough here.

for a middle

You should explain each level's blind spot without prompting: no existing infrastructure at the static level, no provider answer at the plan level, no provider at all when it is faked. Interviewers listen for that reasoning, not the list.

for a senior

Show which level you would gate a merge on, which failures you have actually seen escape each level, and how you keep a slow suite trustworthy — flaky teardown, orphaned resources, and alerting on the scheduled run.

for a principal

Own the economics: what assurance each level buys per dollar and per minute of engineer waiting time, where to invest when the estate grows, and when to accept that a whole level is not worth maintaining for your kind of infrastructure.

## Why infrastructure code gets its own pyramid A test pyramid for application code is shaped by *speed*: many fast unit tests, fewer slow end-to-end tests. For infrastructure code the shape is the same but the forces are different. Infrastructure code is mostly declarations, so there is less logic to unit-test; the interesting failures live in the gap between what you declared and what a real cloud provider will accept. That means the expensive top of the pyramid is more valuable than in application testing — and simultaneously more painful, because it costs real money, real minutes, and a real account. The levels, from bottom to top: ``` real apply into a throwaway account slow, costs money, catches provider reality module unit tests (faked provider) seconds, catches logic across many inputs assertions on the generated plan catches the intended change before it happens static checks on the files instant, catches shape and obvious risk ``` ## Level 1 — static checks on the files This level only reads the source. It can parse the configuration, validate it against the tool's schema, flag unknown or misspelled attributes, enforce formatting and naming conventions, spot deprecated constructs, and pattern-match risky declarations — a storage bucket declared public, an unencrypted volume, a credential committed as a literal. Its blindness is structural: it has no idea what exists today, and no idea what values the variables will take. A configuration that is perfectly valid text can still be an illegal request, a duplicate name, or a change that destroys production. Cheap enough to run on every save and every commit. ## Level 2 — assertions on the plan or diff Here the tool has already resolved inputs, read the recorded state, and queried the provider for what exists, so it can produce the concrete set of changes it intends to make. Asserting against that artifact is enormously more powerful than asserting against the source, because you are now testing the *change*, not the text: is anything being deleted or replaced, how many resources are affected, does the resulting configuration carry the mandatory tags, is the machine size within the approved list. This is the level with the best value-per-second in the whole pyramid, and it is what most mature pipelines gate a pull request on. Its limit is that a plan is a prediction. The provider has not been asked to do the work yet, so anything the provider decides at apply time — quota, permission, a combination the schema allows but the API rejects — is invisible here. ## Level 3 — module unit tests against a faked provider When a module has genuine logic — conditionals, loops, computed names, defaults that interact — you want to run it across many input combinations without paying for infrastructure. Substituting a fake for the provider lets the tool build its plan against invented responses, so a test can assert "with this input the module produces three subnets with these names" in a second or two. These are fast and deterministic, and they are the only economical way to cover the combinatorial edges of a widely reused module. They also test only your code: a fake provider always says yes, so they can never tell you the provider would have refused. ## Level 4 — a real apply into a throwaway environment Create the resources for real in an isolated account, project or subscription, assert something about what came up, then destroy it. This is the only level that sees the provider disagree — quota exhaustion, insufficient permissions, mutually exclusive arguments, resources that take twenty minutes to become ready, or a delete that fails because something else attached itself. Many teams add a thin behavioural check on top: can I actually reach the endpoint, does the database accept a connection, does the health check pass. That validates the *system*, not just the declarations. The price is the whole reason the pyramid exists. A run takes minutes to an hour, costs money, consumes account quota, and can leave orphaned resources behind when teardown fails midway. ## What this buys you in an interview The answer that lands is not the list of levels — it is naming the blind spot of each. Static checks cannot see existing infrastructure. Plan assertions cannot see the provider's answer. Faked-provider tests cannot see the provider at all. Only a real apply can, and that is exactly why it runs on a schedule rather than on every push.

  • Where does a manual code review sit in that pyramid?
    Alongside the bottom levels, not instead of them. Review is good at intent — is this the right architecture, is this blast radius acceptable, should this resource be replaced at all — and bad at the mechanical checks a linter or plan assertion does perfectly every time. Teams that rely on reviewers to spot missing tags or a public bucket eventually miss one; automate those and spend review attention on judgment.
  • If your modules are almost pure declarations with no conditionals, is the unit-test level still worth building?
    Often not. With no branching there is little logic to exercise, and a faked-provider test mostly re-asserts what the file already says, which rots on every edit. In that case invest in the plan-assertion layer and one periodic real apply instead. Unit tests earn their keep when a module is reused widely and has real input-dependent behaviour.
  • How do you keep the expensive top level from becoming permanently red and ignored?
    Treat it as a product: run it on a fixed schedule, alert an owner rather than a channel nobody reads, quarantine known-flaky cases explicitly instead of letting the whole suite be amber, and make teardown failures visible with their own alarm. A suite everyone has learned to ignore provides no assurance while still costing money.

saying these in an interview costs you the question

  • Claiming static linting can catch a destroy-and-recreate
  • Treating a plan as proof the apply will succeed
  • Assuming mocked provider tests validate real cloud behaviour
  • Wanting a full apply test on every single commit
  • Describing IaC testing as just running the tool twice

context