skip to content

IaC Concepts

Every IaC tool is a different dialect of the same handful of ideas: desired state, reconciliation, drift, immutability and a preview before you commit. Learning them once here means Terraform, Pulumi, CloudFormation and Ansible all become variations rather than separate subjects.

on this pageshow

questions

page 2 of 2

An infrastructure-as-code preview was generated and reviewed twenty minutes ago, and the apply runs now. What can make the applied result differ from what was reviewed, and how do teams narrow that window?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A preview describes the world at the moment it was computed, so anything that changes afterwards causes skew: another apply, a console edit, an autoscaler, or values that only resolve during apply. Teams narrow it by serializing applies behind a lock, applying the reviewed artifact, and keeping approval-to-apply short.

open as a page

Two engineers start an apply at the same time against the same shared IaC state record. What can go wrong, and why is locking the standard answer?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Each run reads the record, changes real infrastructure, then writes the record back. Interleaved, the later write overwrites the earlier one, so resources the first run created vanish from the record and become unmanaged orphans. A mutual-exclusion lock serialises the write-bearing operation.

open as a page

Some infrastructure defects survive every cheap check and only surface when the configuration is genuinely applied into a live account. Which classes of failure are those, and why can no plan or mocked test predict them?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Only a real apply exposes what the provider decides at request time: quota and limit refusals, missing permissions, argument combinations the schema allows but the API rejects, name collisions, region or capacity unavailability, slow eventual consistency, and dependency-ordered failures including destroys that hang or leave orphans.

open as a page

You inherit an estate where some infrastructure is reconciled continuously by an agent, some by a nightly job, and some only when an engineer opens a pull request. How do you decide which loop each resource should belong to?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide per resource by comparing the cost of being wrong with the cost of being corrected automatically. Reversible, numerous, cheap-to-recreate resources belong on the fast loop; destructive or data-bearing ones stay human-gated. Above all, give every field exactly one owning loop.

open as a page

You own the human approval gate for infrastructure-as-code changes across many teams. How do you decide which changes require an approval, and what makes such an approval meaningful rather than a rubber stamp?

level: principalimportance: should knowfreq 34%

basics

~20 s

Gate on the content of the diff rather than on every change: destructive, stateful, security-relevant and production changes need a human, additive and reversible ones do not. An approval is meaningful only when the approver can read the diff, is accountable for the resource, and can realistically decline.

open as a page

Any set of infrastructure guardrails eventually meets a change that legitimately has to break one of the rules. How would you design the exemption path so that exceptions remain possible without the guardrail degrading into a formality?

level: principalimportance: should knowfreq 32%

basics

~20 s

Make exemptions explicit, narrowly scoped to one rule and one resource, time-bounded with an expiry that re-fails, attributed to a named requester and an independent approver, and visible in aggregate. Blanket skips and permanent unowned exceptions are what hollow out a guardrail.

open as a page

Most teams refuse to run a full apply-and-destroy integration test on every commit to their infrastructure repository. Why, and how would you decide what runs on a pull request versus nightly versus before a release?

level: principalimportance: should knowfreq 45%

basics

~20 s

A real apply costs minutes to hours, real money, and finite account quota, and it fails intermittently for provider reasons. So pull requests get fast, deterministic checks, a scheduled run pays for the expensive real deployment, and pre-release runs the full path once against realistic scale.

open as a page

If infrastructure is written in a general-purpose programming language such as Python or TypeScript, does that make the approach imperative rather than declarative?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

No. The authoring language and the deployment model are separate axes. A program with loops and conditionals can still produce a complete description of desired state, which an engine then compares against reality before changing anything.

open as a page

Infrastructure automation can be written to react to change notifications, or to periodically compare the entire declared state against reality. Explain the difference between edge-triggered and level-triggered designs, and why reconciliation systems are built the level-triggered way.

level: middleimportance: nice to knowfreq 36%

basics

~20 s

Edge-triggered automation acts on change events; level-triggered automation repeatedly inspects the current state and closes whatever gap it finds. Reconciliation is level-triggered because a missed, duplicated or out-of-order event leaves an edge-triggered system permanently wrong, while a level-triggered one self-heals on the next pass.

open as a page

Ansible keeps no state file, while Terraform and Pulumi record one. How does a tool with no recorded state decide what to do on each run, and what does each approach trade away?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

A stateless tool re-derives reality from the target on every run: each step reads the current condition and acts only if it differs. It has nothing to lose, lock or leak, but it cannot notice that something you deleted from the code should be removed.

open as a page

Why do mature infrastructure-as-code pipelines execute a saved preview artifact from the review step instead of recomputing the diff at apply time, and what does that approach cost?

level: seniorimportance: nice to knowfreq 42%

basics

~20 s

Saving the reviewed diff and executing exactly that artifact guarantees the change a human approved is the change that runs. Recomputing at apply time silently applies a diff nobody saw. The costs are artifact plumbing between jobs, a sensitive file to protect, staleness failures, and a hard version pin.

open as a page

An auditor asks you to demonstrate that no publicly readable storage bucket was created in your production environment over the last twelve months. What can a policy-as-code pipeline offer as evidence, and what does a year of green policy runs genuinely not prove?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A policy pipeline evidences that every change passing through it was evaluated against named, versioned rules, with verdicts, overrides and approvers recorded per commit. It proves nothing about changes made outside that path, or about rules you never wrote.

open as a page

A shared infrastructure module is consumed by a dozen teams. What is a contract test on that module's interface, and what breakage does it catch that the module's own internal tests do not?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A contract test pins the module's published interface — its input names, types, defaults, required-ness and its outputs — and fails when a change breaks a consumer even though the module still works internally. Internal tests only check behaviour, not compatibility.

open as a page

Would you keep every environment's infrastructure code in one repository or give each environment its own repository, and how does that choice affect who can approve production changes?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

Usually one repository, because promotion across repositories degrades into copying and the environments drift. Separate repositories buy a hard permission boundary, but the same boundary is cheaper to get from path-based ownership rules plus per-environment credentials — and the credentials, not the repository, are what actually gate production.

open as a page

You are asked to set up continuous drift detection across a large estate of infrastructure-as-code repositories. How do you design it so it produces action rather than noise?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Scan on a cadence matched to blast radius using read-only credentials, route each finding to the team that owns the code rather than a central channel, suppress known-owned attributes deliberately, and treat drift rate as a process metric rather than auto-reverting.

open as a page

You own an estate that includes stateless API fleets, a self-managed stateful database cluster, CI build agents, and a vendor appliance configurable only through its own console. How would you decide where to mandate immutable replacement and where to keep in-place management, and what does pushing immutability everywhere cost?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Mandate replacement where it is cheap and state-free, and where the fleet is large enough that divergence is unmanageable. Where state, rebuild cost or a vendor blocks replacement, require reproducible builds and detectable divergence instead, as an explicit, audited exception.

open as a page

showing 31–46 of 46