IaC Concepts
Every IaC tool is a different dialect of the same handful of ideas: desired state, reconciliation, drift, immutability and a preview before you commit. Learning them once here means Terraform, Pulumi, CloudFormation and Ansible all become variations rather than separate subjects.
on this pageshowhide
explore
- Declarative vs Imperative5 questions
- Desired State and Reconciliation5 questions
- Immutable vs Mutable Infrastructure5 questions
- Plan and Preview Lifecycle5 questions
- Drift Detection and Remediation5 questions
- State Management Theory4 questions
- Composition and Environment Patterns6 questions
- Policy as Code6 questions
- IaC Testing Strategies5 questions
questions
page 2 of 2An infrastructure-as-code preview was generated and reviewed twenty minutes ago, and the apply runs now. What can make the applied result differ from what was reviewed, and how do teams narrow that window?
basics
~20 sA preview describes the world at the moment it was computed, so anything that changes afterwards causes skew: another apply, a console edit, an autoscaler, or values that only resolve during apply. Teams narrow it by serializing applies behind a lock, applying the reviewed artifact, and keeping approval-to-apply short.
Two engineers start an apply at the same time against the same shared IaC state record. What can go wrong, and why is locking the standard answer?
basics
~20 sEach run reads the record, changes real infrastructure, then writes the record back. Interleaved, the later write overwrites the earlier one, so resources the first run created vanish from the record and become unmanaged orphans. A mutual-exclusion lock serialises the write-bearing operation.
Some infrastructure defects survive every cheap check and only surface when the configuration is genuinely applied into a live account. Which classes of failure are those, and why can no plan or mocked test predict them?
basics
~20 sOnly a real apply exposes what the provider decides at request time: quota and limit refusals, missing permissions, argument combinations the schema allows but the API rejects, name collisions, region or capacity unavailability, slow eventual consistency, and dependency-ordered failures including destroys that hang or leave orphans.
You inherit an estate where some infrastructure is reconciled continuously by an agent, some by a nightly job, and some only when an engineer opens a pull request. How do you decide which loop each resource should belong to?
basics
~20 sDecide per resource by comparing the cost of being wrong with the cost of being corrected automatically. Reversible, numerous, cheap-to-recreate resources belong on the fast loop; destructive or data-bearing ones stay human-gated. Above all, give every field exactly one owning loop.
You own the human approval gate for infrastructure-as-code changes across many teams. How do you decide which changes require an approval, and what makes such an approval meaningful rather than a rubber stamp?
basics
~20 sGate on the content of the diff rather than on every change: destructive, stateful, security-relevant and production changes need a human, additive and reversible ones do not. An approval is meaningful only when the approver can read the diff, is accountable for the resource, and can realistically decline.
Any set of infrastructure guardrails eventually meets a change that legitimately has to break one of the rules. How would you design the exemption path so that exceptions remain possible without the guardrail degrading into a formality?
basics
~20 sMake exemptions explicit, narrowly scoped to one rule and one resource, time-bounded with an expiry that re-fails, attributed to a named requester and an independent approver, and visible in aggregate. Blanket skips and permanent unowned exceptions are what hollow out a guardrail.
Most teams refuse to run a full apply-and-destroy integration test on every commit to their infrastructure repository. Why, and how would you decide what runs on a pull request versus nightly versus before a release?
basics
~20 sA real apply costs minutes to hours, real money, and finite account quota, and it fails intermittently for provider reasons. So pull requests get fast, deterministic checks, a scheduled run pays for the expensive real deployment, and pre-release runs the full path once against realistic scale.
If infrastructure is written in a general-purpose programming language such as Python or TypeScript, does that make the approach imperative rather than declarative?
basics
~20 sNo. The authoring language and the deployment model are separate axes. A program with loops and conditionals can still produce a complete description of desired state, which an engine then compares against reality before changing anything.
Infrastructure automation can be written to react to change notifications, or to periodically compare the entire declared state against reality. Explain the difference between edge-triggered and level-triggered designs, and why reconciliation systems are built the level-triggered way.
basics
~20 sEdge-triggered automation acts on change events; level-triggered automation repeatedly inspects the current state and closes whatever gap it finds. Reconciliation is level-triggered because a missed, duplicated or out-of-order event leaves an edge-triggered system permanently wrong, while a level-triggered one self-heals on the next pass.
Ansible keeps no state file, while Terraform and Pulumi record one. How does a tool with no recorded state decide what to do on each run, and what does each approach trade away?
basics
~20 sA stateless tool re-derives reality from the target on every run: each step reads the current condition and acts only if it differs. It has nothing to lose, lock or leak, but it cannot notice that something you deleted from the code should be removed.
Why do mature infrastructure-as-code pipelines execute a saved preview artifact from the review step instead of recomputing the diff at apply time, and what does that approach cost?
basics
~20 sSaving the reviewed diff and executing exactly that artifact guarantees the change a human approved is the change that runs. Recomputing at apply time silently applies a diff nobody saw. The costs are artifact plumbing between jobs, a sensitive file to protect, staleness failures, and a hard version pin.
An auditor asks you to demonstrate that no publicly readable storage bucket was created in your production environment over the last twelve months. What can a policy-as-code pipeline offer as evidence, and what does a year of green policy runs genuinely not prove?
basics
~20 sA policy pipeline evidences that every change passing through it was evaluated against named, versioned rules, with verdicts, overrides and approvers recorded per commit. It proves nothing about changes made outside that path, or about rules you never wrote.
A shared infrastructure module is consumed by a dozen teams. What is a contract test on that module's interface, and what breakage does it catch that the module's own internal tests do not?
basics
~20 sA contract test pins the module's published interface — its input names, types, defaults, required-ness and its outputs — and fails when a change breaks a consumer even though the module still works internally. Internal tests only check behaviour, not compatibility.
Would you keep every environment's infrastructure code in one repository or give each environment its own repository, and how does that choice affect who can approve production changes?
basics
~20 sUsually one repository, because promotion across repositories degrades into copying and the environments drift. Separate repositories buy a hard permission boundary, but the same boundary is cheaper to get from path-based ownership rules plus per-environment credentials — and the credentials, not the repository, are what actually gate production.
You are asked to set up continuous drift detection across a large estate of infrastructure-as-code repositories. How do you design it so it produces action rather than noise?
basics
~20 sScan on a cadence matched to blast radius using read-only credentials, route each finding to the team that owns the code rather than a central channel, suppress known-owned attributes deliberately, and treat drift rate as a process metric rather than auto-reverting.
You own an estate that includes stateless API fleets, a self-managed stateful database cluster, CI build agents, and a vendor appliance configurable only through its own console. How would you decide where to mandate immutable replacement and where to keep in-place management, and what does pushing immutability everywhere cost?
basics
~20 sMandate replacement where it is cheap and state-free, and where the fleet is large enough that divergence is unmanageable. Where state, rebuild cost or a vendor blocks replacement, require reproducible builds and detectable divergence instead, as an explicit, audited exception.
showing 31–46 of 46