skip to content

IaC Concepts

Every IaC tool is a different dialect of the same handful of ideas: desired state, reconciliation, drift, immutability and a preview before you commit. Learning them once here means Terraform, Pulumi, CloudFormation and Ansible all become variations rather than separate subjects.

on this pageshow

questions

page 1 of 2

Why is infrastructure-as-code usually split into reusable components or modules instead of describing the whole estate in one large configuration?

level: juniorimportance: must knowfreq 68%

answer

  1. a function over infrastructure
  2. inputs, outputs, hidden implementation
  3. write the pattern once, call it many times
  4. smaller reviews, faster runs, contained mistakes
  5. one-resource wrappers buy nothing

basics

~20 s

A component packages a pattern once behind a small input interface, so the tenth service is a few lines instead of a hundred. It also shrinks what each change touches: smaller reviews, faster runs, and a mistake that stops at one component instead of the whole estate.

solid answer

~50 s

Three reasons, in the order I would argue them. **Reuse**: the standard shape of a service — network wiring, identity, logging, alarms — is written and reviewed once, and every consumer gets the reviewed version rather than a fresh copy-paste with its own mistakes. **Comprehension and review**: a reviewer reads a change to one named component with a declared interface, not a diff inside a thousand-line file. **Blast radius and speed**: when components are applied as separate units, a change to the service layer computes a diff over service resources only, so it is faster, and a bad change cannot enumerate the network or the database for destruction. The counterweight is that too many tiny components create wiring and ordering work of their own — decomposition is a judgement call, not a rule to maximise.

go deeper

for a junior

Be ready to name reuse and readability, and to describe a component as something with inputs and outputs that hides its implementation. A concrete example — a service component called once per service — lands better than an abstract definition.

for a middle

Explain that the interface is where standards get enforced: what varies becomes an input, what should not vary becomes a baked-in default. Note that code decomposition and apply-unit decomposition are different decisions.

for a senior

Show the counterweight. Argue when not to decompose, describe the maintenance cost of an interface you can never break, and connect apply-unit boundaries to run time, lock contention and containment of a bad change.

for a principal

Own the estate-level version: how a component library is governed, versioned and adopted across teams, who owns breaking changes, and how you avoid a shared library that becomes a bottleneck every team routes around.

## What a component actually is In every infrastructure-as-code tool the same idea appears under a different name — module, construct, component resource, role, stack template. Strip the naming and it is the same thing: **a named unit of infrastructure with a declared set of inputs, a declared set of outputs, and an implementation the caller does not need to read.** It is a function over infrastructure. Everything below follows from that. ## Reason one: reuse of a reviewed pattern An estate contains the same shape many times. "A service" means a compute unit, a load balancer target, an identity with a scoped policy, log retention, a couple of alarms, and consistent tags. Written inline, that is roughly a hundred lines per service, and the twelfth copy will quietly omit the log retention because someone was in a hurry. Written as a component, it is one reviewed implementation and twelve short call sites, each supplying a name, a size and a couple of flags. The subtler benefit is that improvements propagate. When you add a required tag or fix an over-broad permission in the component, every consumer picks it up the next time they take a new version. With copy-paste, you have twelve places to find and a grep that misses two of them. ## Reason two: the interface is where the thinking happens A component forces you to answer "what legitimately varies here?" The answer becomes the input list. What does *not* vary becomes an opinion baked into the implementation — encryption on, public access off, retention at ninety days. That is how organisational standards actually get enforced in practice: not by a document, but by being the default inside the component that everyone calls. A good interface is small and value-shaped: sizes, counts, names, CIDRs, retention. A poor one exposes every underlying knob, at which point the component is not an abstraction, it is a passthrough with extra indirection. ## Reason three: review scope A reviewer's attention is finite. A change described as "bump the service component from 2.3.0 to 2.4.0 in the payments environment" is reviewable: they read the component's changelog and the preview of what will change. A forty-line diff buried in the middle of a monolithic configuration file is reviewed by scrolling, which is not review. ## Reason four: blast radius and run time This reason applies to *separately applied* units rather than merely separately written ones, and the distinction matters. Writing components does not by itself split the recorded state; you also have to apply them as independent units. Once you do, three things improve: each run reads and diffs fewer resources, so it finishes in seconds instead of many minutes; a mistake in one unit produces a plan that cannot even name the resources in another; and two teams can change their own areas concurrently instead of queueing behind one lock. ``` monolith: one run, 900 resources, 11 minutes, everyone waits composed: network | platform | service-a | service-b four runs, minutes each, independent owners ``` ## The honest counterweight Decomposition is not free, and an interviewer will respect you more for saying so. Every boundary you introduce becomes an interface you must version and keep stable, and every cross-boundary dependency becomes wiring — component A's output has to reach component B somehow, and now there is an ordering relationship a human has to know about. A component wrapping a single resource with a one-to-one input list adds indirection and buys nothing; you now read two files to learn what one resource does. The practical heuristic: create a component when the pattern is used more than once, or when it encodes a decision you want enforced. Create a separately applied unit when the pieces have different change rates, different owners, or different consequences of failure. Otherwise leave it inline and split later — merging code is easy, and the split can be done when the pain is real. ## What a weak answer sounds like "Modules keep the files shorter." File length is a symptom, not the reason. Reuse of a *reviewed* pattern, an interface that encodes standards, reviewable change units and contained blast radius are the reasons; shorter files are what you notice on the way past.

  • When is wrapping something in a component the wrong call?
    When it is used once, or when the component's inputs map one-to-one onto the resource it wraps. That is indirection without abstraction — a reader now opens two files to learn what one resource does, and the wrapper's interface has to be maintained forever. Write it inline and extract it the second time you need it.
  • Does splitting code into components automatically reduce blast radius?
    No. Components are a code-organisation move; blast radius follows the unit you apply and the record it diffs against. If ten components are still applied together from one composition against one recorded state, a bad change still computes over everything. Reducing blast radius means applying them as independent units with independent records.

saying these in an interview costs you the question

  • Modules exist mainly to make files shorter.
  • Wrap every single resource in its own module.
  • Expose every underlying option as a module input, just in case.
  • Splitting code into modules automatically shrinks blast radius.
  • One big configuration is fine because the tool works out the order.

context

open as a page

What is the difference between declarative and imperative infrastructure code, and why do mainstream infrastructure-as-code tools choose the declarative model?

level: juniorimportance: must knowfreq 88%

basics

~20 s

Declarative code describes the end state you want and lets the tool work out the steps to reach it; imperative code lists the steps themselves. IaC tools are declarative so the same file can be applied repeatedly and still describe one known result.

open as a page

In infrastructure-as-code, what does it mean for a tool to be convergent, and why is it normally safe to simply re-run the same configuration after a run failed halfway through?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Convergence means each run moves the system toward the declared state and stops when it matches, so a second run with no changes does nothing. A half-finished run is safe to repeat: the tool redoes only the work still outstanding.

open as a page

In an infrastructure-as-code workflow, what does it mean for infrastructure to have "drifted", and what are the most common ways drift appears in a live cloud estate?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Drift is divergence between what your infrastructure code declares and what actually exists at the provider. Common causes: console or emergency edits, autoscalers and other controllers changing fields, provider-applied defaults, and a second tool managing the same resource.

open as a page

In infrastructure operations, what is the difference between mutable and immutable infrastructure, and what happens to a running server under each model when a configuration or package change has to ship?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Mutable infrastructure is changed in place — you patch and reconfigure the servers you already run. Immutable infrastructure never edits a running server: you build a new image, launch replacements from it, and destroy the old instances.

open as a page

An infrastructure-as-code change preview lists actions such as create, update in place, replace, and destroy. What does each mean, and why is replace the one to look at hardest?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Create adds a new object, update in place modifies an existing one, destroy removes it, and replace destroys the existing object and creates a new one because an attribute changed that the provider cannot alter after creation. Replace is destructive: data and identity do not survive it.

open as a page

Infrastructure-as-code tools such as Terraform and Pulumi keep a recorded state of what they created. Why does a tool need that record at all, and what could it not do without one?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A state record maps each declared resource to the real object the tool created. Without it the tool cannot tell an update from a create, cannot know which real objects to delete when code is removed, and cannot order a destroy.

open as a page

How would you lay out an infrastructure-as-code repository so the same infrastructure can be deployed to dev, staging and production?

level: middleimportance: must knowfreq 78%

basics

~20 s

Define each piece of infrastructure once as a reusable component and vary only the inputs per environment. Give every environment its own thin top-level configuration, its own values file and its own isolated state and credentials, so an apply in dev cannot touch production.

open as a page

Why is idempotency the property that makes declarative infrastructure code usable, and what typically makes a hand-written provisioning shell script non-idempotent?

level: middleimportance: must knowfreq 66%

basics

~20 s

Idempotency means applying the same configuration again leaves the same end state, so a run can be repeated or resumed safely. Scripts break it by using create-and-append operations that stack up on every execution instead of asserting a condition.

open as a page

A colleague says an infrastructure-as-code tool "just compares my code to what exists in the cloud". Which three distinct states does a reconciliation run actually involve, and which comparisons does the tool make between them?

level: middleimportance: must knowfreq 72%

basics

~20 s

Three states, not two: desired state (your code), recorded state (the tool's snapshot of the objects it manages) and actual state (what the provider reports now). A run refreshes the record from reality, then diffs desired against it.

open as a page

Walk through the phases an infrastructure-as-code tool goes through when it previews a change, from reading the current world to applying it. What happens in each?

level: middleimportance: must knowfreq 78%

basics

~20 s

A preview runs in four phases: refresh, where the tool re-reads live infrastructure; diff, where it compares your declared configuration against that reality and lists proposed actions; approval, where a human or policy gate reviews them; and apply, which executes the approved actions in dependency order.

open as a page

Policy checks in an infrastructure-as-code pipeline usually run against the proposed plan or diff rather than against the resources that are already running. Why is evaluating the proposed change the decisive design choice, and what does each of the two placements catch that the other misses?

level: middleimportance: must knowfreq 62%

basics

~20 s

Evaluating the proposed change is preventive: the violation is judged before it exists, so the bad resource is never created. Scanning what is already running is detective — it finds violations only after they exist and after the exposure window has opened.

open as a page

Infrastructure-as-Code repositories are usually tested at several levels, from cheap file checks up to real deployments. What are those levels, and what does each one catch that the cheaper level below it cannot?

level: middleimportance: must knowfreq 62%

basics

~20 s

IaC testing layers, cheap to expensive: static checks on the files, assertions on the generated plan or diff, unit tests of modules against faked providers, then a real apply into a throwaway account. Each layer sees more reality at more cost.

open as a page

An engineer fixed a production outage by changing a resource directly in the cloud console overnight, and your infrastructure code no longer matches. Walk through how you decide what to do about it.

level: seniorimportance: must knowfreq 66%

basics

~20 s

First establish what was changed and why, and whether the change is still load-bearing. Then choose one of three honest outcomes: revert to the code, codify the change into the code and apply, or adopt an unmanaged resource into the tool's record. Never re-apply blindly.

open as a page

You are introducing automated policy checks to teams that already ship infrastructure changes daily. What enforcement levels sit between "record the finding" and "block the apply", and how would you sequence the rollout so the guardrails survive contact with delivery pressure?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Three levels are standard: advisory reports only, soft-mandatory fails but a named owner can override with the override recorded, and hard-mandatory cannot be overridden at all. Roll out advisory first to find false positives and the real violation rate, then promote rules one at a time.

open as a page

In infrastructure-as-code work, what does "policy as code" mean, and what does expressing a rule such as "no storage bucket may be publicly readable" as an executable check give you that the same rule written in a standards document does not?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Policy as code expresses organisational rules as machine-executable checks that run automatically on every proposed infrastructure change. Unlike a written standard, the rule is unambiguous, never forgotten by a reviewer, version-controlled like the code it governs, and leaves a pass/fail record.

open as a page

Before infrastructure code is sent to any cloud API, a static check can be run over the files themselves. What kinds of mistakes does that layer catch, and what is it structurally unable to catch?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Static checks read only the source, so they catch syntax and schema errors, unknown or misspelled attributes, style and naming violations, deprecated constructs, hardcoded secrets, and obviously risky settings. They cannot know what infrastructure exists or what a change would actually do.

open as a page

A shared infrastructure component has grown a boolean or environment-name input for nearly every per-environment difference. Why is that a problem, and what would you do instead?

level: middleimportance: should knowfreq 42%

basics

~20 s

Every flag doubles the number of code paths, and only the combinations your environments actually use are ever exercised — so the production path is the least tested one. Prefer parameterising values, pushing structural differences up into the caller, or accepting two explicit definitions.

open as a page

What is the difference between infrastructure as code and configuration as code, and how does that difference show up in the kinds of tools a team ends up running?

level: middleimportance: should knowfreq 50%

basics

~20 s

Infrastructure as code declares the resources themselves — networks, machines, managed services. Configuration as code declares the state inside or on top of them, such as installed packages, files and services. The two overlap but answer different questions.

open as a page

How does an infrastructure-as-code tool actually detect drift — what does it compare against what, and how does that differ between tools that keep a recorded state and tools that do not?

level: middleimportance: should knowfreq 58%

basics

~20 s

Detection re-reads each managed resource from the provider API and compares three things: the configuration you declared, the record of what the tool last created, and the observed live values. A difference between the record and the observation is drift.

open as a page

What is a golden (baked) machine image, and how do you decide what to bake into the image at build time versus what to configure when the instance boots?

level: middleimportance: should knowfreq 52%

basics

~20 s

A golden image is built once by a pipeline with the OS packages, runtime, agents and application already installed. Bake whatever is slow, version-sensitive or downloaded from the network; leave only per-instance and per-environment data to be injected at boot.

open as a page

A configuration-management tool runs successfully against every host every night, yet two hosts that should be identical behave differently. Why does repeated in-place convergence still leave divergent servers, and how does replacing hosts instead of patching them remove that class of problem?

level: middleimportance: should knowfreq 44%

basics

~20 s

A convergence run only asserts what it currently declares. It does not remove what it stopped declaring, does not pin what was never pinned, and cannot undo manual edits, so hosts accumulate divergent history. Replacement discards that history on every change.

open as a page

A policy engine such as Open Policy Agent or HashiCorp Sentinel is normally fed a machine-readable description of the change an infrastructure tool intends to make, rather than the source files the engineer wrote. Why do teams evaluate that representation instead of the source, and what is it still unable to tell the engine?

level: middleimportance: should knowfreq 38%

basics

~20 s

Source files are a template, not an outcome — variables, loops and module inputs are only resolved when the tool computes the change. The engine is given that resolved description so it judges the concrete resources and values that will exist, not the text that produces them.

open as a page

An IaC tool creates a database with a generated password. Why does that plaintext password end up in the tool's recorded state, and what follows for how you store and access that record?

level: middleimportance: should knowfreq 48%

basics

~20 s

The record stores every attribute the provider returned, and a generated password is just another attribute. Any recorded state therefore contains secrets in the clear, so its storage must be encrypted, tightly access-controlled and audited exactly like a secret store.

open as a page

You are breaking a large infrastructure estate into separately applied components. What criteria decide where the boundaries go?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Cut along change cadence, ownership and consequence of failure: things that change together and are owned by the same people belong in one unit, and the seam goes where you want a mistake to stop. Keep cross-boundary dependencies few and pointing one way.

open as a page

A shared infrastructure component is consumed by dev, staging and production. How do you roll a change to that component out across the environments safely?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Publish the component as an immutable version, have every environment pin an exact version, and promote by bumping the pin one environment at a time — dev, then staging, then production — with each bump a reviewed change carrying its own preview of what will change.

open as a page

Your team has adopted a declarative IaC tool. Where does imperative scripting against a cloud SDK or CLI still legitimately win, and how do you keep those scripts from becoming invisible operations?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Scripting still wins for one-shot actions rather than durable resources: operational tasks, bulk data work, bootstrapping the tooling itself, and anything with no resource model. Keep those scripts in the repository, reviewed, re-runnable and logged like any other change.

open as a page

Compare infrastructure that is reconciled only when a person triggers a run with infrastructure reconciled by an agent that runs the same loop continuously. What does the continuous model buy you, and what new failure modes does it introduce?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Continuous reconciliation shrinks the time infrastructure spends wrong from days to minutes and removes the human from the critical path. In exchange, a bad commit reaches production unreviewed, emergency console fixes get reverted automatically, and the loop itself becomes a system you must observe, rate-limit and be able to pause.

open as a page

A resource in your infrastructure code shows drift on the same attribute after every deployment because another system — an autoscaler or a controller — legitimately writes that field. How do you resolve the conflict?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Give the field exactly one owner. Either your code stops managing that attribute and hands it to the other system, or you disable the other writer and keep it declared. Two writers on one field produce endless flapping and eventually ignored drift reports.

open as a page

In an immutable infrastructure model, rollback is often described as simply redeploying the previous image. What property of the model makes that possible, and which parts of a real production system does it fail to restore?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Rollback is cheap because the previous version still exists untouched as an addressable artifact, so you launch it rather than reverse a change. It restores compute and configuration only — not migrated data, consumed messages, external side effects or deleted infrastructure.

open as a page

showing 1–30 of 46