Infrastructure as Code & Config Management
Describing infrastructure and machine configuration in files that get reviewed and version-controlled instead of clicked together in a console. Interviewers ask about it because reproducibility, drift, and 'who changed prod' are ops questions with tooling answers.
on this pageshowhide
explore
- Terraform (has its own guide)155 questions
- HCL & Resources38 questions
- Providers & Registry16 questions
- State & Backends30 questions
- Variables & Outputs24 questions
- Modules16 questions
- Plan/Apply Workflow31 questions
- Pulumi15 questions
- Languages and Programming Model5 questions
- Stacks, State, and Secrets5 questions
- Pulumi vs Terraform5 questions
- AWS CDK16 questions
- Constructs and Levels5 questions
- Apps, Stacks, and Environments5 questions
- Synth, Bootstrap, and Deploy6 questions
- CloudFormation16 questions
- Template Anatomy6 questions
- Stacks, Change Sets, and Rollback5 questions
- Drift Detection and Import5 questions
- Ansible33 questions
- Inventory and Hosts5 questions
- Playbooks and Tasks6 questions
- Variables and Templating6 questions
- Roles and Reuse6 questions
- Modules and Plugins5 questions
- Vault and Secrets5 questions
- Chef4 questions
- Puppet4 questions
- Salt4 questions
- IaC Concepts46 questions
- Declarative vs Imperative5 questions
- Desired State and Reconciliation5 questions
- Immutable vs Mutable Infrastructure5 questions
- Plan and Preview Lifecycle5 questions
- Drift Detection and Remediation5 questions
- State Management Theory4 questions
- Composition and Environment Patterns6 questions
- Policy as Code6 questions
- IaC Testing Strategies5 questions
→ has its own guide
questions
293 · 9 sectionsIn Terraform, what does a `data` block do, and how is it different from a `resource` block?
basics
~20 sA data block reads information about infrastructure Terraform does not manage and exposes it to the configuration. Terraform never creates, updates or destroys what a data block reads; a resource block declares an object Terraform owns for its whole lifecycle.
In Terraform, what does a dynamic block do, and how would you use one to generate a security group's ingress blocks from a variable?
basics
~20 sA dynamic block produces repeated nested blocks from a collection: for_each supplies one element per block and the content block holds the block body. For a security group you iterate a list of rule objects to emit one ingress block each.
In Terraform HCL, when is the ${ ... } interpolation syntax actually required, and why does writing "${var.name}" as an entire argument value produce a deprecation warning?
basics
~20 sSince Terraform 0.12, expressions are first-class, so you write var.name directly. The ${ } syntax is only for embedding an expression inside a larger string, such as "app-${var.env}". Quoting a whole expression adds nothing and Terraform warns that it is deprecated.
In Terraform, how do you create resources in two AWS regions from a single configuration, and how does an individual resource choose the second region?
basics
~20 sDeclare a second aws provider block with an alias argument and its own region, then set the meta-argument provider = aws.<alias> on each resource that belongs in that region. Resources with no provider argument use the unaliased default block.
In a Terraform aws provider block, why should you not set access_key and secret_key, and where does the provider get credentials instead?
basics
~20 sKeys written into HCL get committed, cloned and copied, and cannot be rotated per run. Omit access_key and secret_key; the AWS provider then resolves credentials itself from environment variables, the shared AWS config files, or an attached instance or CI role.
In a Pulumi program, why does reading a property off a resource — say a bucket's arn — give you an Output<T> rather than a plain string, and how do you build other values from it?
basics
~20 sPulumi runs your program before the cloud has created anything, so resource attributes are Output<T> placeholders for values known only after deployment. Derive new values with apply(), pulumi.all() or pulumi.interpolate — never by concatenating an Output as a string.
In Pulumi, what is a stack, and how does one Pulumi program deploy separate dev, staging and production environments?
basics
~20 sA Pulumi stack is one independently configurable instance of a program, with its own config file, its own state and its own outputs. Running the same program against dev, staging and prod stacks gives three isolated deployments and no duplicated code.
What is a Pulumi ComponentResource, and what must a ComponentResource subclass do in its constructor for the component to behave correctly?
basics
~20 sA ComponentResource is a logical grouping of other resources exposed as one reusable class — Pulumi's module equivalent. Its constructor must call super with a type token and name, create every child with the parent option set to itself, and finish with registerOutputs.
How does Pulumi keep secret values out of plaintext in stack config and in state, and what is a stack's "secrets provider"?
basics
~20 sPulumi encrypts values you mark as secret before they are written anywhere. pulumi config set --secret stores ciphertext in Pulumi.<stack>.yaml, secretness propagates into the state checkpoint, and the secrets provider — a passphrase, a cloud KMS key, or the Pulumi Cloud per-stack key — holds the encryption key.
Pulumi lets you define infrastructure in TypeScript, Python, Go or C#, while Terraform uses HCL. What do you actually gain, and what do you give up, by choosing a general-purpose language?
basics
~20 sA general-purpose language buys real control flow, static types, IDE completion and unit tests over infrastructure logic. It costs reviewability: a pull request now shows the program that produces resources rather than the resources themselves.
In the AWS CDK, what are an App, a Stack and a Construct, and which of the three corresponds to something AWS actually creates when you deploy?
basics
~20 sConstructs are the composable building blocks of the tree; a Stack is the deployment unit that synthesizes into exactly one CloudFormation stack; the App is the root object holding every stack. Only stacks become real AWS deployments.
In the AWS CDK, what are L1, L2 and L3 constructs, how do you recognise an L1 class in code, and what do you give up by dropping to one?
basics
~20 sAWS CDK constructs come in three levels: L1 Cfn classes mirror CloudFormation resources one-to-one with no defaults, L2 constructs add sensible defaults and helper methods such as the grant helpers, and L3 patterns assemble many resources into one opinionated component.
In the AWS CDK, what does the `cdk synth` command actually do, and what does it leave behind in the cdk.out directory?
basics
~10 scdk synth runs your CDK program and writes a cloud assembly into cdk.out: one CloudFormation template per stack, asset manifests, and staged asset files. It deploys nothing and changes no AWS resources.
You define an AWS CDK stack without passing the `env` property. Where does it deploy, what do `stack.account` and `stack.region` resolve to, and what stops working?
basics
~20 sAn environment-agnostic stack deploys wherever the CDK CLI's credentials point. Its account and region become unresolved tokens backed by CloudFormation pseudo-parameters, so synth-time lookups such as Vpc.fromLookup fail and the stack assumes only two availability zones.
In an AWS CDK app, what does calling bucket.grantRead(myFunction) actually put into the synthesized CloudFormation template, and why is that preferred over hand-writing the policy document?
basics
~20 sCalling grantRead adds a policy statement to the Lambda function's execution role in the synthesized template, scoped to that bucket's ARN and its objects, and it also grants decrypt on the bucket's encryption key when the CDK app knows about one.
In AWS CloudFormation, what exactly does stack drift detection compare, and what does CloudFormation do about the drift it finds?
basics
~20 sDrift detection compares a stack's live resource configuration against the values its template and parameters expect, then reports the differences. It is read-only and on demand: CloudFormation never reverts drift for you and never watches resources continuously.
In an AWS CloudFormation template, what is the difference between Ref and Fn::GetAtt when applied to a resource, and how do you know which one gives you that resource's ARN?
basics
~20 sRef returns one default value chosen by the resource type — usually its name or physical ID — while Fn::GetAtt returns a specific named attribute of that resource. Only the resource type's documentation says which value each one produces.
In AWS CloudFormation, what is a change set, and which part of the described change set tells you a resource will be destroyed and recreated rather than updated in place?
basics
~20 sA change set is a stored, named preview of what an AWS CloudFormation stack update would do, computed without applying anything. Its Replacement field is the thing to read: True means the resource is deleted and recreated with a new physical ID.
An AWS CloudFormation stack update fails and the stack ends up in UPDATE_ROLLBACK_FAILED. What does that status actually mean, and how do you get the stack back to a state where you can deploy again?
basics
~20 sUPDATE_ROLLBACK_FAILED means CloudFormation could not restore the stack's previous state, so it refuses further updates. Fix whatever blocked the rollback, then call ContinueUpdateRollback — skipping unrecoverable resources only as a last resort, since that leaves records inaccurate.
Your first `aws cloudformation create-stack` call fails and the stack now shows the status ROLLBACK_COMPLETE. Why can you neither retry the create nor update it, and what do you do next?
basics
~10 sROLLBACK_COMPLETE means the creation failed and CloudFormation deleted everything it had made, leaving an empty stack record that only accepts deletion. Delete the stack, fix what the failure event reported, and create it again.
What is an Ansible inventory, and what do the built-in `all` and `ungrouped` groups contain?
basics
~20 sAn Ansible inventory is the list of managed hosts and the groups holding them, written as static INI or YAML files or produced by a dynamic plugin. Every host is automatically in the group all; a host in no other group is also in ungrouped.
What is the difference between Ansible's command and shell modules, and which should you reach for by default?
basics
~20 sAnsible's command module executes a program directly with no shell, so pipes, redirects, globs and environment-variable expansion do not work. The shell module runs the command line through a shell on the target, so they do. Prefer command unless you need shell features.
In an Ansible role created by `ansible-galaxy init`, what does each standard directory (tasks, handlers, defaults, vars, files, templates, meta) hold, and which files does Ansible load automatically?
basics
~10 sAn Ansible role is a fixed directory tree - tasks/, handlers/, defaults/, vars/, files/, templates/, meta/. Ansible auto-loads main.yml from each of those directories, and resolves copy and template sources against files/ and templates/.
In an Ansible playbook, what does the Jinja2 expression `{{ myvar | default('fallback') }}` do, and how do `default('fallback', true)` and `default(omit)` differ from it?
basics
~20 sThe default filter supplies a value only when the variable is undefined. Passing true as a second argument also replaces empty or otherwise falsy values. Passing the special omit value removes the parameter from the task entirely, so the module applies its own default.
What is Ansible Vault, and what does encrypting a file with it protect you against — and what does it not protect you against?
basics
~20 sAnsible Vault symmetrically encrypts variable files or single values with a password so they can be committed to git. Ansible decrypts them in memory at run time. It protects secrets at rest only — anyone holding the password reads everything.
In a Chef recipe, `package 'nginx'` is a resource declaration rather than a shell command. What does Chef actually do with that declaration, and how does it stay idempotent across repeated runs?
basics
~20 sChef adds the resource to the run's resource collection with a desired state. During converge, the matching provider inspects the node, compares current state to desired, and acts only when they differ — so a second run changes nothing.
Chef Infra Client executes a run in two phases, compile and converge. What happens in each, and why does plain Ruby written inside a recipe so often run earlier than its author expected?
basics
~20 sCompile evaluates every recipe as Ruby and builds an ordered resource collection; converge then walks that collection and lets each provider act. Bare Ruby in a recipe runs during compile — before any resource has taken effect — so it sees the node's pre-run state.
A Chef cookbook lists `depends` entries in its metadata.rb. How do those dependency versions get resolved and pinned for a node, and what did Policyfiles change compared with Berkshelf plus environment version constraints?
basics
~20 sBerkshelf resolves the metadata.rb constraints into a Berksfile.lock and uploads the cookbooks, while the Chef Infra Server pins versions per environment and expands the run list at run time. Policyfiles replace both with one artefact that locks the run list and every cookbook version together.
You inherit a fleet of long-lived Linux servers configured by a large set of Chef cookbooks, maintained by a team fluent in Ruby. How do you decide whether to keep converging with chef-client or move that configuration somewhere else?
basics
~20 sJudge the cookbooks, not the tool's reputation: whether they are declarative and idempotent, whether anyone left can debug a compile-versus-converge bug, and whether the Chef Infra Server plus agent estate is worth operating. Working, quiet cookbooks are a poor migration candidate.
Walk through a single Puppet agent run, from the agent waking up on its interval to the node actually changing. Where is the catalog compiled, and what goes into it?
basics
~20 sOn each run interval the Puppet agent collects Facter facts and sends them to the Puppet Server, which compiles a node-specific catalog from the manifests and Hiera data for that node. The agent applies that catalog locally, then posts a report back.
In a Puppet manifest, does the order you declare resources decide the order they are applied? How would you make a service restart when its config file changes?
basics
~20 sPuppet applies a dependency graph, not the file top to bottom. Ordering comes from require, before, notify and subscribe, the chaining arrows, and autorequire; notify or subscribe also sends a refresh event, which is what restarts a service after its config changes.
Puppet enforces configuration with an agent on each node that pulls a catalog on an interval, rather than an operator pushing over SSH the way Ansible does. What do you gain and what do you give up by choosing the pull model?
basics
~20 sPull buys continuous, unattended enforcement: every node re-applies its catalog on an interval, so manual changes get corrected and rebuilt nodes converge without anyone driving them. The cost is an agent and certificate per node, propagation latency, and no ordered cross-host orchestration.
How do you find out what Puppet would change on a node or fleet without letting it change anything, and what are the limits of that report?
basics
~20 sRun the agent in no-op mode — puppet agent -t --noop, or noop = true in puppet.conf. Puppet still compiles the catalog and inspects every managed resource, but reports out-of-sync ones instead of correcting them. It says nothing about resources the catalog does not declare.
How does Salt's master/minion architecture work, and why can a command like `salt '*' cmd.run 'uptime'` come back from thousands of hosts in seconds?
basics
~20 sSalt minions hold a persistent ZeroMQ connection out to the master. A command is published once onto that bus, every minion that matches the target expression evaluates it locally and runs it, then returns its own result, so fan-out time barely grows with host count.
In Salt, what is the difference between grains and pillars, and which of the two should hold a database password?
basics
~20 sGrains are facts the minion discovers about itself (OS, CPU, IP) and reports upward. Pillars are data the master compiles and sends down to specific minions. Secrets belong in pillar, because pillar is master-controlled and delivered only to the minions it targets.
You have roughly 5,000 servers to keep configured and you also need ad-hoc commands across them. What does adopting Salt's persistent minion agent buy you over a purely SSH-push approach, and what does it cost?
basics
~20 sThe agent buys constant-time fan-out and event-driven reaction: connections are established once, so a command is one publish rather than 5,000 SSH handshakes, and minions can push events the master reacts to. It costs an agent lifecycle, key management and a master that is root-everywhere.
In Salt, how do beacons and the reactor turn something happening on a minion into an automated response, and what goes wrong when the reaction is heavy or self-triggering?
basics
~20 sA beacon runs on the minion, watches something local — a service, a file, disk usage — and fires a tagged event onto Salt's event bus. The master's reactor matches event tags to reactor SLS files and issues the response, which is what makes Salt the event-driven configuration tool.
Why is infrastructure-as-code usually split into reusable components or modules instead of describing the whole estate in one large configuration?
basics
~20 sA component packages a pattern once behind a small input interface, so the tenth service is a few lines instead of a hundred. It also shrinks what each change touches: smaller reviews, faster runs, and a mistake that stops at one component instead of the whole estate.
What is the difference between declarative and imperative infrastructure code, and why do mainstream infrastructure-as-code tools choose the declarative model?
basics
~20 sDeclarative code describes the end state you want and lets the tool work out the steps to reach it; imperative code lists the steps themselves. IaC tools are declarative so the same file can be applied repeatedly and still describe one known result.
In infrastructure-as-code, what does it mean for a tool to be convergent, and why is it normally safe to simply re-run the same configuration after a run failed halfway through?
basics
~20 sConvergence means each run moves the system toward the declared state and stops when it matches, so a second run with no changes does nothing. A half-finished run is safe to repeat: the tool redoes only the work still outstanding.
In an infrastructure-as-code workflow, what does it mean for infrastructure to have "drifted", and what are the most common ways drift appears in a live cloud estate?
basics
~20 sDrift is divergence between what your infrastructure code declares and what actually exists at the provider. Common causes: console or emergency edits, autoscalers and other controllers changing fields, provider-applied defaults, and a second tool managing the same resource.
In infrastructure operations, what is the difference between mutable and immutable infrastructure, and what happens to a running server under each model when a configuration or package change has to ship?
basics
~20 sMutable infrastructure is changed in place — you patch and reconfigure the servers you already run. Immutable infrastructure never edits a running server: you build a new image, launch replacements from it, and destroy the old instances.