skip to content

Infrastructure as Code & Config Management

Describing infrastructure and machine configuration in files that get reviewed and version-controlled instead of clicked together in a console. Interviewers ask about it because reproducibility, drift, and 'who changed prod' are ops questions with tooling answers.

on this pageshow

explore

→ has its own guide

questions

293 · 9 sections

In Terraform, what does a `data` block do, and how is it different from a `resource` block?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A data block reads information about infrastructure Terraform does not manage and exposes it to the configuration. Terraform never creates, updates or destroys what a data block reads; a resource block declares an object Terraform owns for its whole lifecycle.

open as a page

In Terraform, what does a dynamic block do, and how would you use one to generate a security group's ingress blocks from a variable?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A dynamic block produces repeated nested blocks from a collection: for_each supplies one element per block and the content block holds the block body. For a security group you iterate a list of rule objects to emit one ingress block each.

open as a page

In Terraform HCL, when is the ${ ... } interpolation syntax actually required, and why does writing "${var.name}" as an entire argument value produce a deprecation warning?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Since Terraform 0.12, expressions are first-class, so you write var.name directly. The ${ } syntax is only for embedding an expression inside a larger string, such as "app-${var.env}". Quoting a whole expression adds nothing and Terraform warns that it is deprecated.

open as a page

In Terraform, how do you create resources in two AWS regions from a single configuration, and how does an individual resource choose the second region?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Declare a second aws provider block with an alias argument and its own region, then set the meta-argument provider = aws.<alias> on each resource that belongs in that region. Resources with no provider argument use the unaliased default block.

open as a page

In a Terraform aws provider block, why should you not set access_key and secret_key, and where does the provider get credentials instead?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Keys written into HCL get committed, cloned and copied, and cannot be rotated per run. Omit access_key and secret_key; the AWS provider then resolves credentials itself from environment variables, the shared AWS config files, or an attached instance or CI role.

open as a page

In a Pulumi program, why does reading a property off a resource — say a bucket's arn — give you an Output<T> rather than a plain string, and how do you build other values from it?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Pulumi runs your program before the cloud has created anything, so resource attributes are Output<T> placeholders for values known only after deployment. Derive new values with apply(), pulumi.all() or pulumi.interpolate — never by concatenating an Output as a string.

open as a page

In Pulumi, what is a stack, and how does one Pulumi program deploy separate dev, staging and production environments?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A Pulumi stack is one independently configurable instance of a program, with its own config file, its own state and its own outputs. Running the same program against dev, staging and prod stacks gives three isolated deployments and no duplicated code.

open as a page

What is a Pulumi ComponentResource, and what must a ComponentResource subclass do in its constructor for the component to behave correctly?

level: middleimportance: must knowfreq 50%
basics
~20 s

A ComponentResource is a logical grouping of other resources exposed as one reusable class — Pulumi's module equivalent. Its constructor must call super with a type token and name, create every child with the parent option set to itself, and finish with registerOutputs.

open as a page

How does Pulumi keep secret values out of plaintext in stack config and in state, and what is a stack's "secrets provider"?

level: middleimportance: must knowfreq 60%
basics
~20 s

Pulumi encrypts values you mark as secret before they are written anywhere. pulumi config set --secret stores ciphertext in Pulumi.<stack>.yaml, secretness propagates into the state checkpoint, and the secrets provider — a passphrase, a cloud KMS key, or the Pulumi Cloud per-stack key — holds the encryption key.

open as a page

Pulumi lets you define infrastructure in TypeScript, Python, Go or C#, while Terraform uses HCL. What do you actually gain, and what do you give up, by choosing a general-purpose language?

level: middleimportance: must knowfreq 75%
basics
~20 s

A general-purpose language buys real control flow, static types, IDE completion and unit tests over infrastructure logic. It costs reviewability: a pull request now shows the program that produces resources rather than the resources themselves.

open as a page

In the AWS CDK, what are an App, a Stack and a Construct, and which of the three corresponds to something AWS actually creates when you deploy?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Constructs are the composable building blocks of the tree; a Stack is the deployment unit that synthesizes into exactly one CloudFormation stack; the App is the root object holding every stack. Only stacks become real AWS deployments.

open as a page

In the AWS CDK, what are L1, L2 and L3 constructs, how do you recognise an L1 class in code, and what do you give up by dropping to one?

level: juniorimportance: must knowfreq 78%
basics
~20 s

AWS CDK constructs come in three levels: L1 Cfn classes mirror CloudFormation resources one-to-one with no defaults, L2 constructs add sensible defaults and helper methods such as the grant helpers, and L3 patterns assemble many resources into one opinionated component.

open as a page

In the AWS CDK, what does the `cdk synth` command actually do, and what does it leave behind in the cdk.out directory?

level: juniorimportance: must knowfreq 70%
basics
~10 s

cdk synth runs your CDK program and writes a cloud assembly into cdk.out: one CloudFormation template per stack, asset manifests, and staged asset files. It deploys nothing and changes no AWS resources.

open as a page

You define an AWS CDK stack without passing the `env` property. Where does it deploy, what do `stack.account` and `stack.region` resolve to, and what stops working?

level: middleimportance: must knowfreq 68%
basics
~20 s

An environment-agnostic stack deploys wherever the CDK CLI's credentials point. Its account and region become unresolved tokens backed by CloudFormation pseudo-parameters, so synth-time lookups such as Vpc.fromLookup fail and the stack assumes only two availability zones.

open as a page

In an AWS CDK app, what does calling bucket.grantRead(myFunction) actually put into the synthesized CloudFormation template, and why is that preferred over hand-writing the policy document?

level: middleimportance: must knowfreq 62%
basics
~20 s

Calling grantRead adds a policy statement to the Lambda function's execution role in the synthesized template, scoped to that bucket's ARN and its objects, and it also grants decrypt on the bucket's encryption key when the CDK app knows about one.

open as a page

In AWS CloudFormation, what exactly does stack drift detection compare, and what does CloudFormation do about the drift it finds?

level: juniorimportance: must knowfreq 50%
basics
~20 s

Drift detection compares a stack's live resource configuration against the values its template and parameters expect, then reports the differences. It is read-only and on demand: CloudFormation never reverts drift for you and never watches resources continuously.

open as a page

In an AWS CloudFormation template, what is the difference between Ref and Fn::GetAtt when applied to a resource, and how do you know which one gives you that resource's ARN?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Ref returns one default value chosen by the resource type — usually its name or physical ID — while Fn::GetAtt returns a specific named attribute of that resource. Only the resource type's documentation says which value each one produces.

open as a page

In AWS CloudFormation, what is a change set, and which part of the described change set tells you a resource will be destroyed and recreated rather than updated in place?

level: middleimportance: must knowfreq 68%
basics
~20 s

A change set is a stored, named preview of what an AWS CloudFormation stack update would do, computed without applying anything. Its Replacement field is the thing to read: True means the resource is deleted and recreated with a new physical ID.

open as a page

An AWS CloudFormation stack update fails and the stack ends up in UPDATE_ROLLBACK_FAILED. What does that status actually mean, and how do you get the stack back to a state where you can deploy again?

level: seniorimportance: must knowfreq 52%
basics
~20 s

UPDATE_ROLLBACK_FAILED means CloudFormation could not restore the stack's previous state, so it refuses further updates. Fix whatever blocked the rollback, then call ContinueUpdateRollback — skipping unrecoverable resources only as a last resort, since that leaves records inaccurate.

open as a page

Your first `aws cloudformation create-stack` call fails and the stack now shows the status ROLLBACK_COMPLETE. Why can you neither retry the create nor update it, and what do you do next?

level: juniorimportance: should knowfreq 58%
basics
~10 s

ROLLBACK_COMPLETE means the creation failed and CloudFormation deleted everything it had made, leaving an empty stack record that only accepts deletion. Delete the stack, fix what the failure event reported, and create it again.

open as a page

What is an Ansible inventory, and what do the built-in `all` and `ungrouped` groups contain?

level: juniorimportance: must knowfreq 80%
basics
~20 s

An Ansible inventory is the list of managed hosts and the groups holding them, written as static INI or YAML files or produced by a dynamic plugin. Every host is automatically in the group all; a host in no other group is also in ungrouped.

open as a page

What is the difference between Ansible's command and shell modules, and which should you reach for by default?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Ansible's command module executes a program directly with no shell, so pipes, redirects, globs and environment-variable expansion do not work. The shell module runs the command line through a shell on the target, so they do. Prefer command unless you need shell features.

open as a page

In an Ansible role created by `ansible-galaxy init`, what does each standard directory (tasks, handlers, defaults, vars, files, templates, meta) hold, and which files does Ansible load automatically?

level: juniorimportance: must knowfreq 72%
basics
~10 s

An Ansible role is a fixed directory tree - tasks/, handlers/, defaults/, vars/, files/, templates/, meta/. Ansible auto-loads main.yml from each of those directories, and resolves copy and template sources against files/ and templates/.

open as a page

In an Ansible playbook, what does the Jinja2 expression `{{ myvar | default('fallback') }}` do, and how do `default('fallback', true)` and `default(omit)` differ from it?

level: juniorimportance: must knowfreq 58%
basics
~20 s

The default filter supplies a value only when the variable is undefined. Passing true as a second argument also replaces empty or otherwise falsy values. Passing the special omit value removes the parameter from the task entirely, so the module applies its own default.

open as a page

What is Ansible Vault, and what does encrypting a file with it protect you against — and what does it not protect you against?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Ansible Vault symmetrically encrypts variable files or single values with a password so they can be committed to git. Ansible decrypts them in memory at run time. It protects secrets at rest only — anyone holding the password reads everything.

open as a page

In a Chef recipe, `package 'nginx'` is a resource declaration rather than a shell command. What does Chef actually do with that declaration, and how does it stay idempotent across repeated runs?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Chef adds the resource to the run's resource collection with a desired state. During converge, the matching provider inspects the node, compares current state to desired, and acts only when they differ — so a second run changes nothing.

open as a page

Chef Infra Client executes a run in two phases, compile and converge. What happens in each, and why does plain Ruby written inside a recipe so often run earlier than its author expected?

level: middleimportance: should knowfreq 58%
basics
~20 s

Compile evaluates every recipe as Ruby and builds an ordered resource collection; converge then walks that collection and lets each provider act. Bare Ruby in a recipe runs during compile — before any resource has taken effect — so it sees the node's pre-run state.

open as a page

A Chef cookbook lists `depends` entries in its metadata.rb. How do those dependency versions get resolved and pinned for a node, and what did Policyfiles change compared with Berkshelf plus environment version constraints?

level: seniorimportance: should knowfreq 42%
basics
~20 s

Berkshelf resolves the metadata.rb constraints into a Berksfile.lock and uploads the cookbooks, while the Chef Infra Server pins versions per environment and expands the run list at run time. Policyfiles replace both with one artefact that locks the run list and every cookbook version together.

open as a page

You inherit a fleet of long-lived Linux servers configured by a large set of Chef cookbooks, maintained by a team fluent in Ruby. How do you decide whether to keep converging with chef-client or move that configuration somewhere else?

level: principalimportance: nice to knowfreq 28%
basics
~20 s

Judge the cookbooks, not the tool's reputation: whether they are declarative and idempotent, whether anyone left can debug a compile-versus-converge bug, and whether the Chef Infra Server plus agent estate is worth operating. Working, quiet cookbooks are a poor migration candidate.

open as a page

Walk through a single Puppet agent run, from the agent waking up on its interval to the node actually changing. Where is the catalog compiled, and what goes into it?

level: middleimportance: must knowfreq 75%
basics
~20 s

On each run interval the Puppet agent collects Facter facts and sends them to the Puppet Server, which compiles a node-specific catalog from the manifests and Hiera data for that node. The agent applies that catalog locally, then posts a report back.

open as a page

In a Puppet manifest, does the order you declare resources decide the order they are applied? How would you make a service restart when its config file changes?

level: juniorimportance: should knowfreq 66%
basics
~20 s

Puppet applies a dependency graph, not the file top to bottom. Ordering comes from require, before, notify and subscribe, the chaining arrows, and autorequire; notify or subscribe also sends a refresh event, which is what restarts a service after its config changes.

open as a page

Puppet enforces configuration with an agent on each node that pulls a catalog on an interval, rather than an operator pushing over SSH the way Ansible does. What do you gain and what do you give up by choosing the pull model?

level: seniorimportance: should knowfreq 58%
basics
~20 s

Pull buys continuous, unattended enforcement: every node re-applies its catalog on an interval, so manual changes get corrected and rebuilt nodes converge without anyone driving them. The cost is an agent and certificate per node, propagation latency, and no ordered cross-host orchestration.

open as a page

How do you find out what Puppet would change on a node or fleet without letting it change anything, and what are the limits of that report?

level: seniorimportance: nice to knowfreq 42%
basics
~20 s

Run the agent in no-op mode — puppet agent -t --noop, or noop = true in puppet.conf. Puppet still compiles the catalog and inspects every managed resource, but reports out-of-sync ones instead of correcting them. It says nothing about resources the catalog does not declare.

open as a page

How does Salt's master/minion architecture work, and why can a command like `salt '*' cmd.run 'uptime'` come back from thousands of hosts in seconds?

level: middleimportance: must knowfreq 62%
basics
~20 s

Salt minions hold a persistent ZeroMQ connection out to the master. A command is published once onto that bus, every minion that matches the target expression evaluates it locally and runs it, then returns its own result, so fan-out time barely grows with host count.

open as a page

In Salt, what is the difference between grains and pillars, and which of the two should hold a database password?

level: juniorimportance: should knowfreq 58%
basics
~20 s

Grains are facts the minion discovers about itself (OS, CPU, IP) and reports upward. Pillars are data the master compiles and sends down to specific minions. Secrets belong in pillar, because pillar is master-controlled and delivered only to the minions it targets.

open as a page

You have roughly 5,000 servers to keep configured and you also need ad-hoc commands across them. What does adopting Salt's persistent minion agent buy you over a purely SSH-push approach, and what does it cost?

level: principalimportance: should knowfreq 34%
basics
~20 s

The agent buys constant-time fan-out and event-driven reaction: connections are established once, so a command is one publish rather than 5,000 SSH handshakes, and minions can push events the master reacts to. It costs an agent lifecycle, key management and a master that is root-everywhere.

open as a page

In Salt, how do beacons and the reactor turn something happening on a minion into an automated response, and what goes wrong when the reaction is heavy or self-triggering?

level: seniorimportance: nice to knowfreq 30%
basics
~20 s

A beacon runs on the minion, watches something local — a service, a file, disk usage — and fires a tagged event onto Salt's event bus. The master's reactor matches event tags to reactor SLS files and issues the response, which is what makes Salt the event-driven configuration tool.

open as a page

Why is infrastructure-as-code usually split into reusable components or modules instead of describing the whole estate in one large configuration?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A component packages a pattern once behind a small input interface, so the tenth service is a few lines instead of a hundred. It also shrinks what each change touches: smaller reviews, faster runs, and a mistake that stops at one component instead of the whole estate.

open as a page

What is the difference between declarative and imperative infrastructure code, and why do mainstream infrastructure-as-code tools choose the declarative model?

level: juniorimportance: must knowfreq 88%
basics
~20 s

Declarative code describes the end state you want and lets the tool work out the steps to reach it; imperative code lists the steps themselves. IaC tools are declarative so the same file can be applied repeatedly and still describe one known result.

open as a page

In infrastructure-as-code, what does it mean for a tool to be convergent, and why is it normally safe to simply re-run the same configuration after a run failed halfway through?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Convergence means each run moves the system toward the declared state and stops when it matches, so a second run with no changes does nothing. A half-finished run is safe to repeat: the tool redoes only the work still outstanding.

open as a page

In an infrastructure-as-code workflow, what does it mean for infrastructure to have "drifted", and what are the most common ways drift appears in a live cloud estate?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Drift is divergence between what your infrastructure code declares and what actually exists at the provider. Common causes: console or emergency edits, autoscalers and other controllers changing fields, provider-applied defaults, and a second tool managing the same resource.

open as a page

In infrastructure operations, what is the difference between mutable and immutable infrastructure, and what happens to a running server under each model when a configuration or package change has to ship?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Mutable infrastructure is changed in place — you patch and reconfigure the servers you already run. Immutable infrastructure never edits a running server: you build a new image, launch replacements from it, and destroy the old instances.

open as a page