skip to content

HCL & Resources

The language itself — resources, data sources, locals, expressions, and the meta-arguments controlling how many instances exist and how they change. This is where day-to-day Terraform writing happens, so interviewers probe it for depth rather than syntax recall.

part ofTerraformoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

In Terraform, what does a `data` block do, and how is it different from a `resource` block?

level: juniorimportance: must knowfreq 78%

answer

  1. read, never create
  2. one block owns, one block asks
  3. the `data.` prefix in references
  4. deleting it destroys nothing real
  5. re-read on every plan

basics

~20 s

A data block reads information about infrastructure Terraform does not manage and exposes it to the configuration. Terraform never creates, updates or destroys what a data block reads; a resource block declares an object Terraform owns for its whole lifecycle.

solid answer

~50 s

A `resource` block is a declaration of ownership: Terraform will create the object, update it when the configuration changes, and destroy it when you remove the block. A `data` block is read-only — it asks a provider a question, such as "which VPC has this tag" or "what account am I in", and exposes the answer as attributes you reference with the `data.` prefix, for example `data.aws_vpc.main.id`. Terraform re-reads data sources on every plan, so they always reflect the current state of the world rather than a value captured once. Deleting a data block deletes nothing real; it just stops the lookup. The rule of thumb is that you use a data source for things somebody else owns — another team's VPC, a vendor's AMI, a value that exists in the console — and reference the resource directly for anything this configuration creates itself.

code

hcl · 13 lines
hcl
data "aws_caller_identity" "current" {}

data "aws_availability_zones" "available" {
  state = "available"
}

resource "aws_s3_bucket" "artifacts" {
  bucket = "artifacts-${data.aws_caller_identity.current.account_id}"
}

output "first_az" {
  value = data.aws_availability_zones.available.names[0]
}

go deeper

for a junior

Be able to say plainly that a data block reads and a resource block owns, and reference a data source with the data. prefix. Know that deleting a data block destroys nothing.

for a middle

Explain that Terraform re-reads data sources on every plan, that the result is cached in state only as a read result, and that the natural use is the boundary of your configuration — things another team, a vendor or the console owns.

for a senior

Show judgment about where the boundary belongs: an unfiltered lookup that can match many objects or none is a production hazard, and referencing what you create directly is always better than looking it up. Be ready to justify a data source over an input variable.

for a principal

Own the estate-level rule: which facts are looked up at runtime and which are pinned as inputs. Runtime lookups make configurations self-configuring but non-reproducible; pinned inputs make the pipeline responsible for supplying truth. Set that convention deliberately rather than per repository.

## The two kinds of blocks Terraform's configuration language has two blocks that both talk to a provider about real infrastructure, and they mean opposite things. A `resource` block is a claim of ownership. It says: this object should exist, with these arguments, and Terraform is responsible for making that true. Terraform records the object in state, compares the recorded object with the configuration on every plan, and will create, update, replace or destroy it accordingly. Remove the block and the next apply destroys the object. A `data` block — often called a data source — is a question, not a claim. It says: go ask the provider about something that already exists and give me its attributes. Terraform reads it, never writes it. There is no create, no update, no destroy, and removing the block from the configuration has zero effect on real infrastructure. ```hcl resource "aws_instance" "app" { # Terraform owns this instance ami = data.aws_ami.al2023.id instance_type = "t3.micro" } data "aws_ami" "al2023" { # Terraform only reads this image most_recent = true owners = ["amazon"] filter { name = "name" values = ["al2023-ami-*-x86_64"] } } ``` ## Addressing A managed resource is addressed `<type>.<name>.<attribute>` — `aws_instance.app.id`. A data source takes the same shape with a `data.` prefix — `data.aws_ami.al2023.id`. That prefix is the whole reason both can share a type name without colliding, and forgetting it is the single most common beginner error. The same prefix appears in CLI addresses, so `terraform state list` shows entries like `data.aws_ami.al2023`. ## Who provides them Data sources ship inside providers, exactly like resource types do. The AWS provider gives you `aws_caller_identity` (which account and role am I running as, via `account_id` and `arn`), `aws_availability_zones` (the `names` list for the current region), `aws_ami`, `aws_vpc`, `aws_subnets`, and `aws_iam_policy_document`, which builds policy JSON from HCL blocks and exposes it as `json`. A few data sources come from utility providers rather than a cloud: `terraform_remote_state` reads another configuration's outputs, and `external` and `http` are generic escape hatches. Not every data source reads infrastructure. `aws_iam_policy_document` contacts nothing at all — it is a pure document builder that happens to be packaged as a data source. That is worth knowing because it breaks the mental model that "data source means API call". ## What Terraform does with the result The result is cached in state so that plan output can be rendered and so downstream values are stable within a run, but that cache is not authoritative the way a managed resource's state entry is. Terraform re-reads the data source at the start of each plan. If the underlying object changed since last time, the new value flows straight into whatever references it — which is a feature when you want to track something external, and a hazard when the value feeds an argument that forces replacement. Because the read happens during planning, a data source whose arguments are all known gives you concrete values in the plan. If an argument depends on something Terraform has not created yet, the read cannot happen at plan time and is deferred to apply. ## When not to use one The common design mistake is using a data source to look up something the same configuration creates. If `aws_vpc.main` is in this root module, write `aws_vpc.main.id`, not a `data "aws_vpc"` block filtered by tag. The direct reference creates an implicit dependency and always resolves; the lookup has no dependency edge and will run before the object exists on a first apply. Data sources are for the boundary of your configuration, not its interior. The second mistake is treating a data source as a way to run a step. Data sources read; they do not act. A read that has side effects will be executed on every plan, including plans that are never applied.

  • Does a data source appear in the state file, and does `terraform destroy` do anything to it?
    Yes, the read result is recorded in state so plan rendering and downstream references are stable within a run, but it is a cache, not ownership. `terraform destroy` removes the entry from state and makes no API call to delete anything — the underlying object is untouched because Terraform never claimed it.
  • Your team owns a VPC in a separate Terraform configuration. What are your options for referencing it here?
    Either look it up with a data source such as `aws_vpc` filtered by a stable tag, or pass the id in as an input variable that the pipeline supplies. Both keep this configuration from owning the VPC. Which one you choose is a handoff-strategy decision; the data-source option trades an explicit input for a runtime lookup that can match zero or many objects.
  • Is every data source an API call?
    No. `aws_iam_policy_document` builds policy JSON purely locally from its `statement` blocks and exposes it as `json`; nothing is contacted. Terraform packages some pure functions as data sources, so "data source" means "read-only value producer", not necessarily "remote read".

A resource block is a deed of ownership; a data block is a lookup in the public register. Tearing up your copy of the register entry does not demolish the building.

saying these in an interview costs you the question

  • Says removing a data block will destroy the object it read
  • Writes `aws_ami.ubuntu.id` instead of `data.aws_ami.ubuntu.id`
  • Thinks a data source can create the object if it is missing
  • Uses a data source to look up a resource the same config creates
  • Believes the data source value is captured once and frozen

context

open as a page

In Terraform, what does a dynamic block do, and how would you use one to generate a security group's ingress blocks from a variable?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A dynamic block produces repeated nested blocks from a collection: for_each supplies one element per block and the content block holds the block body. For a security group you iterate a list of rule objects to emit one ingress block each.

open as a page

In Terraform HCL, when is the ${ ... } interpolation syntax actually required, and why does writing "${var.name}" as an entire argument value produce a deprecation warning?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Since Terraform 0.12, expressions are first-class, so you write var.name directly. The ${ } syntax is only for embedding an expression inside a larger string, such as "app-${var.env}". Quoting a whole expression adds nothing and Terraform warns that it is deprecated.

open as a page

In Terraform, what is the practical difference between the count and for_each meta-arguments, and why can removing one element from the middle of a list used with count destroy resources you never meant to touch?

level: middleimportance: must knowfreq 78%

basics

~20 s

count addresses instances by numeric position, so deleting a middle list element shifts every later index and Terraform destroys and recreates those resources. for_each addresses instances by map key or set value, which stays stable, so unrelated instances are untouched.

open as a page

In Terraform, when is a data source actually read, and why does a plan sometimes show a data source and everything downstream of it as `(known after apply)`?

level: middleimportance: must knowfreq 56%

basics

~20 s

Terraform reads a data source during plan whenever all of its arguments are already known. If any argument depends on a value that does not exist yet, the read is deferred to apply, its attributes become unknown, and every value derived from them renders as (known after apply).

open as a page

In Terraform, how is the resource dependency graph built from a configuration, and when is an explicit depends_on genuinely required?

level: middleimportance: must knowfreq 78%

basics

~20 s

Terraform infers dependency edges from attribute references — writing aws_vpc.main.id inside another resource creates the edge automatically. depends_on is only for hidden ordering that exists at the provider but is never expressed as a reference in the configuration.

open as a page

What does a Terraform for expression produce, and how do you write one that yields a map instead of a list?

level: middleimportance: must knowfreq 72%

basics

~20 s

A for expression transforms one collection into another. Square brackets produce a list-like tuple: [for s in var.names : upper(s)]. Curly braces with key => value produce a map-like object: {for s in var.names : s => upper(s)}. An optional if clause filters elements out.

open as a page

In Terraform, what does `create_before_destroy = true` in a resource's lifecycle block change about how that resource is replaced, and what must be true of the resource for it to work?

level: middleimportance: must knowfreq 70%

basics

~20 s

By default Terraform destroys a resource before creating its replacement. Setting create_before_destroy = true inverts that order — the new object is created first, then the old one is destroyed — so it only works when both objects can briefly exist at once.

open as a page

In Terraform, an autoscaler keeps changing an attribute that your configuration also sets, so every plan shows a diff. How does the lifecycle block's `ignore_changes` fix that, and what does it cost you?

level: middleimportance: must knowfreq 74%

basics

~20 s

ignore_changes lists attributes whose real-world value Terraform should stop reacting to, so the plan no longer proposes reverting them. The cost is silence: drift on those attributes is never reported again, and the configured value is used only when the object is first created.

open as a page

During terraform apply an instance is created successfully but its remote-exec provisioner fails halfway through. What state is that resource left in, and what does the next apply do?

level: middleimportance: must knowfreq 55%

basics

~20 s

The apply errors and Terraform marks the resource tainted in state: the real object still exists, but Terraform treats it as unusable. The next plan schedules it for destroy and recreate, so the whole provisioner sequence runs again on a fresh object.

open as a page

HashiCorp's documentation calls Terraform provisioners a last resort. What are the concrete reasons for that advice, and what should you reach for instead?

level: middleimportance: must knowfreq 72%

basics

~20 s

Provisioners are opaque to Terraform: their effects are not in state, not visible in the plan, and not idempotent, and they need network reachability plus credentials from wherever Terraform runs. Prefer a baked image, cloud-init/user_data, or a config-management handoff.

open as a page

In Terraform, what does the idiom count = var.enabled ? 1 : 0 on a resource block do, and how do you then reference that resource's attributes elsewhere in the configuration?

level: juniorimportance: should knowfreq 62%

basics

~20 s

It makes the resource conditional: one instance when the flag is true, zero when false. Because the block is still expanded by count, every reference needs an index — aws_s3_bucket.logs[0].id — or a splat wrapped in one() to yield null when the resource is absent.

open as a page

What does `prevent_destroy = true` in a Terraform resource's lifecycle block actually do, and what does it fail to protect against?

level: juniorimportance: should knowfreq 55%

basics

~20 s

prevent_destroy makes Terraform reject any plan that would destroy that resource, with an error instead of a deletion. It is a guardrail in the configuration only: delete the resource block itself and the guard disappears with it.

open as a page

In Terraform, what is the difference between the local-exec and remote-exec provisioners, and what does each one need in order to run?

level: juniorimportance: should knowfreq 58%

basics

~20 s

local-exec runs a command on the machine running Terraform itself. remote-exec runs commands on the resource that was just created, over SSH or WinRM, so it additionally needs a connection block with a reachable host and working credentials.

open as a page

A Terraform plan fails with "Invalid for_each argument: the for_each value depends on resource attributes that cannot be determined until apply". Why does Terraform refuse, and how do you restructure the configuration instead of reaching for -target?

level: middleimportance: should knowfreq 48%

basics

~20 s

Terraform must expand a resource into addressed instances before it can render a diff, so for_each keys have to be known during plan. Fix it by keying on static input — variables, locals, or another resource's own keys — and letting the unknown values sit in each.value.

open as a page

In a Terraform root module, a `data "aws_subnets"` lookup filtered by a tag on a VPC that the same configuration creates fails on the first apply in a fresh account but succeeds when you re-run. Why, and how do you fix it?

level: middleimportance: should knowfreq 38%

basics

~20 s

The filter is a literal string, so Terraform sees no dependency on the VPC and reads the data source during plan, before the VPC exists. The re-run succeeds only because the partial first apply left the objects behind. Reference the created resources directly instead of looking them up.

open as a page

When Terraform destroys a set of resources, what order does it use, and how does it work that order out?

level: middleimportance: should knowfreq 45%

basics

~20 s

Terraform destroys in reverse dependency order, walking the same graph with its edges inverted: nothing is deleted until everything that depends on it is gone. Subnets go before the VPC, because the subnet referenced the VPC's id.

open as a page

In Terraform, what is the iterator argument of a dynamic block for, and what do key and value hold when for_each is a list versus a set?

level: middleimportance: should knowfreq 42%

basics

~20 s

The iterator argument renames the temporary variable inside a dynamic block, which otherwise takes the block's label. For a list, key is the element index and value the element; for a set, key is identical to value and should not be used.

open as a page

In Terraform, which parts of a resource can a dynamic block NOT generate, and why?

level: middleimportance: should knowfreq 36%

basics

~20 s

A dynamic block can only generate repeatable nested blocks defined by the resource, data source, provider or provisioner schema. It cannot produce plain arguments such as tags, nor meta-argument blocks like lifecycle, which Terraform must process before evaluating expressions.

open as a page

What does the splat expression aws_instance.web[*].id evaluate to in Terraform, and when do you have to use a for expression instead?

level: middleimportance: should knowfreq 50%

basics

~20 s

It evaluates to a list of the id attribute of every element in aws_instance.web, in order. A splat only traverses attributes of a list-like value, so once you need filtering, a computed key, or output shaped as a map, you must write a for expression instead.

open as a page

A live Terraform configuration manages 30 EC2 instances with count over a list, and you need to switch it to for_each keyed by instance name without destroying anything. How do you carry out that migration?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Rewrite the block to use for_each, then add a moved block for every instance mapping the old index address to the new key address. Terraform re-addresses the existing objects in state and the plan should report zero adds, changes and destroys.

open as a page

A Terraform configuration selects an EC2 image with `data "aws_ami"` and `most_recent = true`. Why is that a hazard in production, and what would you do instead?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because the data source is re-read on every plan, most_recent = true resolves to whatever image the vendor published last, so an unrelated apply can silently replace running instances. Pin the image id as a reviewed input instead of resolving it at plan time.

open as a page

A terraform apply fails with an error beginning "Error: Cycle:" that names two aws_security_group resources. What has gone wrong, and how do you break it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Two objects reference each other, so Terraform's dependency graph is no longer acyclic and no valid order exists. Break the loop by moving one direction of the relationship into a separate resource that both groups depend on, rather than referencing each other inline.

open as a page

Terraform's documentation calls depends_on a last resort. What does adding one actually cost you compared with an ordinary attribute reference?

level: seniorimportance: should knowfreq 42%

basics

~20 s

An explicit edge is whole-object and untyped: Terraform cannot see which value matters, so it must plan conservatively, marking more attributes as "(known after apply)". It also serialises work that could have run concurrently and hides the real reason for the ordering.

open as a page

When should you avoid a Terraform dynamic block and instead write the nested blocks out literally or declare separate resources?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Write blocks literally when the set is small and known; reach for a dynamic block mainly to hide repetition behind a reusable module's interface. When each item needs its own diff, lifecycle or targeted replacement, prefer separate resources instead.

open as a page

How does Terraform's templatefile() function render a file, and what is allowed to be referenced from inside the template?

level: seniorimportance: should knowfreq 48%

basics

~20 s

templatefile(path, vars) reads a file from disk at plan or apply time and renders it as a string template. Only the names in the vars map are in scope inside the template - it cannot see Terraform variables, locals or resources - and referencing a name you did not pass is an error.

open as a page

In Terraform, what class of errors does the try() function actually catch, and when is lookup() or coalesce() the better choice?

level: seniorimportance: should knowfreq 42%

basics

~20 s

try() evaluates its arguments in order and returns the first that does not raise a dynamic type or traversal error - a missing attribute, a missing index, a failed conversion. It is not a general exception handler. Use lookup() for a missing map key with a default and coalesce() for a value that is present but null or empty.

open as a page

Inside a Terraform resource's `lifecycle` block, what do `precondition` and `postcondition` do, and when is each one evaluated?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Both assert something must be true and fail the run with your own error message if it is not. A precondition is checked before Terraform creates or updates the resource; a postcondition is checked afterwards and can inspect the resource's own attributes through self.

open as a page

Your Terraform module needs `prevent_destroy` on the database in production but not in dev. Why can you not write `prevent_destroy = var.is_prod`, and what do you do instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Terraform processes the lifecycle block before it evaluates expressions, so prevent_destroy and create_before_destroy accept only literal true or false — a variable reference is a configuration error. Environment-specific protection has to come from a provider argument, separate configurations, or pipeline policy.

open as a page

A Terraform config with a remote-exec provisioner works from a developer's laptop but times out when the same code runs on a CI runner. What is going wrong, and how would you fix it properly?

level: seniorimportance: should knowfreq 42%

basics

~20 s

remote-exec connects from wherever Terraform runs, so the CI runner must have a network route to the host, an open port, and valid credentials. The laptop had all three; the runner usually does not. The proper fix is to stop provisioning over the network.

open as a page

showing 1–30 of 38