In Terraform, what does a `data` block do, and how is it different from a `resource` block?
answer
- read, never create
- one block owns, one block asks
- the `data.` prefix in references
- deleting it destroys nothing real
- re-read on every plan
basics
~20 sA data block reads information about infrastructure Terraform does not manage and exposes it to the configuration. Terraform never creates, updates or destroys what a data block reads; a resource block declares an object Terraform owns for its whole lifecycle.
solid answer
~50 sA `resource` block is a declaration of ownership: Terraform will create the object, update it when the configuration changes, and destroy it when you remove the block. A `data` block is read-only — it asks a provider a question, such as "which VPC has this tag" or "what account am I in", and exposes the answer as attributes you reference with the `data.` prefix, for example `data.aws_vpc.main.id`. Terraform re-reads data sources on every plan, so they always reflect the current state of the world rather than a value captured once. Deleting a data block deletes nothing real; it just stops the lookup. The rule of thumb is that you use a data source for things somebody else owns — another team's VPC, a vendor's AMI, a value that exists in the console — and reference the resource directly for anything this configuration creates itself.
code
hcl · 13 linesdata "aws_caller_identity" "current" {}
data "aws_availability_zones" "available" {
state = "available"
}
resource "aws_s3_bucket" "artifacts" {
bucket = "artifacts-${data.aws_caller_identity.current.account_id}"
}
output "first_az" {
value = data.aws_availability_zones.available.names[0]
}go deeper
Be able to say plainly that a data block reads and a resource block owns, and reference a data source with the data. prefix. Know that deleting a data block destroys nothing.
Explain that Terraform re-reads data sources on every plan, that the result is cached in state only as a read result, and that the natural use is the boundary of your configuration — things another team, a vendor or the console owns.
Show judgment about where the boundary belongs: an unfiltered lookup that can match many objects or none is a production hazard, and referencing what you create directly is always better than looking it up. Be ready to justify a data source over an input variable.
Own the estate-level rule: which facts are looked up at runtime and which are pinned as inputs. Runtime lookups make configurations self-configuring but non-reproducible; pinned inputs make the pipeline responsible for supplying truth. Set that convention deliberately rather than per repository.
## The two kinds of blocks Terraform's configuration language has two blocks that both talk to a provider about real infrastructure, and they mean opposite things. A `resource` block is a claim of ownership. It says: this object should exist, with these arguments, and Terraform is responsible for making that true. Terraform records the object in state, compares the recorded object with the configuration on every plan, and will create, update, replace or destroy it accordingly. Remove the block and the next apply destroys the object. A `data` block — often called a data source — is a question, not a claim. It says: go ask the provider about something that already exists and give me its attributes. Terraform reads it, never writes it. There is no create, no update, no destroy, and removing the block from the configuration has zero effect on real infrastructure. ```hcl resource "aws_instance" "app" { # Terraform owns this instance ami = data.aws_ami.al2023.id instance_type = "t3.micro" } data "aws_ami" "al2023" { # Terraform only reads this image most_recent = true owners = ["amazon"] filter { name = "name" values = ["al2023-ami-*-x86_64"] } } ``` ## Addressing A managed resource is addressed `<type>.<name>.<attribute>` — `aws_instance.app.id`. A data source takes the same shape with a `data.` prefix — `data.aws_ami.al2023.id`. That prefix is the whole reason both can share a type name without colliding, and forgetting it is the single most common beginner error. The same prefix appears in CLI addresses, so `terraform state list` shows entries like `data.aws_ami.al2023`. ## Who provides them Data sources ship inside providers, exactly like resource types do. The AWS provider gives you `aws_caller_identity` (which account and role am I running as, via `account_id` and `arn`), `aws_availability_zones` (the `names` list for the current region), `aws_ami`, `aws_vpc`, `aws_subnets`, and `aws_iam_policy_document`, which builds policy JSON from HCL blocks and exposes it as `json`. A few data sources come from utility providers rather than a cloud: `terraform_remote_state` reads another configuration's outputs, and `external` and `http` are generic escape hatches. Not every data source reads infrastructure. `aws_iam_policy_document` contacts nothing at all — it is a pure document builder that happens to be packaged as a data source. That is worth knowing because it breaks the mental model that "data source means API call". ## What Terraform does with the result The result is cached in state so that plan output can be rendered and so downstream values are stable within a run, but that cache is not authoritative the way a managed resource's state entry is. Terraform re-reads the data source at the start of each plan. If the underlying object changed since last time, the new value flows straight into whatever references it — which is a feature when you want to track something external, and a hazard when the value feeds an argument that forces replacement. Because the read happens during planning, a data source whose arguments are all known gives you concrete values in the plan. If an argument depends on something Terraform has not created yet, the read cannot happen at plan time and is deferred to apply. ## When not to use one The common design mistake is using a data source to look up something the same configuration creates. If `aws_vpc.main` is in this root module, write `aws_vpc.main.id`, not a `data "aws_vpc"` block filtered by tag. The direct reference creates an implicit dependency and always resolves; the lookup has no dependency edge and will run before the object exists on a first apply. Data sources are for the boundary of your configuration, not its interior. The second mistake is treating a data source as a way to run a step. Data sources read; they do not act. A read that has side effects will be executed on every plan, including plans that are never applied.
- Does a data source appear in the state file, and does `terraform destroy` do anything to it?Yes, the read result is recorded in state so plan rendering and downstream references are stable within a run, but it is a cache, not ownership. `terraform destroy` removes the entry from state and makes no API call to delete anything — the underlying object is untouched because Terraform never claimed it.
- Your team owns a VPC in a separate Terraform configuration. What are your options for referencing it here?Either look it up with a data source such as `aws_vpc` filtered by a stable tag, or pass the id in as an input variable that the pipeline supplies. Both keep this configuration from owning the VPC. Which one you choose is a handoff-strategy decision; the data-source option trades an explicit input for a runtime lookup that can match zero or many objects.
- Is every data source an API call?No. `aws_iam_policy_document` builds policy JSON purely locally from its `statement` blocks and exposes it as `json`; nothing is contacted. Terraform packages some pure functions as data sources, so "data source" means "read-only value producer", not necessarily "remote read".
A resource block is a deed of ownership; a data block is a lookup in the public register. Tearing up your copy of the register entry does not demolish the building.
saying these in an interview costs you the question
- Says removing a data block will destroy the object it read
- Writes `aws_ami.ubuntu.id` instead of `data.aws_ami.ubuntu.id`
- Thinks a data source can create the object if it is missing
- Uses a data source to look up a resource the same config creates
- Believes the data source value is captured once and frozen