Terraform writes a state file (terraform.tfstate) after every apply. What does that file actually record, and why can't Terraform work from your configuration plus live cloud API queries alone?
answer
- a ledger, not a cache
- addresses on one side, real IDs on the other
- the cloud never learns your resource names
- attributes and dependencies ride along
- no entry means Terraform creates a second one
basics
~20 sTerraform state binds each configuration address, such as aws_instance.web, to the real object ID the provider returned, and stores that object's last-known attributes plus its recorded dependencies. Without the binding, Terraform cannot tell which cloud objects are its own.
solid answer
~50 sState is a JSON snapshot that answers one question the cloud API cannot: which real object does each address in my configuration correspond to. For every managed resource instance it records `mode`, `type`, `name`, the provider it came from, the full `attributes` object returned at the last apply (including the provider-assigned `id`), and a `dependencies` list of the addresses it referenced. Top-level fields carry bookkeeping: the state format `version`, the `terraform_version` that wrote it, a `serial` counter and a `lineage` UUID. Config plus API is not enough because an AWS account does not tell you which of its fifty instances is `aws_instance.web` — the cloud has no notion of your resource addresses, and nothing in the API says "Terraform created this." State is that ledger, which is also why it is the thing you protect.
code
json · 25 lines{
"version": 4,
"terraform_version": "1.9.5",
"serial": 17,
"lineage": "3f1c9a5e-7b2d-4a11-9c2f-6d8e4a0b1c23",
"outputs": {},
"resources": [
{
"mode": "managed",
"type": "aws_instance",
"name": "web",
"provider": "provider[\"registry.terraform.io/hashicorp/aws\"]",
"instances": [
{
"schema_version": 1,
"attributes": {
"id": "i-0abc123def456",
"instance_type": "t3.micro"
},
"dependencies": ["aws_security_group.web"]
}
]
}
]
}go deeper
Be able to say plainly that state maps each resource in your code to the real object Terraform created, and that it is written after apply. Knowing the file is JSON named terraform.tfstate is expected.
Explain the contents: address to provider-assigned ID, the cached attribute object, the recorded dependency list, and the serial and lineage header. Be ready to argue why querying the cloud API cannot replace it.
Show you treat state as the operational asset it is: the reason a lost or mismatched state means duplicates and undeletable objects, and the reason attribute caching makes the file sensitive and shared rather than incidental.
Own the consequence for estate design: the state file defines the unit of ownership and blast radius for a configuration, so where its boundaries fall is an organizational decision about who can change what, not just a file-layout choice.
## What the file is `terraform.tfstate` is a plain JSON document that Terraform writes after every apply (and reads before every plan). It is Terraform's private record of what it has created. Conceptually it is a ledger with one entry per resource instance Terraform manages, plus a small header of bookkeeping fields. A trimmed-down real example: ```json { "version": 4, "terraform_version": "1.9.5", "serial": 17, "lineage": "3f1c9a5e-7b2d-4a11-9c2f-6d8e4a0b1c23", "outputs": {}, "resources": [ { "mode": "managed", "type": "aws_instance", "name": "web", "provider": "provider[\"registry.terraform.io/hashicorp/aws\"]", "instances": [ { "schema_version": 1, "attributes": { "id": "i-0abc123def456", "instance_type": "t3.micro" }, "dependencies": ["aws_security_group.web"] } ] } ] } ``` ## The binding that matters The single most important thing in that document is the pairing of a **resource address** — `aws_instance.web`, or `module.network.aws_subnet.private["a"]` — with a **real object identifier**, here `i-0abc123def456`. Your configuration only speaks in addresses. The cloud only speaks in IDs. Nothing on the AWS side knows that this particular instance is the one your `web` block describes; there is no field on the object saying "managed by Terraform, address aws_instance.web." State is the only place that correspondence exists. That is why the frequent suggestion "why not just query the API and compare?" fails. Querying tells Terraform what exists in the account. It cannot tell Terraform which of those objects it is responsible for, which were created by a colleague's stack, and which were clicked together by hand. Tag conventions are a heuristic people sometimes reach for, but they are not identity — tags are mutable, not universally supported, and would still not distinguish two instances declared by the same module with different keys. ## Cached attributes Each instance also stores the whole attribute object as it stood after the last successful apply, not just the ID. Terraform uses those cached values to compute the diff ("what I recorded" versus "what the config now says"), to resolve references from other resources and from outputs at plan time without an API round trip, and to hold values the API will never hand back — a generated password, for example. `schema_version` records which version of the provider's schema shaped those attributes, so a newer provider can upgrade the stored object rather than misread it. ## Recorded dependency edges The `dependencies` list on each instance is the set of addresses that resource referenced at the last apply. It exists so ordering survives a configuration that no longer describes the resource: if you delete a block, Terraform still has to destroy the object in an order that respects what it depended on, and the config can no longer tell it. ## Bookkeeping `version` is the state **format** version (4 for all of Terraform 0.12 onward), independent of the Terraform release. `terraform_version` records which CLI wrote the snapshot — an older CLI refuses to read state written by a newer one, because the newer one may have upgraded structures it does not understand. `serial` increments on every write, and `lineage` is a UUID minted when the state was first created, so Terraform can tell "a newer version of the same state" from "a completely unrelated state." ## The consequences you are really being asked about Everything downstream follows from the ledger: - **Deletion is only possible through state.** When you remove a block, Terraform knows to destroy `i-0abc123def456` solely because state says that address owned it. No state entry, no destroy — the object simply keeps running and nothing points at it. - **A missing entry means a duplicate.** Terraform sees a declared resource with no recorded object and creates a new one, next to the old one it can no longer see. - **State is sensitive and shared.** Because it caches attribute values, it is not a scratch file; and because it is the shared source of truth for the whole root module, everyone applying that configuration must be reading and writing the same copy. ## The register to answer in A strong answer names the mapping first, then the cached attributes and dependency edges, then the fact that the file is JSON with a documented top-level shape but an internal format you read with Terraform's own commands rather than a text editor. The weak answer describes state as "a cache to make plans faster" — speed is a side effect, identity is the purpose.
- If state already holds the real object ID, why does it also need to record the resource address?The address is the object's identity inside your configuration; the ID is its identity in the cloud. Terraform matches config to state by address, so renaming `aws_instance.web` to `aws_instance.frontend` without moving the state entry makes Terraform see one resource gone and one new resource declared, and it plans a destroy plus a create.
- Are data sources recorded in state as well?Yes. They appear as entries with `mode` set to `"data"`, holding the attributes read at the last refresh so other expressions can resolve against them at plan time. They are re-read rather than created, so Terraform never destroys them; removing the block just drops the entry.
- What is the `outputs` section of state for?It stores the last computed value of each root-module output, so other tooling can read what the configuration produced without running a plan. It is also what a consumer reads when another configuration pulls values out of this state rather than re-querying the provider.
saying these in an interview costs you the question
- State is just a performance cache for faster plans
- Terraform could discover its resources by tagging them
- State stores your HCL configuration files
- Only resource IDs are stored, not attribute values
- Losing state is harmless because the config is the source of truth